Measure and improve shell-to-English accuracy - #18
Merged
Merged
Conversation
…ld-out gold Teacher/judge default to claude-haiku-4-5; teacher-v2 passes the parser's structure and names, which lifts agreement with the Opus gold from 74.0% to 78.9% (Opus's own second candidates score 72.0%). Jev Choice matches the judge only 52% of the time, so selection stays with the judge model. Candidates are keyed per teacher model, judge dedup hashes candidate text, and test/validation labels are frozen in datasets/gold_lock.json. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
… out of DPO Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Jev returns one of the options it is given, so the student samples 4-5 candidates and Jev picks: 77.4% accepted versus 69.1% greedy on the v1 test set (oracle 83.9%). Adds labeling/self_label.py (free labelling gated at 0.30 => 66% coverage at 89.6% precision, rest queued for a paid pass), --best-of N in the service, and self-labels as training-only dataset rows. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Noisemaker111
force-pushed
the
fix/live-status-accuracy-20260920
branch
from
September 21, 2026 00:29
2026695 to
aaa941c
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The existing evaluation made shell-to-English quality look better than the evidence supports: deterministic target recall covers only 46 of 587 held-out commands, and interrupted remote grading discarded completed local inference. This change adds a newest-first Hugging Face benchmark sweep, structured parser-assisted training and inference, whole-set lexical regression reporting, resumable generation checkpoints, and head-tail retention for long commands.
A structured LoRA run on LFM2.5-350M produced a Q8 model measured at 437 MB residency, 90 ms p50, and 425 output tokens/s versus 655 MB, 105 ms, and 335 tokens/s for the current Qwen3 0.6B Q4 model. Its whole-set lexical F1 is lower (0.418 versus 0.525), so the candidate is recorded but not promoted pending independent semantic action and invention grading.
Validation:
uvx pytest live-status/tests -q(16 passed); Python compileall; Git diff check; local inference across all 587 held-out rows for merged HF, Q8_0, and Q4_K_M exports.