Skip to content

Measure and improve shell-to-English accuracy - #18

Merged
Noisemaker111 merged 6 commits into
agentsfrom
fix/live-status-accuracy-20260920
Sep 21, 2026
Merged

Noisemaker111 merged 6 commits into
agentsfrom
fix/live-status-accuracy-20260920

Conversation

@Noisemaker111

Copy link
Copy Markdown
Owner

The existing evaluation made shell-to-English quality look better than the evidence supports: deterministic target recall covers only 46 of 587 held-out commands, and interrupted remote grading discarded completed local inference. This change adds a newest-first Hugging Face benchmark sweep, structured parser-assisted training and inference, whole-set lexical regression reporting, resumable generation checkpoints, and head-tail retention for long commands.

A structured LoRA run on LFM2.5-350M produced a Q8 model measured at 437 MB residency, 90 ms p50, and 425 output tokens/s versus 655 MB, 105 ms, and 335 tokens/s for the current Qwen3 0.6B Q4 model. Its whole-set lexical F1 is lower (0.418 versus 0.525), so the candidate is recorded but not promoted pending independent semantic action and invention grading.

Validation: uvx pytest live-status/tests -q (16 passed); Python compileall; Git diff check; local inference across all 587 held-out rows for merged HF, Q8_0, and Q4_K_M exports.

Noisemaker111 and others added 6 commits September 20, 2026 20:29
…ld-out gold

Teacher/judge default to claude-haiku-4-5; teacher-v2 passes the parser's structure and
names, which lifts agreement with the Opus gold from 74.0% to 78.9% (Opus's own second
candidates score 72.0%). Jev Choice matches the judge only 52% of the time, so selection
stays with the judge model. Candidates are keyed per teacher model, judge dedup hashes
candidate text, and test/validation labels are frozen in datasets/gold_lock.json.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
… out of DPO

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Jev returns one of the options it is given, so the student samples 4-5 candidates and Jev
picks: 77.4% accepted versus 69.1% greedy on the v1 test set (oracle 83.9%). Adds
labeling/self_label.py (free labelling gated at 0.30 => 66% coverage at 89.6% precision,
rest queued for a paid pass), --best-of N in the service, and self-labels as training-only
dataset rows.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@Noisemaker111
Noisemaker111 force-pushed the fix/live-status-accuracy-20260920 branch from 2026695 to aaa941c Compare September 21, 2026 00:29
@Noisemaker111
Noisemaker111 merged commit 06941e0 into agents Sep 21, 2026
1 check passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant