Skip to content

Benchmark Live Status on recent session commands - #21

Merged
Noisemaker111 merged 2 commits into
agentsfrom
feature/recent-session-benchmark-20260920
Sep 21, 2026
Merged

Noisemaker111 merged 2 commits into
agentsfrom
feature/recent-session-benchmark-20260920

Conversation

@Noisemaker111

Copy link
Copy Markdown
Owner

Adds a resumable benchmark for redacted, novel commands from recent Codex sessions; supports explicit local or gateway judges with blinded candidate identities and confidence intervals. Documents the 120-row out-of-time result: the current LoRA remains the deployment choice, DPO is statistically tied, and both regress sharply on complex recent PowerShell.\n\nVerification:\n- python -m unittest discover -s live-status/tests -p 'test_*.py' -v\n- python -m unittest discover -s experiments/command_model -p test_bindings.py -v\n- python -m unittest discover -s experiments/command_model -p test_contract.py -v\n- public benchmark CLI completed and saved/reopened 120 generated rows plus 120 OpenCode Go judgments\n- Ollama unloaded after judging

@Noisemaker111
Noisemaker111 merged commit d327d62 into agents Sep 21, 2026
1 check passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant