Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
28 changes: 19 additions & 9 deletions live-status/ARCHITECTURE.md
Original file line number Diff line number Diff line change
Expand Up @@ -8,15 +8,17 @@
Cursor, Grok, │
PSReadLine) parsers/shell.py ─► structure, difficulty tags ───────────┤
▼
labeling/teacher.py (Opus 5, two candidates per command, batched, resumable) ──► labels/teacher.jsonl
judging/judge.py (Opus 5, separate prompt, scores teacher + heuristic candidates,
labeling/teacher.py (Haiku 4.5, two candidates per command, batched, resumable) ──► labels/teacher.jsonl
labeling/self_label.py (student samples k candidates, Jev selects and gates; no paid model)
──► labels/self_labels.jsonl
judging/judge.py (Haiku 4.5, separate prompt, scores teacher + heuristic candidates,
writes recommended_output) ──► labels/judged.jsonl
dataset_build/build.py (accept / review / reject; MinHash families; frozen test) ──► datasets/<v>/*.jsonl
│
benchmarks/baseline.py ◄── untuned tiny models through Ollama ◄─────────────────────────────┤
training/train.py ──► models/<name> (merged HF weights) ◄────────────────────────────┘
training/export_gguf.py──► GGUF f16/q8_0 + Ollama quantized tags
evaluation/evaluate.py ──► validators + Opus 5 grading + latency/memory ──► evaluation/registry.json
evaluation/evaluate.py ──► validators + jev/Opus grading + latency/memory ──► evaluation/registry.json
evaluation/active.py ──► student failures on unlabeled commands ──► teacher/judge ──► prefs.jsonl (DPO)

client ──► api/server.py ──► redact ─► cache ─► model (Ollama) ─► validate ─► heuristic fallback
Expand All @@ -38,12 +40,20 @@
- **Heuristic as fallback, not fast path.** The deterministic describer answers in under a millisecond, but its confidence ≥ 0.9 outputs cover only 8.7% of executions and the judge rated just 19% of them ≥ 80 (they are correct but generic: "Reviewing the Git diff." for `git diff --stat`). The service therefore always asks the model and uses the heuristic when the model fails, times out or produces an invalid sentence; `--fast-path` re-enables the shortcut.
- **Plain completion format.** The student learns `Command:\n…\n\nStatus: <sentence>` with no system prompt, so each request costs only the command's tokens.
- **Ollama/llama.cpp for serving.** It already runs on this machine, serves GGUF at every quantization level, and keeps models warm. `llama-server` is supported by the same backend interface.
- **jev for grading, Opus for writing.** jev (TypeSafe System One) returns only probabilities,
choices and scores, in milliseconds, at $0.042 per million input tokens. Given the gold
status it agrees with Opus evaluations at AUC 0.92, so it grades every benchmark and ranks
student outputs for failure mining. Without a reference it is too weak (AUC 0.68) to
replace the Opus judge when labels are created.
- **Opus 5 as both teacher and judge**, with different prompts. Label diversity comes from two candidates per command plus the heuristic, not from different models.
- **Jev selects what a local model wrote.** Jev cannot generate, but a Choice or Score returns
one of the options it was handed, so the student samples 4-5 candidates locally and Jev picks
one. On the v1 test set that beats greedy decoding by 8 points (77.4% vs 69.1%; oracle 83.9%),
which powers both `self_label` (free labelling, no paid model) and the service's `--best-of`
quality mode. Selection only helps when the candidates actually differ: choosing between two
near-identical teacher sentences matched the judge just 52% of the time.
- **Haiku writes, jev grades, Opus audits.** Teacher and judge run on Haiku 4.5 with different
prompts; label diversity comes from two candidates per command plus the heuristic. jev
(TypeSafe System One) returns only probabilities, choices and scores, in milliseconds, at
$0.042 per million input tokens: given the gold status it agrees with Opus evaluations at
AUC 0.92, so it grades every benchmark and ranks student outputs for failure mining. It
cannot replace the judge, because without a reference it reaches only AUC 0.68 and its own
Choice between near-equal candidates matches the judge just 52% of the time. Opus is kept
for spot checks and for the calibration sets both graders are measured against.
- **Frequency weighting.** Each template group counts `min(4, 1 + log2(count))` times in training, so common patterns are learned first without drowning the long tail.

## Adding a transcript format
Expand Down
37 changes: 33 additions & 4 deletions live-status/DATASET.md
Original file line number Diff line number Diff line change
Expand Up @@ -42,13 +42,42 @@ ordinary text look secret-shaped).

## Labels

- Teacher: Opus 5, `teacher-v1`, temperature 0.4, two candidates per command (concise and
complete), batches of ≤20 commands / 30k characters.
- Judge: Opus 5, `judge-v1`, temperature 0, scores each candidate plus the heuristic,
picks the best and writes `recommended_output`.
- Teacher: `teacher-v2`, temperature 0.4, two candidates per command (concise and complete),
batches of ≤20 commands / 30k characters. Items carry the parser's `structure` and `names`
so the model keeps concrete names instead of "the script".
- Judge: `judge-v1`, temperature 0, scores each candidate plus the heuristic, picks the best
and writes `recommended_output`.
- Decision: accepted when the recommended score ≥ 85, validators pass and the judge is not
uncertain; manual review at 70–84, uncertain, or validator failure; rejected for secret
leakage or score < 70. The judge rewrites weak candidates, so v1 has no rejections.
- Model: **Haiku 4.5** by default (`LIVE_STATUS_TEACHER`). v1's labels were written by Opus 5;
rows record `teacher_model` and `judge_model`, and candidate keys are prefixed per model
(`ota` = Opus teacher candidate a, `hta` = Haiku), so mixed provenance stays traceable.

### Why Haiku

On 300 test commands that already had Opus labels, Haiku relabelled from scratch and its
labels were graded against the Opus gold:

| Teacher | Accepted vs Opus gold | Same actions | Invented |
| --- | ---: | ---: | ---: |
| Haiku, `teacher-v1` prompt | 74.0% | 97.6% | 4.5% |
| Haiku, `teacher-v2` prompt (+structure, +names) | **78.9%** | 96.7% | 3.7% |
| Opus's own second-choice candidate (reference point) | 72.0% | — | — |

Haiku with the structure hints matches Opus's own alternate candidates, at a fraction of the
cost and without exhausting the Claude subscription that the interactive session shares.
Opus stays available (`--model claude-opus-5`) for spot checks.

### Self-labels (no paid model)

`labeling/self_label.py` has the fine-tuned student write 4-5 candidates for a command and
Jev select one, keeping the label when the selector's score clears 0.30. Measured against gold
on the v1 test set, that gate keeps 66% of commands at 89.6% precision (0.25 → 77% at 87.7%,
0.35 → 55% at 90.0%; `evaluation/self_label_threshold.json`). Kept rows join **training only**,
carry `label_source: "self"`, and are dropped when their family belongs to a held-out split.
Rejected commands land in `labels/needs_teacher.txt` for a paid pass, so the teacher only sees
what the local loop could not label.

## v1 splits

Expand Down
22 changes: 19 additions & 3 deletions live-status/EVALUATION.md
Original file line number Diff line number Diff line change
@@ -1,7 +1,7 @@
# Evaluation

```powershell
python live-status/cli.py evaluate --backend ollama:<model>[:plain|:long|:instruct][:cpu] --name <run> [--grader jev|opus|both|none] [--promote]
python live-status/cli.py evaluate --backend ollama:<model>[:plain|:long|:instruct|:structured|:structured-plain][:cpu] --name <run> [--grader jev|opus|both|none] [--promote]
python live-status/cli.py evaluate --regrade --name <run> --grader opus # re-grade saved outputs
```

Expand All @@ -20,19 +20,24 @@ recall for simple read/process/delete commands.
| Grader | What it sees | Cost / speed | Agreement with Opus 5 |
| --- | --- | --- | --- |
| `opus` | command, reference, output; returns correct, score, missing/hallucinated actions, secret leak, injection followed | ~20 outputs per call, ~60 s | — |
| `opus` on Haiku's labels | same | same | Haiku labels reach 78.9% of the Opus gold standard (DATASET.md) |
| `jev` (default) | command, reference, output; answers `same_actions`, `invented`, `quality` | 700 outputs in ~3 s, ~$0.02 | AUC 0.92, 86.4% agreement at score ≥ 0.45 (700 Opus-graded outputs) |

A jev-accepted output passes validators and has
`same × (1 − invented) × quality/4 ≥ 0.45`. Calibration lives in
`evaluation/jev_eval_calibration.json`; rerun it with `python live-status/judging/jev.py calibrate-eval`.
Grading a status without a reference is much weaker (AUC 0.68 on 3,000 teacher candidates), so
jev only *ranks* unlabeled outputs during failure mining and Opus writes the labels.
Grading a status without a reference is much weaker (AUC 0.68 on 3,000 teacher candidates), and
asking jev to *choose* between near-equal candidates matches the judge only 52% of the time
(1,200 commands, `judging/jev.py calibrate-choice`). So jev ranks unlabeled outputs during
failure mining, and a judge model still writes the labels.

jev is stricter than Opus on good outputs: Opus's alternate teacher candidates on 200 test
commands pass 86% of Opus evaluations and 72% of jev evaluations. Compare runs only within
one grader.

Both graders cache verdicts by (command, output), so re-evaluating an unchanged output is free.
Generation checkpoints every 25 rows to `evaluation/outputs/<run>.generated.jsonl`. A retry
resumes from that checkpoint, and a completed run removes it after saving the final output.

The `no_secret` validator also fails outputs that echo a redaction placeholder
(`<TOKEN>`, `<SECRET>`). Two of the round-1 runs did this once each, which blocks promotion;
Expand All @@ -44,6 +49,17 @@ A run is promoted into `evaluation/registry.json` only if nothing leaks, its acc
least the current best, hallucination does not rise by more than a point, and it was graded
by the same grader as the current best.

## Service under load

```powershell
python live-status/benchmarks/service_load.py --url http://127.0.0.1:8765 --clients 8 --requests 240
```

Qwen3-0.6B q4_K_M behind the service, 8 concurrent clients, cache bypassed: 10.0 requests/s,
p50 0.80 s, p99 1.00 s, no failures, 239/240 answered by the model and one by the fallback.
Single-client latency is 0.10 s p50; the gap is queueing, since one Ollama runner serves the
default concurrency of 4.

## Reports

Each report includes validator rates, accepted %, invented/hallucination %, omission % (Opus),
Expand Down
21 changes: 14 additions & 7 deletions live-status/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -7,9 +7,10 @@ Get-Process opencode2,node,powershell -ErrorAction SilentlyContinue | Select-Obj
→ Checking running opencode2, node, and powershell processes.
```

Current best: **Qwen3-0.6B + LoRA, GGUF q4_K_M (397 MB)**, 67.8% of held-out statuses
accepted by the grader (the teacher's own second choices score 72%), 100 ms p50 on an
RTX 3070 and 393 ms on CPU only. Untuned models of the same size score at most 23.5%.
Current best: **Qwen3-0.6B + LoRA, GGUF q4_K_M (397 MB)**, 68.8% of held-out statuses
accepted by the grader (the teacher's own second choices score 72%), 105 ms p50 on an
RTX 3070 and 393 ms on CPU only; behind the service, 10 requests/s with 8 clients. Untuned
models of the same size score at most 23.5%.
Details: [TRAINING.md](TRAINING.md), [EVALUATION.md](EVALUATION.md),
[DATASET.md](DATASET.md), [ARCHITECTURE.md](ARCHITECTURE.md).

Expand All @@ -20,6 +21,9 @@ python live-status/cli.py serve --backend ollama:live-status-v1-qwen3-06b-lora-q
python live-status/api/client.py "git fetch origin && git status -sb"
```

`--best-of 4` samples several candidates and lets jev pick the best: +8 points of quality for
~1 s per request instead of ~0.1 s.

`POST /v1/summarize-command` with `{"command": "...", "shell": "powershell", "cwd": "optional"}`
returns `{"status": "..."}` (add `"debug": true` for source and latency). The service redacts
before inference, caches by normalised command, bounds concurrency, falls back to a
Expand All @@ -35,8 +39,9 @@ Run from the repository root. GPU steps use the training venv (see TRAINING.md).
| --- | --- |
| Inventory transcript sources | `python live-status/cli.py inventory_sources` |
| Extract, redact, deduplicate, report | `python live-status/cli.py extract_commands` |
| Teacher labels (Opus 5) | `python live-status/cli.py generate_labels --limit 4000` |
| Judge labels (Opus 5) | `python live-status/cli.py judge_labels` |
| Self-label (student + jev, free) | `python live-status/cli.py self_label --backend ollama:... --pool 8000` |
| Teacher labels (Haiku) | `python live-status/cli.py generate_labels --limit 4000` |
| Judge labels (Haiku) | `python live-status/cli.py judge_labels` |
| Build splits | `python live-status/cli.py build_dataset --version v1` |
| Baseline tiny models | `python live-status/cli.py benchmark_base_models --models smollm2:135m qwen3:0.6b` |
| Train | `live-status/cli.py train --base Qwen/Qwen3-0.6B --method lora --name ...` |
Expand All @@ -52,8 +57,10 @@ commands stay in its `private/` folder.

## Models and keys

- Teacher and judge: `claude-opus-5` through the local CLIProxyAPI (`127.0.0.1:8317`).
It shares the Claude subscription; keep `--workers` at 4 or below. On a long cooldown
- Teacher and judge: `claude-haiku-4-5-20251001` through the local CLIProxyAPI
(`127.0.0.1:8317`), overridable with `--model` or `LIVE_STATUS_TEACHER`. Haiku labels match
Opus's own second-choice candidates (DATASET.md), so Opus is reserved for spot checks.
These share the Claude subscription; keep `--workers` at 4 or below. On a long cooldown
the run stops and resumes on rerun.
- Grader: TypeSafe `jev` through Vercel AI Gateway (`AI_GATEWAY_API_KEY` in the repo
`.env`) or directly (`TYPESAFE_API_KEY`). Needs Bun; `bun install` in `live-status/jev`.
Expand Down
26 changes: 22 additions & 4 deletions live-status/TRAINING.md
Original file line number Diff line number Diff line change
Expand Up @@ -18,6 +18,8 @@ Status: <gold sentence><eos>

Loss covers only the status and EOS. `--prompt instruct` prepends the long instruction
(serve it with the `:long` backend suffix); `mixed` uses it on 30% of rows.
`--prompt structured` adds a deterministic action parse and bounded raw command; serve that
adapter with the `:structured-plain` backend suffix.

## Mechanics

Expand Down Expand Up @@ -62,14 +64,30 @@ For scale: Opus's own alternate (non-selected) candidates for the same 200 test
score 72.0% accepted under jev (86% under the Opus evaluator), so Qwen3-0.6B at 69.4% is
close to the teacher's second-choice quality under the same grader.

## Round 2: failure mining, more data, DPO (v1 test grown to 583 commands)

The 89 commands added to test by round-1 mining are the previous model's own failures, so
scores on the enlarged set are lower and only comparable within it.

| Checkpoint | Train rows | Accepted | Invented | Same actions | p50 |
| --- | ---: | ---: | ---: | ---: | ---: |
| Qwen3-0.6B q4_K_M (round 1) | 3,113 | 65.9% | 6.3% | 94.3% | 100 ms |
| + round-1 corrections (`v1b`) | 3,628 | **68.8%** | 6.0% | 94.3% | 105 ms |
| + DPO on 517 preference pairs | 3,628 | 68.6% | 6.3% | **96.1%** | 107 ms |

Mining the model's own failures is worth ~3 points. DPO left acceptance flat while raising
same-action agreement and the mean grade (0.548 → 0.559); it is kept as an optional stage.
Preference pairs whose command later landed in test or validation are dropped automatically
(148 of 665 here).

## Iteration loop

```
build_dataset -> train -> export_gguf -> evaluate (jev) -> mine_failures -> build_dataset ...
```

`mine_failures` runs the current student over unlabeled commands, ranks its outputs with
jev, sends the weakest 40% (plus a random fifth as many) to the Opus teacher and judge,
and appends (chosen, rejected) pairs to `labels/prefs.jsonl`. Round r1 (SmolLM2-135M, 1,500
commands) sent 726 to the teacher; the judge stopped after 104 when the Opus subscription
entered a cooldown, and resumes on rerun.
jev, sends the weakest 40% (plus a random fifth as many) to the teacher and judge, and
appends (chosen, rejected) pairs to `labels/prefs.jsonl`. Round r1 (SmolLM2-135M, 1,500
commands) sent 726 to the teacher; judging finished on Haiku after the Opus subscription
hit a cooldown. `mine_failures --rebuild-prefs` re-derives pairs from existing judgements.
Loading
Loading