Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
4 changes: 3 additions & 1 deletion .github/workflows/agent-checks.yml
Original file line number Diff line number Diff line change
Expand Up @@ -23,7 +23,9 @@ jobs:
with:
python-version: '3.12'
- name: Validate syntax
run: python -m compileall -q scripts skills experiments/command_model
run: python -m compileall -q scripts skills experiments/command_model live-status
- name: Live-status redaction and service boundary invariants
run: python -m unittest discover -s live-status/tests -t live-status -v
- name: Path binding and saved evidence invariants
run: python -m unittest discover -s experiments/command_model -p test_bindings.py -v
- name: Native and PowerShell contract invariants
Expand Down
3 changes: 3 additions & 0 deletions .gitignore
Original file line number Diff line number Diff line change
Expand Up @@ -5,3 +5,6 @@ work/
.worktrees/
experiments/command_model/js_parser/node_modules/
*.log
.env
.env.*
live-status/jev/node_modules/
1 change: 1 addition & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -19,6 +19,7 @@ scopes; the purpose document preserves what each actually demonstrated.
| Execution verification and datasets | Check candidate behavior and save examples with frozen held-out partitions |
| Model training | Train for the actual delegation protocol, preserving original adapters |
| Local execution loop and frontier evaluation | Execute grounded English jobs and measure verified whole-operation accuracy, time and tokens |
| [Live status](live-status/README.md) | Tiny local model and service that turn commands into one-sentence live status text |

The gatherer does not certify training examples. See the
[data pipeline](experiments/command_model/DATA_PIPELINE.md) and
Expand Down
51 changes: 51 additions & 0 deletions live-status/ARCHITECTURE.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,51 @@
# Live status architecture

```
transcripts ──► data_miner ──► private/records.jsonl ──► dedupe ──► private/commands.jsonl
(Claude Code, (sources.py raw + redacted template │ raw representative
Codex, OpenCode, one parser groups, ▼
OpenCode2, per format) counts commands_redacted.jsonl ──┐
Cursor, Grok, │
PSReadLine) parsers/shell.py ─► structure, difficulty tags ───────────┤
▼
labeling/teacher.py (Opus 5, two candidates per command, batched, resumable) ──► labels/teacher.jsonl
judging/judge.py (Opus 5, separate prompt, scores teacher + heuristic candidates,
writes recommended_output) ──► labels/judged.jsonl
dataset_build/build.py (accept / review / reject; MinHash families; frozen test) ──► datasets/<v>/*.jsonl
│
benchmarks/baseline.py ◄── untuned tiny models through Ollama ◄─────────────────────────────┤
training/train.py ──► models/<name> (merged HF weights) ◄────────────────────────────┘
training/export_gguf.py──► GGUF f16/q8_0 + Ollama quantized tags
evaluation/evaluate.py ──► validators + Opus 5 grading + latency/memory ──► evaluation/registry.json
evaluation/active.py ──► student failures on unlabeled commands ──► teacher/judge ──► prefs.jsonl (DPO)

client ──► api/server.py ──► redact ─► cache ─► model (Ollama) ─► validate ─► heuristic fallback
```

## Boundaries

| Boundary | Rule | Enforced by |
| --- | --- | --- |
| Raw commands | Only `private/` (owner-only ACL) holds unredacted text | `common.private_dir`, `commands_redacted.jsonl` has no raw field |
| Teacher/judge | Only redacted text is sent | `labeling.llm.chat` refuses any prompt where `find_secrets` matches |
| Model output | Must pass validators and must not contain any secret from the raw command | `api.server.Service`, `evaluation.validators.check` |
| Logs | Service logs hold a hash, source, latency and status only; no request lines | `Service._log`, `Handler.log_message` |
| Network | Loopback by default; non-loopback bind requires `LIVE_STATUS_API_TOKEN` | `api.server.main` |
| Held-out data | Test/validation families are frozen on first build | `datasets/frozen_families.json` |

## Why these choices

- **Heuristic as fallback, not fast path.** The deterministic describer answers in under a millisecond, but its confidence ≥ 0.9 outputs cover only 8.7% of executions and the judge rated just 19% of them ≥ 80 (they are correct but generic: "Reviewing the Git diff." for `git diff --stat`). The service therefore always asks the model and uses the heuristic when the model fails, times out or produces an invalid sentence; `--fast-path` re-enables the shortcut.
- **Plain completion format.** The student learns `Command:\n…\n\nStatus: <sentence>` with no system prompt, so each request costs only the command's tokens.
- **Ollama/llama.cpp for serving.** It already runs on this machine, serves GGUF at every quantization level, and keeps models warm. `llama-server` is supported by the same backend interface.
- **jev for grading, Opus for writing.** jev (TypeSafe System One) returns only probabilities,
choices and scores, in milliseconds, at $0.042 per million input tokens. Given the gold
status it agrees with Opus evaluations at AUC 0.92, so it grades every benchmark and ranks
student outputs for failure mining. Without a reference it is too weak (AUC 0.68) to
replace the Opus judge when labels are created.
- **Opus 5 as both teacher and judge**, with different prompts. Label diversity comes from two candidates per command plus the heuristic, not from different models.
- **Frequency weighting.** Each template group counts `min(4, 1 + log2(count))` times in training, so common patterns are learned first without drowning the long tail.

## Adding a transcript format

Write a function in `data_miner/sources.py` decorated with `@source(name, description, discover)` that yields `_record(...)` dictionaries, then rerun `extract_commands`.
72 changes: 72 additions & 0 deletions live-status/DATASET.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,72 @@
# Dataset

## Sources (mined 2026-09-16)

| Source | Parser | Records |
| --- | --- | ---: |
| Codex CLI sessions (incl. JS `exec` cells) | `codex` | 33,515 |
| Claude Code transcripts (`Bash`, `PowerShell`) | `claude-code` | 12,002 |
| OpenCode SQLite store | `opencode` | 12,500 |
| OpenCode2 host stores (live + release evidence) | `opencode2` | 5,369 |
| Grok CLI chat histories | `grok` | 933 |
| PSReadLine history (human-typed) | `psreadline` | 857 |
| Cursor agent chats | `cursor` | 555 |
| Curated adversarial fixtures | `adversarial` | 36 |

65,767 executions → 59,812 exact-unique → 58,109 template groups (numbers, UUIDs, hex ids,
timestamps and temp paths normalised). 95% of groups occur once; 54 occur 20+ times.
Known but unparsed locations (T3 state, legacy `.opencode`, the OpenCode2 request ledger) are
listed in `inventory.json`. Full distributions: `reports/mining_stats.md`.

Each record keeps the schema requested in the brief (`id, source_file, source_type, timestamp,
shell, command_raw, command_redacted, working_directory, preceding_context, following_context,
existing_model_text, exit_code, tags`). `existing_model_text` is the agent's own description
when the tool recorded one (12,000 groups); it is kept for analysis but never shown to the
teacher, because it carries intent the command does not show.

The shell is taken from the recording tool when known (Claude `Bash` → bash, Codex/OpenCode2/
Cursor on a Windows cwd → PowerShell, Codex `shell: cmd.exe` → cmd) and from syntax otherwise.

## Redaction

`redaction/redact.py` replaces vendor keys, bearer/basic tokens, cookies, URL credentials,
connection-string passwords, credential flags and env assignments, private keys and
high-entropy tokens with `<API_KEY> <TOKEN> <PASSWORD> <SECRET> <COOKIE> <PRIVATE_KEY>`.
Variable references (`$env:X`, `$secret`), code expressions and hex digests are kept.
An audit of the first pass found 4,336 groups flagged, mostly false positives (`-Pattern`
read as `mysql -p`, `git checkout -b` read as a curl cookie, `sessionID ===`); after fixes
180 groups are redacted. Raw commands exist only in `private/` (owner-only ACL).
The LLM client refuses any prompt in which `find_secrets` still matches, and payloads are
checked in their serialised form (9 commands are withheld because JSON escaping makes
ordinary text look secret-shaped).

## Labels

- Teacher: Opus 5, `teacher-v1`, temperature 0.4, two candidates per command (concise and
complete), batches of ≤20 commands / 30k characters.
- Judge: Opus 5, `judge-v1`, temperature 0, scores each candidate plus the heuristic,
picks the best and writes `recommended_output`.
- Decision: accepted when the recommended score ≥ 85, validators pass and the judge is not
uncertain; manual review at 70–84, uncertain, or validator failure; rejected for secret
leakage or score < 70. The judge rewrites weak candidates, so v1 has no rejections.

## v1 splits

| Split | Rows |
| --- | ---: |
| train | 3,113 |
| validation | 425 |
| test | 494 |
| manual_review | 102 |
| rejected | 0 |

Selection: the 40% most frequent templates, then round-robin over (shell, first action,
complexity) buckets with tagged hard examples first, plus every secret/injection/synthetic
group. Families (identical template or MinHash Jaccard ≥ 0.7) never cross splits; 0 template
collisions between train and validation/test. Test/validation ids are frozen in
`datasets/frozen_families.json`. Half of the secret, injection-like and synthetic families
go to test: test holds 82 secret-bearing, 38 injection-like and 16 synthetic commands, plus
124 long PowerShell, 66 loops, 45 conditionals, 60 natural-language and 52 malformed commands.
Per-split distributions are in `datasets/v1/report.json`.

Rows carry `weight = min(4, 1 + log2(count))`; training repeats a row that many times.
51 changes: 51 additions & 0 deletions live-status/EVALUATION.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,51 @@
# Evaluation

```powershell
python live-status/cli.py evaluate --backend ollama:<model>[:plain|:long|:instruct][:cpu] --name <run> [--grader jev|opus|both|none] [--promote]
python live-status/cli.py evaluate --regrade --name <run> --grader opus # re-grade saved outputs
```

Every run writes `evaluation/<run>.json` (report) and `evaluation/outputs/<run>.jsonl`
(per-command output, metrics, validator results and grades).

## Deterministic validators (`evaluation/validators.py`)

one sentence · 2–22 words (hard limit 30) · starts with an `-ing` verb · no boilerplate
("This command…") · no shell syntax (`|`, `&&`, `$env:`, `2>&1`…) · no secret: nothing that
`find_secrets` matches, no placeholder, and no secret substring of the raw command · target
recall for simple read/process/delete commands.

## Graders

| Grader | What it sees | Cost / speed | Agreement with Opus 5 |
| --- | --- | --- | --- |
| `opus` | command, reference, output; returns correct, score, missing/hallucinated actions, secret leak, injection followed | ~20 outputs per call, ~60 s | — |
| `jev` (default) | command, reference, output; answers `same_actions`, `invented`, `quality` | 700 outputs in ~3 s, ~$0.02 | AUC 0.92, 86.4% agreement at score ≥ 0.45 (700 Opus-graded outputs) |

A jev-accepted output passes validators and has
`same × (1 − invented) × quality/4 ≥ 0.45`. Calibration lives in
`evaluation/jev_eval_calibration.json`; rerun it with `python live-status/judging/jev.py calibrate-eval`.
Grading a status without a reference is much weaker (AUC 0.68 on 3,000 teacher candidates), so
jev only *ranks* unlabeled outputs during failure mining and Opus writes the labels.

jev is stricter than Opus on good outputs: Opus's alternate teacher candidates on 200 test
commands pass 86% of Opus evaluations and 72% of jev evaluations. Compare runs only within
one grader.

Both graders cache verdicts by (command, output), so re-evaluating an unchanged output is free.

The `no_secret` validator also fails outputs that echo a redaction placeholder
(`<TOKEN>`, `<SECRET>`). Two of the round-1 runs did this once each, which blocks promotion;
the service replaces such outputs with the heuristic fallback.

## Promotion (`--promote`)

A run is promoted into `evaluation/registry.json` only if nothing leaks, its accepted rate is at
least the current best, hallucination does not rise by more than a point, and it was graded
by the same grader as the current best.

## Reports

Each report includes validator rates, accepted %, invented/hallucination %, omission % (Opus),
secret leakage %, injection-followed % (Opus), latency p50/p90/p99, tokens/s, model RAM/VRAM
from Ollama, and accepted % per difficulty tag.
70 changes: 70 additions & 0 deletions live-status/README.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,70 @@
# Live status

A tiny local model that turns a shell or tool command into one live status sentence:

```
Get-Process opencode2,node,powershell -ErrorAction SilentlyContinue | Select-Object Name,Id
→ Checking running opencode2, node, and powershell processes.
```

Current best: **Qwen3-0.6B + LoRA, GGUF q4_K_M (397 MB)**, 67.8% of held-out statuses
accepted by the grader (the teacher's own second choices score 72%), 100 ms p50 on an
RTX 3070 and 393 ms on CPU only. Untuned models of the same size score at most 23.5%.
Details: [TRAINING.md](TRAINING.md), [EVALUATION.md](EVALUATION.md),
[DATASET.md](DATASET.md), [ARCHITECTURE.md](ARCHITECTURE.md).

## Use it

```powershell
python live-status/cli.py serve --backend ollama:live-status-v1-qwen3-06b-lora-q4_k_m --port 8765
python live-status/api/client.py "git fetch origin && git status -sb"
```

`POST /v1/summarize-command` with `{"command": "...", "shell": "powershell", "cwd": "optional"}`
returns `{"status": "..."}` (add `"debug": true` for source and latency). The service redacts
before inference, caches by normalised command, bounds concurrency, falls back to a
deterministic describer on timeout or invalid output, and never logs commands. For remote use
set `LIVE_STATUS_API_TOKEN` and pass `--tls-cert/--tls-key`; a non-loopback bind without a
token is refused. `api/client.py` works unchanged against `https://my-server.example`.

## Pipeline

Run from the repository root. GPU steps use the training venv (see TRAINING.md).

| Step | Command |
| --- | --- |
| Inventory transcript sources | `python live-status/cli.py inventory_sources` |
| Extract, redact, deduplicate, report | `python live-status/cli.py extract_commands` |
| Teacher labels (Opus 5) | `python live-status/cli.py generate_labels --limit 4000` |
| Judge labels (Opus 5) | `python live-status/cli.py judge_labels` |
| Build splits | `python live-status/cli.py build_dataset --version v1` |
| Baseline tiny models | `python live-status/cli.py benchmark_base_models --models smollm2:135m qwen3:0.6b` |
| Train | `live-status/cli.py train --base Qwen/Qwen3-0.6B --method lora --name ...` |
| Export GGUF + Ollama | `live-status/cli.py export_gguf --name ... --quants q8_0 q4_K_M` |
| Evaluate / promote | `python live-status/cli.py evaluate --backend ollama:... --name ... --promote` |
| Mine failures | `python live-status/cli.py mine_failures --backend ollama:... --round r2` |
| Mining → dataset in one go | `python live-status/cli.py run_full_pipeline` |
| Redact any JSONL | `python live-status/cli.py redact_dataset in.jsonl out.jsonl` |

`live-status/scripts/<step>.py` wraps each command. Private data lives in
`LIVE_STATUS_HOME` (default `<main checkout>/work/live-status`, ignored by Git); raw
commands stay in its `private/` folder.

## Models and keys

- Teacher and judge: `claude-opus-5` through the local CLIProxyAPI (`127.0.0.1:8317`).
It shares the Claude subscription; keep `--workers` at 4 or below. On a long cooldown
the run stops and resumes on rerun.
- Grader: TypeSafe `jev` through Vercel AI Gateway (`AI_GATEWAY_API_KEY` in the repo
`.env`) or directly (`TYPESAFE_API_KEY`). Needs Bun; `bun install` in `live-status/jev`.
- Serving: Ollama; k-quants are produced with llama.cpp's `llama-quantize`.

## Tests

```powershell
python -m unittest discover -s live-status/tests -t live-status -v
```

Redaction (secrets never reach teacher prompts, stored redacted data, service output or
logs) and service boundaries (auth, size and rate limits, loopback-only default,
invalid-output fallback).
Loading
Loading