Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
16 changes: 15 additions & 1 deletion live-status/ARCHITECTURE.md
Original file line number Diff line number Diff line change
Expand Up @@ -21,7 +21,9 @@
evaluation/evaluate.py ──► validators + jev/Opus grading + latency/memory ──► evaluation/registry.json
evaluation/active.py ──► student failures on unlabeled commands ──► teacher/judge ──► prefs.jsonl (DPO)

client ──► api/server.py ──► redact ─► cache ─► model (Ollama) ─► validate ─► heuristic fallback
client ──► api/server.py ──► redact ─► cache ─┬► PowerShell AST ─► mapped renderer ─┐
└► model (Ollama) ───────────────────┤
validate ─► heuristic fallback
```

## Boundaries
Expand All @@ -37,6 +39,18 @@

## Why these choices

- **Parser-first PowerShell prototype.** `parser:powershell` keeps one local PowerShell
process warm, extracts commands with PowerShell's real AST, and renders only mapped facts.
With a working directory, a semantic resolver can ground a test action in its declared test
name. Parse errors and dynamic invocation abstain; unknown executables remain literal. It
uses no model or GPU at runtime. A development-only linguistic trainer places every mapping
in simulated combinations and asks an explicitly selected LLM for proposals; only reviewed,
accepted phrases enter `linguistic_map.json`. Coverage and grounding are reported separately
from human usefulness.
The product target is the actual command cells agents execute: test runners, Python modules
and scripts, file searches, Git/GitHub operations, and their meaningful targets and settings.
A passing mapping must retain those details; naming only the executable is a fallback, not a
useful result. Bash needs its own parser and is not covered by this PowerShell prototype.
- **Heuristic as fallback, not fast path.** The deterministic describer answers in under a millisecond, but its confidence ≥ 0.9 outputs cover only 8.7% of executions and the judge rated just 19% of them ≥ 80 (they are correct but generic: "Reviewing the Git diff." for `git diff --stat`). The service therefore always asks the model and uses the heuristic when the model fails, times out or produces an invalid sentence; `--fast-path` re-enables the shortcut.
- **Plain completion format.** The student learns `Command:\n…\n\nStatus: <sentence>` with no system prompt, so each request costs only the command's tokens.
- **Ollama/llama.cpp for serving.** It already runs on this machine, serves GGUF at every quantization level, and keeps models warm. `llama-server` is supported by the same backend interface.
Expand Down
19 changes: 18 additions & 1 deletion live-status/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -17,16 +17,33 @@ Details: [TRAINING.md](TRAINING.md), [EVALUATION.md](EVALUATION.md),
## Use it

```powershell
python live-status/cli.py serve --backend parser:powershell --port 8765
python live-status/cli.py serve --backend ollama:live-status-v1-qwen3-06b-lora-q4_k_m --port 8765
python live-status/api/client.py "git fetch origin && git status -sb"
```

`parser:powershell` is the CPU-only prototype. It uses PowerShell's own AST, maps only
known commands, repeats unknown executable names literally, and abstains on malformed or
dynamic invocation. When `cwd` is supplied, it can resolve a named test file and report its
declared test purpose instead of restating the filename. It does not load a model or use the GPU.
The recent-session result and
its limitations are recorded in
[the model sweep](research/2026-09-20-model-sweep.md#parser-first-powershell-prototype).

An LLM may improve the linguistic mapping during development without joining the runtime path.
`python live-status/linguistic_training.py simulate --output <file>` creates stable single,
pair, and triple phrase combinations. `propose --route <provider-identity> --model
<exact-user-selection> --output <file>` asks that explicitly selected CLIProxyAPI route and
model for grounded wording proposals. Every proposal
defaults to `accepted: false`; after review, `apply --input <file>` writes only accepted entries
to `inference/linguistic_map.json`.

`--best-of 4` samples several candidates and lets jev pick the best: +8 points of quality for
~1 s per request instead of ~0.1 s.

`POST /v1/summarize-command` with `{"command": "...", "shell": "powershell", "cwd": "optional"}`
returns `{"status": "..."}` (add `"debug": true` for source and latency). The service redacts
before inference, caches by normalised command, bounds concurrency, falls back to a
before inference, caches by shell, working directory, and normalised command, bounds concurrency, falls back to a
deterministic describer on timeout or invalid output, and never logs commands. For remote use
set `LIVE_STATUS_API_TOKEN` and pass `--tls-cert/--tls-key`; a non-loopback bind without a
token is refused. `api/client.py` works unchanged against `https://my-server.example`.
Expand Down
14 changes: 9 additions & 5 deletions live-status/api/server.py
Original file line number Diff line number Diff line change
Expand Up @@ -59,7 +59,7 @@ def _log(self, row: dict) -> None:
def summarize(self, command: str, shell: str | None, cwd: str | None) -> dict:
t0 = time.perf_counter()
red = redact(command[:MAX_COMMAND])
key = sha(f"{shell}|{normalize_ws(red)}")
key = sha(f"{shell}|{cwd}|{normalize_ws(red)}")
with self.lock:
hit = self.cache.get(key)
if hit is not None:
Expand All @@ -69,12 +69,16 @@ def summarize(self, command: str, shell: str | None, cwd: str | None) -> dict:
h_text, h_conf = describe(red, shell)
if (self.fast_path and h_conf >= FAST_PATH) or self.backend is None:
return self._done(h_text, "heuristic" if h_conf >= FAST_PATH else "fallback", key, t0, cache=h_conf >= FAST_PATH)
status, source = None, "model"
status = None
source = getattr(self.backend, "source", "model")
if self.slots.acquire(timeout=self.queue_wait):
try:
model_in = red if len(red) <= MODEL_INPUT_CHARS else red[:MODEL_INPUT_CHARS] + " …"
status, _ = self.backend.generate(model_in)
if self.best_of > 1:
if getattr(self.backend, "accepts_cwd", False):
status, _ = self.backend.generate(model_in, cwd=cwd)
else:
status, _ = self.backend.generate(model_in)
if self.best_of > 1 and source == "model":
picked, source = self._best_of(model_in, status)
status = picked or status
except Exception as exc: # timeouts, backend down
Expand Down Expand Up @@ -197,7 +201,7 @@ def do_POST(self):
def main(argv=None):
p = argparse.ArgumentParser(prog="serve")
p.add_argument("--backend", default=os.environ.get("LIVE_STATUS_BACKEND", "ollama:live-status"),
help="ollama:<model> | llama-server:<url> | hf:<base>@<adapter> | none")
help="parser:powershell | ollama:<model> | llama-server:<url> | hf:<base>@<adapter> | none")
p.add_argument("--host", default="127.0.0.1")
p.add_argument("--port", type=int, default=8765)
p.add_argument("--timeout", type=float, default=8.0)
Expand Down
6 changes: 5 additions & 1 deletion live-status/inference/backends.py
Original file line number Diff line number Diff line change
Expand Up @@ -150,8 +150,12 @@ def _stop_ids(self) -> list[int]:


def from_spec(spec: str):
"""ollama:<model>[:mode][:cpu] | llama-server:<url> | hf:<base>[@<adapter>][:structured]"""
"""Create a configured inference or deterministic parser backend."""
kind, _, rest = spec.partition(":")
if kind == "parser" and rest == "powershell":
from inference.parser_first import PowerShellAstBackend

return PowerShellAstBackend()
if kind == "ollama":
mode, cpu = "plain", False
if rest.endswith(":cpu"):
Expand Down
4 changes: 4 additions & 0 deletions live-status/inference/linguistic_map.json
Original file line number Diff line number Diff line change
@@ -0,0 +1,4 @@
{
"version": 1,
"action_overrides": {}
}
Loading
Loading