Skip to content

fix(cli): return parsed claim repair arguments in one packet - #5586

Open
loopx-agent wants to merge 3 commits into
mainfrom
codex/claim-protocol-recovery
Open

loopx-agent wants to merge 3 commits into
mainfrom
codex/claim-protocol-recovery

Conversation

@loopx-agent

Copy link
Copy Markdown
Collaborator

Goal And Delivered Outcome

Malformed parsed todo claim calls returned only the first textual error. A call with an actor, missing executor and unsupported Turn/metadata flags could therefore require several repairs or a full Todo help lookup. The error now returns one recovery packet: supplied legal argv, all missing inputs and the flags removed from that argv. The caller supplies explicit missing values and reviews removals before retrying from the original working directory.

  • Basis: trajectory-led long-horizon command recovery, qualified with public-safe synthetic reproductions under the existing S2/S10 caller boundary.
  • Before → after: test_one_packet_repairs_all_grammar_errors_and_preserves_route rejects without a write, executes the returned repair, verifies preview, adoption and canonical replay on Legacy/File/SQLite.
  • Intended base: main at f35978e3a1289a146687346467b848bfca8fa2b0; no issue is closed.

Author Declaration

  • Written by: model_agent — OpenAI GPT-6 (Codex).
  • Implemented against: docs/architecture/rfcs/typescript-control-plane-migration-v0.md (§2.5 and long-goal caller boundary) and docs/development/testing-and-quality.md#roadmap-aligned-optimization, at the base above.
Criterion Disposition Owner / evidence
Keep one claim admission owner and validate the input boundary implemented Existing CLI grammar formats a typed local usage error; the TypeScript claim transaction and provider admission are unchanged. Real repair and rejection tests exercise the owning transaction.
Separate caller repeats/context from provider cost and qualify correctness implemented Matched scripted recovery below; independent ownership, source-mode and CAS negatives. Actual model costs remain unknown.
Sustained provider operation and ten-day Goal acceptance out_of_scope No store, routing, retention, soak or benchmark-run change.

Self-check: inspected the final diff, related open claim work and latest main; ran the focused real CLI/backend, parity and static checks below. No raw trace or private probe is included.

Scope And Continuation

The useful slice is parsed claim usage recovery, complete within that boundary. Unknown flags, invalid values and missing parser-required inputs retain existing argparse diagnostics before the handler; no parser redesign is claimed. The argv preserves supplied actor/executor, route, project/state path, preview, operation identity and lease/CAS, including zero. Missing executors and lease keys are never invented. A repaired request still undergoes registration, ownership, source-mode and hard-lease admission. Legacy canonical-only flags are retained and rejected; the formatter does not promote state or relax CAS.

Python remains the existing grammar/formatting adapter, not another domain decision owner. todo_claim_invalid_arguments is a local CLI usage classification, not a persisted or shared authority vocabulary. No capability, provider, setting or RPC is added. The related refactor reuses the claim field set for validation and argv projection and removes the duplicate handler validation call. JSON and Markdown expose the same facts.

Default behavior disclosure: malformed claims gain richer output and claim-specific validation runs before shared validation; their first human diagnostic can differ. Existing claim-specific and Turn diagnostics remain, and successful/unrelated command output stays unchanged. Source-mode recovery and ordinary-short/replan-full context qualification remain in the existing research frontier, not prerequisites for this reversible repair. Maintainer merge is required.

Validation

  • Tested revision: b92c08d13a4c23f034c68a7c1d37f4de30eb1b93; local run state: finished.
  • Input classes: synthetic, authorized_private_read_only. Only generalized findings and synthetic product tests are public.
Check kind Result Evidence / limitation
unit / real_entrypoint passed 205 Python diagnostics, recovery and source/console entrypoint tests; another 63 mutation/claim/lease tests. Includes 19 real repair/negative cases, preview/replay, foreign executor, missing inputs, invalid and zero CAS, hard lease and legacy mode refusal.
real_backend passed Disposable real File and SQLite providers, with Markdown display removed; authority revision/readback unchanged after rejected and preview requests. Legacy Markdown also qualified.
regression_parity passed examples/control_plane/cli-output-budget-regression-smoke.py; ordinary command matrix retains its existing base/head budgets.
static passed Control-plane typecheck, changed Python Ruff, diff check, semantic advisory/full vocabulary smoke, RFC language mirrors and ledger checks. I/O manifest only changes seven line locations; no new sites/classifications.
static failed Full docs governance fails on the same two pre-existing dated headings in external-evidence-research-capability-v0.md and its Chinese mirror, reproduced on the unchanged base. No unrelated repair bundled.
manual passed Candidate public/private scan and ownership/placement review; raw trajectories, local paths and probes excluded.

Matched scripted consumer, four interleaved trials per arm on the same Python executable:

Observation Base Candidate
Initial error stdout bytes 232 1,073
Error-to-success stdout bytes 29,280 3,917
CLI processes in that recovery path 3 2
Python fixture file-open events before admission 0 0

The base policy is error → full claim help → exact valid claim; candidate is error → argv plus explicit required executor → valid claim. The richer error costs 841 extra bytes; a consumer that already repairs from prose alone may not benefit. This proves an available deterministic recovery path, not actual model retry/token savings. Python audit counts exclude Node provider IO and OS byte reads; there is no latency, sustained provider, benchmark-score or backend-IO improvement claim. Actual model tokens are unavailable. Remote CI/merge readiness is not yet qualified.

Frontend / Visual Evidence

UI impact: none. Existing source/console CLI JSON and Markdown change; frontend/Lark controls, packaged assets and configuration do not. No visual layout or new caller-facing capability is introduced. Source data for screenshots: none.

Type of Change / LoopX Area

Bug fix, documentation and focused validation; existing Todo CLI/error projection. Direction: S2/S10 caller recovery under the TS migration boundary.

Shared-authority RFC fixture impact

No production-scale authority schema or semantic dimension changes. Real Legacy/File/SQLite arms run. PostgreSQL three-arm promotion/routing rehearsal is not applicable: no promotion, runtime routing or authority compatibility projection changes; grammar rejection precedes authority access. Existing provider correctness and sustained performance acceptance remain open.

Boundary Checklist

  • Diff/body exclude private state, credentials, raw traces, internal links and local paths.
  • No duplicate benchmark runner or maintainer-owned adapter work.
  • Scope stays on claim grammar recovery; remaining research stays with its existing owner.
  • UI impact is none.
  • All commits carry DCO sign-offs.

中文:一次 claim 语法错误同时返回缺失/非法参数与保留已有值的 argv,调用方补显式输入后重试;不猜执行者、不删有效 lease/CAS 或降低权限。Legacy/File/SQLite 真实恢复与负例已通过,TS 仍是 claim 权威。错误本身更长;上述节省只适用于脚本化的“错误后查完整 help”路径,不代表实际模型 token、重试率或 EdgeBench 得分提升。

Signed-off-by: LoopX Agent <337587101+loopx-agent@users.noreply.github.com>
Signed-off-by: LoopX Agent <337587101+loopx-agent@users.noreply.github.com>
Signed-off-by: LoopX Agent <337587101+loopx-agent@users.noreply.github.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant