Repository navigation
quota: explicit Todo selection can reject the canonical runnable successor #4327
Description
Activity
I'd like to take this bounded fix. Plan: reproduce the monitor-created successor outside compact display lanes; make default and explicit selection consume the same canonical eligible inventory; preserve claim, capability, Goal/Agent and continuation fences; verify deterministic failed-receipt retry and CLI/managed Turn parity.
I'll coordinate with #4061's complete-source fallback work and keep this PR scoped to #4327. First milestone is a failing regression on current main, followed by the fix and real-entrypoint validation in one reviewable PR.
Reproduced the reported
heartbeat_receipt_identity_conflicton current main (7eb4b7bb1661bd5eff63a8725a33169792d5964b) with a real canonical File store, a material monitor-created claimed successor, and an agent lane beyond compact display limits.The observed cause differs from the initial source-divergence hypothesis: the successor reaches the explicit candidate builder, but TypeScript returns
state=deferred,reason=autonomous_replanfor the long-Todo-chain gate. Default quota still displays that same successor whilenormal_delivery_allowed=false. The CLI drops the qualification reason and reports an identity/receipt-write failure even though no conflicting receipt exists.I'm preserving the replan/claim/capability boundaries and fixing this preflight path to return the typed selection reason, avoid a false failed receipt, and provide same-Goal/Agent/Turn guard re-entry. Validation covers normal admitted selection, rejected/deferred selection, retry without duplicate receipts, and actual monitor successors on legacy/File/SQLite. If the original default packet had
normal_delivery_allowed=true, that would indicate an additional source-divergence case; a minimal public-safe fixture or the relevant typed fields would help isolate it without raw state/logs.Additional same-Turn reproduction on main 7eb4b7b, from a real Goal-bound File store:
- Initial guided should-run committed turn guided-start:767b746128d57caca47aaa7d404c6c05.
- It selected an existing in-flight Todo, while capability_gate reported runnable_count=4 and interaction_contract requested explicit selection; agent_channel.delivery_allowed=false.
- A new canonical open advancement Todo was then added in the same Turn with claimed agent, task_repository, target-capability, write scope and action_kind.
- Explicit should-run for that new Todo with the same turn id returned heartbeat_receipt_identity_conflict and write_failed, reason: explicit action selection must name one currently projected agent-scoped, capability-ready Todo.
- The initial heartbeat receipt had no settlement_identity, so this is not an actual identity mismatch. It is a post-guard frontier-staleness/add-after-guard case.
This is distinct from the monitor-successor case but supports the same contract fix: return a typed projection_stale or todo_not_in_turn_frontier outcome, perform no false receipt write, and project one deterministic refresh/re-entry action. It should not silently admit a Todo created after the guard, but it also should not misreport an identity conflict or strand the Turn.
A second real same-Turn reproduction confirms the frontier-staleness branch and rules out capability admission as the reason:
- Initial guided Turn
guided-start:bd53343e23d713716d9109f837b7d6cbreturnedok=true,should_run=true,normal_delivery_allowed=true, andeffective_action=normal_run. - Its
action_portfolio.suggested_actionsexplicitly included the requested open P0 Todo as an alternative;agent_todo_summary.first_executable_itemsincluded the same Todo with matching agent, repository, action kind and required capabilities. capability_gatereported six runnable candidates, emptymissing,blocked_missingandrepair_missing.- Before explicit binding, unrelated canonical Todo planning updates advanced the active-state frontier.
- Reusing the exact projected command, capabilities and Turn ID for the originally suggested Todo returned
heartbeat_receipt_identity_conflictwithwrite_failed; the initial receipt had no settlement identity.
This should not silently bind across a changed frontier. It should preserve the safe refusal but expose
projection_stale/turn_frontier_changed, identify that the requested Todo was valid in the prior projection, perform no fake receipt write, and provide one deterministic refresh/re-entry action. This is the same #4327 repair line, not a new issue.- Initial guided Turn
Implemented the bounded repair in PR #4335 at exact head eb8930d. The TS action-selection reducer remains authoritative; the CLI now exposes deferred/rejected qualification, replays an existing identity-less receipt without mutation, writes no event for a first-call rejection, and preserves true bound-receipt conflicts. Validation: 47/47 settlement CLI tests, 19 adjacent Python tests, 11 TS reducer tests, 9 CLI projection tests, Ruff/diff/docs governance; exact post-latest-main regression 3/3. Review remains required because this changes quota/runtime semantics.
@huangruiteng 感谢补充。看到 #4335 也在处理 #4327,我这边 #4332 已通过全部 CI,包含真实 File/SQLite 回归,以及最终 workspace/scope guard 后重新计算资格的修复。
两份 PR 重叠较多,建议收敛到一份,保留双方互补的修复、测试和提交归属。你倾向以哪份为主?我可以按确认后的方向补齐差异,避免继续重复开发。
@huangruiteng 感谢补充。看到 #4335 也在处理 #4327,我这边 #4332 已通过全部 CI,包含真实 File/SQLite 回归,以及最终 workspace/scope guard 后重新计算资格的修复。
两份 PR 重叠较多,建议收敛到一份,保留双方互补的修复、测试和提交归属。你倾向以哪份为主?我可以按确认后的方向补齐差异,避免继续重复开发。
我的锅,我的 agent 没看到你的认领。。
Problem
A material
quota monitor-pollcan create a claimedadvancement_tasksuccessor that is visible as runnable in the canonical Todo list and selected by ordinaryquota should-run, while explicit same-Turn selection rejects the exact same Todo withheartbeat_receipt_identity_conflict.Minimal observed contract
loopx todo list --todo-id <id>reported the Todo open and included it inagent_todos.executable_backlog_items.loopx quota should-runwithout--todo-idselected that exact Todo fromcapability_gate.runnable_candidateswith no missing capabilities.loopx quota should-run --todo-id <same-id> --turn-instance-id <fresh-id>rejected it as not currently projected agent-scoped/capability-ready.--limitdid not change the rejection.Likely seam
Default selection can consume the capability-gate/planning projection, but
build_explicit_advancement_next_actionqualifies againstselect_quota_todo_source_items. Those projections can diverge after a monitor-created successor, so the explicit action guard fails closed even though the canonical runnable projection selects the Todo.Expected behavior
Explicit selection must qualify from the same canonical authoritative Todo inventory as default selection, or return a typed projection-staleness reason plus a deterministic refresh/retry path. A failed receipt must not leave a newly created runnable successor impossible to settle.
Acceptance