Skip to content

quota: explicit Todo selection can reject the canonical runnable successor #4327

Description

@huangruiteng

Problem

A material quota monitor-poll can create a claimed advancement_task successor that is visible as runnable in the canonical Todo list and selected by ordinary quota should-run, while explicit same-Turn selection rejects the exact same Todo with heartbeat_receipt_identity_conflict.

Minimal observed contract

  • The monitor transition returned a typed successor receipt for a new claimed advancement Todo.
  • loopx todo list --todo-id <id> reported the Todo open and included it in agent_todos.executable_backlog_items.
  • loopx quota should-run without --todo-id selected that exact Todo from capability_gate.runnable_candidates with no missing capabilities.
  • loopx quota should-run --todo-id <same-id> --turn-instance-id <fresh-id> rejected it as not currently projected agent-scoped/capability-ready.
  • Raising --limit did not change the rejection.

Likely seam

Default selection can consume the capability-gate/planning projection, but build_explicit_advancement_next_action qualifies against select_quota_todo_source_items. Those projections can diverge after a monitor-created successor, so the explicit action guard fails closed even though the canonical runnable projection selects the Todo.

Expected behavior

Explicit selection must qualify from the same canonical authoritative Todo inventory as default selection, or return a typed projection-staleness reason plus a deterministic refresh/retry path. A failed receipt must not leave a newly created runnable successor impossible to settle.

Acceptance

  1. Add a regression: material monitor poll creates a successor beyond compact display lanes; default and explicit selection agree on the exact Todo.
  2. Preserve claim, required-capability, Goal/Agent and continuation boundaries.
  3. Retry after a typed failed receipt is deterministic and idempotent.
  4. CLI/managed Turn use the same selection source and return a path-free typed receipt.

Activity

  1. LIHUA919 commented on Sep 13, 2026

    @LIHUA919
    Contributor

    I'd like to take this bounded fix. Plan: reproduce the monitor-created successor outside compact display lanes; make default and explicit selection consume the same canonical eligible inventory; preserve claim, capability, Goal/Agent and continuation fences; verify deterministic failed-receipt retry and CLI/managed Turn parity.

    I'll coordinate with #4061's complete-source fallback work and keep this PR scoped to #4327. First milestone is a failing regression on current main, followed by the fix and real-entrypoint validation in one reviewable PR.

  2. LIHUA919 commented on Sep 13, 2026

    @LIHUA919
    Contributor

    Reproduced the reported heartbeat_receipt_identity_conflict on current main (7eb4b7bb1661bd5eff63a8725a33169792d5964b) with a real canonical File store, a material monitor-created claimed successor, and an agent lane beyond compact display limits.

    The observed cause differs from the initial source-divergence hypothesis: the successor reaches the explicit candidate builder, but TypeScript returns state=deferred, reason=autonomous_replan for the long-Todo-chain gate. Default quota still displays that same successor while normal_delivery_allowed=false. The CLI drops the qualification reason and reports an identity/receipt-write failure even though no conflicting receipt exists.

    I'm preserving the replan/claim/capability boundaries and fixing this preflight path to return the typed selection reason, avoid a false failed receipt, and provide same-Goal/Agent/Turn guard re-entry. Validation covers normal admitted selection, rejected/deferred selection, retry without duplicate receipts, and actual monitor successors on legacy/File/SQLite. If the original default packet had normal_delivery_allowed=true, that would indicate an additional source-divergence case; a minimal public-safe fixture or the relevant typed fields would help isolate it without raw state/logs.

  3. huangruiteng commented on Sep 13, 2026

    @huangruiteng
    CollaboratorAuthor

    Additional same-Turn reproduction on main 7eb4b7b, from a real Goal-bound File store:

    • Initial guided should-run committed turn guided-start:767b746128d57caca47aaa7d404c6c05.
    • It selected an existing in-flight Todo, while capability_gate reported runnable_count=4 and interaction_contract requested explicit selection; agent_channel.delivery_allowed=false.
    • A new canonical open advancement Todo was then added in the same Turn with claimed agent, task_repository, target-capability, write scope and action_kind.
    • Explicit should-run for that new Todo with the same turn id returned heartbeat_receipt_identity_conflict and write_failed, reason: explicit action selection must name one currently projected agent-scoped, capability-ready Todo.
    • The initial heartbeat receipt had no settlement_identity, so this is not an actual identity mismatch. It is a post-guard frontier-staleness/add-after-guard case.

    This is distinct from the monitor-successor case but supports the same contract fix: return a typed projection_stale or todo_not_in_turn_frontier outcome, perform no false receipt write, and project one deterministic refresh/re-entry action. It should not silently admit a Todo created after the guard, but it also should not misreport an identity conflict or strand the Turn.

  4. huangruiteng commented on Sep 13, 2026

    @huangruiteng
    CollaboratorAuthor

    A second real same-Turn reproduction confirms the frontier-staleness branch and rules out capability admission as the reason:

    • Initial guided Turn guided-start:bd53343e23d713716d9109f837b7d6cb returned ok=true, should_run=true, normal_delivery_allowed=true, and effective_action=normal_run.
    • Its action_portfolio.suggested_actions explicitly included the requested open P0 Todo as an alternative; agent_todo_summary.first_executable_items included the same Todo with matching agent, repository, action kind and required capabilities.
    • capability_gate reported six runnable candidates, empty missing, blocked_missing and repair_missing.
    • Before explicit binding, unrelated canonical Todo planning updates advanced the active-state frontier.
    • Reusing the exact projected command, capabilities and Turn ID for the originally suggested Todo returned heartbeat_receipt_identity_conflict with write_failed; the initial receipt had no settlement identity.

    This should not silently bind across a changed frontier. It should preserve the safe refusal but expose projection_stale / turn_frontier_changed, identify that the requested Todo was valid in the prior projection, perform no fake receipt write, and provide one deterministic refresh/re-entry action. This is the same #4327 repair line, not a new issue.

  5. huangruiteng commented on Sep 13, 2026

    @huangruiteng
    CollaboratorAuthor

    Implemented the bounded repair in PR #4335 at exact head eb8930d. The TS action-selection reducer remains authoritative; the CLI now exposes deferred/rejected qualification, replays an existing identity-less receipt without mutation, writes no event for a first-call rejection, and preserves true bound-receipt conflicts. Validation: 47/47 settlement CLI tests, 19 adjacent Python tests, 11 TS reducer tests, 9 CLI projection tests, Ruff/diff/docs governance; exact post-latest-main regression 3/3. Review remains required because this changes quota/runtime semantics.

  6. LIHUA919 commented on Sep 13, 2026

    @LIHUA919
    Contributor

    @huangruiteng 感谢补充。看到 #4335 也在处理 #4327,我这边 #4332 已通过全部 CI,包含真实 File/SQLite 回归,以及最终 workspace/scope guard 后重新计算资格的修复。

    两份 PR 重叠较多,建议收敛到一份,保留双方互补的修复、测试和提交归属。你倾向以哪份为主?我可以按确认后的方向补齐差异,避免继续重复开发。

  7. huangruiteng commented on Sep 13, 2026

    @huangruiteng
    CollaboratorAuthor

    @huangruiteng 感谢补充。看到 #4335 也在处理 #4327,我这边 #4332 已通过全部 CI,包含真实 File/SQLite 回归,以及最终 workspace/scope guard 后重新计算资格的修复。

    两份 PR 重叠较多,建议收敛到一份,保留双方互补的修复、测试和提交归属。你倾向以哪份为主?我可以按确认后的方向补齐差异,避免继续重复开发。

    我的锅,我的 agent 没看到你的认领。。

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions