Skip to content

[Bug] long_todo_chain replan 在合法 vision-patch 写回后 4 分钟即重新武装,且投影指令无法清除该义务 #4667

Description

@huangruiteng

现象

lane 的 long_todo_chain replan 义务在一次合法交付之后 4 分钟就重新武装并换了新的 obligation id,而该 lane 在此之前刚刚按 guard 投影的指令完成了 bounded replan 写回。结果是"每次唤醒都要求重规划"的循环。

时间线(同一 lane codex-side-bypass, goal loopx-meta)

时间 (UTC) 事件
15:42 guard 要求 replan,obligation replan-c471cccdf3d0a276,trigger long_todo_chain(selectable 21 > 阈值 20),frontier_revision todo_frontier_revision_v0:506463d15b19380633011591
15:44 按投影指令更新了一条 lane todo 的证据(todo_4144c9deec0f)
15:45 交付 refresh-state --classification bounded_replan_progress --autonomous-replan-recorded --repair-delta-kind goal_vision_patch --repair-delta-kind blocker --agent-vision-json <packet>:写入 vision patch(含 acceptance_summary 与 goal_path_delta_v0, outcome=replan)并记录 autonomous_replan_ack(recorded: true)
15:46 记账 1 slot;归档已完成 todo
15:49 guard 再次要求 replan,新 obligation replan-ca90e6ddfe013b35,同一 trigger,frontier_revision 变为 todo_frontier_revision_v0:7ba751eda7da58a9f6601e2a,rearmed_after_obligation_id: replan-3363450b0714266d(一个更早的 id,不是刚刚交付的 c471ccc…)

机制

  1. 抑制触发需要 ACK 携带该 trigger 的 frontier 检查点:loopx/control_plane/todos/frontier_revision.ts::classifyAck 要求 delta.trigger_checkpoints[] 中存在 frontier_revision 等于当前 revision,或 frontier_owned_identity 等于当前 owned identity。
  2. 产生这种语义 delta 的路径是创建一个 typed runnable successor:loopx/control_plane/goals/goal_frontier/ack_policy.py 会为它写入 replan_semantic_delta_v0(outcomes new_runnable_successor、obligation id、successor todo id、trigger_checkpoints)。而 replan_successor_semantic_binding 要求 successor 同时具备 action_kind 和 target_key(或 explore refs)。
  3. guard 投影给出的补救指令并不包含这些要素:
    • todo_actions[0].text = "run a bounded long-chain vision replan: compare evidence with the active vision, group or prune the todo chain, and select the next high-value runnable slice" —— 直接 todo add 这段文字不会带上 action_kind/target_key,因此不能被认定为 typed successor;
    • cli_channel.next_cli_actions[0] 只给 refresh-state ... --autonomous-replan-recorded(外加 progress/vision 字段),照做会得到 recorded: true 但没有 trigger_checkpoints 的 ACK,于是 trigger 继续有效。

也就是说:按照 guard 投影的指令字面执行,无法清除 long_todo_chain 义务;只有额外知道"必须落一个带 action_kind+target_key 的 successor todo"才能清除。这一点目前只存在于实现里,没有出现在投影指令中。

影响

lane 每 3–4 分钟被要求重规划一次,真正的推进工作被挤出;同时每次重规划都会写入 run/vision 记录并占用一次 quota 结算。同一 lane 在本次会话的另一条记录里已被用户提醒过该模式。

建议

三者取其一即可,且都属小改动:

  1. 让投影的 todo_actions 直接给出可执行的 typed successor(带 action_kind、target_key、task_repository),使照做即可清除义务;
  2. 或在 todo_actions/cli_channel 中明确写出"需要落一个带 action_kind+target_key 的 successor todo"这一前提;
  3. 若 vision-patch 路径本应也能结算该义务,则让该路径的 ACK 也带上当前 trigger 的检查点(按 frontier_owned_identity 绑定,避免自己的写回改动 revision 后立刻自失效)。

复现

在一个 selectable open todo > 20 的 lane 上:执行 quota should-run 得到 obligation A → 仅按投影指令做 refresh-state --autonomous-replan-recorded(vision patch 或 blocker delta)→ 等待下一次 quota should-run:会出现新的 obligation id,且 rearmed_after_obligation_id 指向更早的 obligation。

本 lane 的绕行做法(已执行):用 todo add 建立带 action_kind=review_pull_requests + target_key=github-pr-review:huangruiteng/loopx:queue-round-<ts> 且 --replan-obligation-id replan-ca90e6ddfe013b35 的 successor,并 supersede 掉同内容的旧 successor,使链条不增长。

Activity

  1. huangruiteng commented on Sep 17, 2026

    @huangruiteng
    CollaboratorAuthor

    追加证据(同日 15:42Z–16:00Z,同一 lane,第三次/第四次重规划)

    在同一 lane 上继续观察,机制比首帖更明确,而且投影给出的所有补救动作都被 delta 校验拒绝:

    1. obligation id 每次求值都会重新生成

    时间 (UTC) obligation id 触发它的 lane 动作
    15:42 replan-c471cccdf3d0a276 —
    15:49 replan-ca90e6ddfe013b35 15:44 更新一条 lane todo 的证据
    15:56 replan-93faac7149153ea3 15:45–15:46 vision patch 写回 + 记账 + 归档
    15:58 replan-ef73bf962100e1da 15:51 新增/取代 successor
    15:59 replan-574027881d7ee84d 15:58 又一次 successor/取代

    也就是说:lane 每做一次写回(哪怕正是为了清除该义务),obligation 就换一个新 id;而 todo add --replan-obligation-id 会校验"当前 open obligation",所以上一刻读到的 id 在下一刻就可能失效(--replan-obligation-id does not match the current open obligation: expected replan-ef73bf962100e1da)。

    2. successor 回执要求"当前 obligation id + 最新一行",两者互相冲突

    loopx/control_plane/goals/goal_frontier/ack_policy.py::replan_successor_transition_ack 的匹配条件是:该 todo 为 open/advancement/本 lane 认领、replan_obligation_id 等于当前 obligation、具备 action_kind+target_key 的 typed 绑定,并且 long_todo_chain_transition_is_fresh(该 successor 的 updated_at ≥ 当前 frontier 中最新一行)。而第 1 点说明"当前 obligation"本身随每次写回变化,于是按上一个 id 建的 successor 在下次求值时已经作废;同时"必须是最新一行"意味着同一轮里任何后续的 todo 变更(例如 guidance 要求的 prune/supersede)都会让刚建好的绑定失效——我们实测到一次 todo supersede --next-agent-todo 生成的自动 successor 就作废了前一刻的绑定。

    3. refresh-state 的 ACK 不带 semantic_delta,永远无法绑定 trigger 检查点

    refresh-state --autonomous-replan-recorded 写出的 autonomous_replan_ack 只有 delta_contract(repair 语义),没有 semantic_delta;而 _acknowledged_replan_obligation_id 读的正是 semantic_delta.obligation_id,normalize_projected_autonomous_replan_ack 也要求 semantic_delta.accepted is True。结果是三次连续 observation 的 rearmed_after_obligation_id 都停在同一个历史 id(replan-3363450b0714266d),说明这些 ACK 从未被当作本轮抑制依据。

    4. delta 校验对三种"按投影照做"的交付都给出拒绝

    交付 拒绝理由
    vision patch + blocker(15:45) 记录成功但 semantic_delta 缺失,trigger 不抑制
    新建 typed successor(15:51、15:58、15:59) successor_or_supersede: no completed todo links a scoped open advancement successor / repeated stall requires a newly linked successor with a different direction
    记录 blocked(16:00) blocker: repeated stall fingerprint cannot close replan; create a new runnable direction or record exploration_exhausted with coverage evidence

    值得注意的是第三个:autonomous_replan_obligation.todo_actions 建议的动作正是"写一个具体 todo 或 vision delta",replan_action_packet.allowed_terminal 又允许 blocked,但两条路都被判为"不能关闭 replan"。

    影响(本次会话实测)

    15:42Z–16:00Z 之间,lane 每 3–7 分钟被要求重规划一次,期间只完成了一次真实评审交付(#4663);其余轮次都在与校验器对齐,并各自消耗一次 quota 结算。这正是"replan 死锁/抖动"的用户可见表现。

    建议(按修复成本从低到高)

    1. 让 obligation id 在触发指纹不变时保持稳定(trigger + frontier owned identity 不变就不换 id)。这样 successor 的 replan_obligation_id 绑定才有意义,replan_successor_transition_ack 才可能命中——这同时覆盖了"同一 lane 自己的写回不该让义务失效"。
    2. 让 refresh-state 的 ACK 也产出 semantic_delta(含 trigger_checkpoints,按 frontier_owned_identity 绑定),使 vision-patch / blocker 路径能像 successor 路径一样被 classifyAck 识别;否则该 flag 对 long_todo_chain 实际上是 no-op。
    3. 让 todo_actions 直接给出可执行的 typed successor(含 action_kind、target_key、task_repository、--replan-obligation-id),使"照投影执行"就能收敛。
    4. 若 blocked 属于允许的终态,则 repeated stall fingerprint cannot close replan 这条规则应与 allowed_terminal 对齐,或明确说明 blocked 只适用于非 replan 场景。

    复现

    同一 lane 连续三次 quota should-run:每次都会得到新的 obligation id;在两次之间按投影做一次 refresh-state --autonomous-replan-recorded(或建一个绑定到当前 id 的 successor),观察第三次的 obligation id 与 rearmed_after_obligation_id 是否变化。

  2. huangruiteng commented on Sep 17, 2026

    @huangruiteng
    CollaboratorAuthor

    Re-measured at 16:22Z with the same wake sequence; the original report still holds, but the remaining cause is narrower than "the ACK is never written".

    What is now ruled out. #4554 (fix(replan): keep checkpoint replan ACKs visible past the run window) is merged, so an ACK is no longer dropped because 20 material runs elapsed. The writeback itself does record: the 16:09Z and 16:22Z rounds both produced autonomous_replan_ack.recorded=true with delta_kinds=["successor_or_supersede"], and both slots settled normally.

    What the churn actually rides on. Three consecutive wakes (16:04Z, 16:09Z, 16:22Z) on this lane:

    wake frontier_revision frontier_owned_identity obligation_id
    16:04Z todo_frontier_revision_v0:0690216829a7807d19c22d92 null replan-328da7a9b816c400
    16:09Z todo_frontier_revision_v0:35e2dc8d61ddeb3ee1ccd7ce null replan-29553c67939e9fd7
    16:22Z todo_frontier_revision_v0:fd5eb5a44e7ee00516c1cc92 null replan-4bce1b3f1c569e2f

    frontier_owned_identity is null every time while this lane holds claimed advancement rows, so the ACK fence falls back to the shared selectable digest. loopx/control_plane/todos/frontier_revision.ts explains why: checkpoint() digests the selectable set (rows that are unclaimed or claimed by this agent) but derives frontier_owned_identity only from rows claim === agent; readIndex() then serves an agent without an owned entry matches[0] ?? index.unclaimed. classifyAck() suppresses the trigger only on an exact frontier_revision match or a non-null matching owned identity, so for this lane the fence is the digest over every unclaimed advancement row in the store.

    Consequence: any lane editing or claiming an unclaimed advancement row moves this lane's frontier_revision, which changes the obligation id (it is hashed from frontier_identity, which includes frontier_revision), which both re-mints the obligation and invalidates the ACK the lane just recorded. The suppression path is therefore unreachable for a lane with no owned rows: satisfying the replan requirement (a todo delta) moves the same digest that fences it. rearmed_after_obligation_id staying pinned at replan-3363450b0714266d across all of these is consistent — no ACK ever survives.

    Narrower statement of the defect. It is not that the ACK is invisible or unwritten; it is that the fence identity is not lane-scoped for an agent with no owned rows, so the edge-trigger contract degenerates into "re-arm on any store activity", and the well-known "excess above the threshold should be spaced, not re-required" behaviour holds only while nothing else touches the store.

    Candidate fix (bounded, for review). Make the fence identity lane-scoped the same way the measured count already is — i.e. derive an owned identity over the rows this agent's selection is computed from, not strictly claim === agent, so another lane's identical-shape churn stops re-arming this lane while a material change to this lane's own chain still does. That needs the row codec to carry the facts the predicate needs (this layer currently sees only id/claim/excluded/updated/serialized/advancement), plus a fixture built from the three observations above and a negative twin proving a material own-row change still re-arms.

    Filed as an agent Todo on lane codex-side-bypass (todo_7deaf5eae8d5, superseding a stale duplicate), so the work is tracked rather than re-derived each wake.

  3. huangruiteng commented on Sep 17, 2026

    @huangruiteng
    CollaboratorAuthor

    Correction/refinement after the next wake (16:37Z), because it changes what the fix is for.

    The previous comment argued the suppression path is unreachable and no ACK survives. That is still true for the revision fence, but it is not the whole story: the churn also has a sanctioned closeout I had not exercised. On the 16:30Z wake I recorded the owner decision as a real blocker todo and wrote back with --progress-result-class blocked --repair-delta-kind blocker (the delta validator accepted delta_kinds=["blocker"], recorded=true). On the next wake the payload contained zero long_todo_chain occurrences and quota should-run returned decision=run with no obligation at all — and its own spend policy names "concrete blocker writeback" as a valid spend basis. So todo add --task-class blocker + a blocked refresh-state is a working way out, which is the remedy the earlier report said was rejected.

    What that leaves as the actual defect is narrower and worth stating precisely, because it decides option A/B/C:

    • The obligation identity is hashed from frontier_identity, which contains the volatile frontier_revision. Observed revisions across four wakes: 0690…, 35e2…, fd5e…, cc2d… — a different id every time (replan-328da7a9, replan-29553c67, replan-4bce1b3f, replan-34bd1a0a), with rearmed_after_obligation_id pinned at replan-3363450b0714266d throughout.
    • The fence is the digest over the shared selectable pool with frontier_owned_identity null (derived only from claim === agent; readIndex serves an owner-less agent index.unclaimed, whose entry was computed with agent=null).
    • Therefore an agent that wants to keep working (rather than declare a blocker) cannot close the replan by producing a typed successor: satisfying the delta moves the same digest that fences it. That is the deadlock the fix has to address; the blocker path is a legitimate stop, not the spacing behaviour.

    So option B (make the fence and the counted quantity the same lane-scoped thing) remains my recommendation for the "keep working" path, with A (material/interval re-arm) as the alternative if the store-scoped trigger is intended. Evidence for the four revisions and the ids is in this issue's previous comment; the lane's tracked todo is todo_7deaf5eae8d5.

  4. huangruiteng commented on Sep 17, 2026

    @huangruiteng
    CollaboratorAuthor

    可用于收口的路径已复现:一个 typed successor + successor_or_supersede delta

    补一条能稳定清掉义务的做法,供决定 A/B/C 时参考(lane codex-side-bypass,2026-09-17T21:18Z):

    触发:本 wake 的 guard 给出 decision=autonomous_replan_required,obligation replan-9a9546983527fdd6,trigger long_todo_chain("current agent lane has a long selectable todo chain"),replan_settlement_contract.settlement_binding = {kind: todo, id: todo_9f68874dd840},semantic_obligation.settlement_bound=false,replan_action_packet.uncovered_frontier.required_any_of 含 new_runnable_successor。

    收口动作(本 turn 内完成,一次成功):

    1. 先落一个typed successor(--task-class advancement_task --action-kind <token> --target-key <token> --claimed-by <lane>),指向真实待办而不是占位;
    2. 再执行 refresh-state,逐字使用本 turn 的 settlement binding(这里是 --todo-id todo_9f68874dd840 --turn-instance-id <本 wake>),并带 --delivery-outcome outcome_progress --repair-delta-kind successor_or_supersede --autonomous-replan-recorded。

    结果:autonomous_replan_ack.recorded=true、delta_kinds=["successor_or_supersede"]、auto_evidence 直接列出刚落的 successor todo id;随后同一个 quota should-run 返回 decision=run、obligation=None、replan_required=false。

    为什么值得写进这里的决策

    • 这条路径证明"只要 lane 真有 successor 可落",投影义务是可以被语义清掉的,不需要走 blocker、也不需要 --replan-obligation-id(实测 --replan-obligation-id 与 --todo-id 互斥,而 settlement binding 恰好给的是 --todo-id,照抄投影会直接报错)。
    • 但前提是本 turn 内完成;跨 turn 无效。上一 wake 我已经落过同一个 successor,但当时 turn 的 writeback 已经 committed,再想用它 ACK 会得到 committed_writeback_payload_conflict: ... do not begin a new Turn or repeat spend to repair it,只能等下一个 wake 重来——这正是本 issue 首帖描述的"每做一次写回就换一个 obligation"的同一机制。
    • 因此若要修"投影指令无法清除义务",最小改动仍然是把这条可执行前提(落 typed successor + 用本 turn 的 --todo-id binding + successor_or_supersede)写进投影,而不是新造一条路径。
  5. huangruiteng commented on Sep 17, 2026

    @huangruiteng
    CollaboratorAuthor

    追加复现(同日 20:07Z–20:16Z,同一 lane)——这次重新武装来自 vision-checkpoint 家族,不是 long_todo_chain,但"投影指令清不掉义务"这一点完全一样

    观测

    1. 本 wake 的 guard 起初是 decision=run(绑定 todo_9f68874dd840),无 replan 义务。
    2. 按投影做一次常规 refresh-state(--classification pr_review_queue_readback --delivery-batch-scale single_surface --delivery-outcome outcome_progress --progress-scope agent_lane --progress-result-class advanced),随后 quota spend-slot 记账(140→141)。
    3. 记账后立刻再跑 quota should-run:decision=autonomous_replan_required,新 obligation replan-af89f76ea5246ba5,rearmed_after_obligation_id=replan-008a3a0b30cfa1e0,trigger_count=4,triggers 为 vision_acceptance_gap + vision_checkpoint_missing + 两个 outcome-checkpoint 类;replan_action_packet.required_any_of = fresh_vision_path_outcome / new_runnable_successor / new_concrete_blocker / coverage_backed_exploration_exhausted / coverage_backed_no_followup。
    4. 触发文本自己说明了原因:refresh-state closed a material segment without a per-agent vision decision; triggers=material_delivery_outcome。也就是说,同一次 refresh-state 只要没带 per-agent vision decision,就会被判成"material segment 未闭合"并立刻重新武装义务。

    投影指令在这个状态下必失败

    cli_channel.next_cli_actions[0] 给的是 --classification bounded_replan_progress ... --repair-delta-kind goal_vision_patch --autonomous-replan-recorded。照抄执行:

    • 带 --replan-obligation-id:--replan-obligation-id cannot be combined with --todo-id;而 replan_settlement_contract.settlement_binding 恰好给的就是 --todo-id,两个投影字段互斥。
    • 换成 --todo-id:committed_writeback_payload_conflict: committed writeback is unchanged; do not begin a new Turn or repeat spend to repair it. Retry the original delivery fields with only the missing vision decision。

    可用解是:逐字复用本 turn 原先的 delivery 字段(--classification pr_review_queue_readback,而不是投影给的 bounded_replan_progress),只额外加 --agent-vision-json,且不能再加 --repair-delta-kind / --autonomous-replan-recorded。这样 vision_checkpoint.decision 才从 missing_required 变为 patched,随后 should-run 回到 decision=run、replan_required=false、无 obligation。

    据此建议的三处小修

    1. 投影的 next_cli_actions 不应把 bounded_replan_progress 当作"修复已提交 writeback"的 classification——它与执行侧 committed_writeback_payload_conflict 契约直接矛盾,照做必失败。
    2. 错误信息要求的 "original delivery fields" 没有出现在投影里;应在修复指令中给出原字段,而不是让 agent 试错。
    3. 更根本的是第 4 点:material segment 的闭合要求要求同一次写回自带 per-agent vision decision,而常规 next_cli_actions 模板不带 --agent-vision-json/--vision-*,于是"按投影完成一次合法交付"本身就会制造下一个 replan 义务。这与 long_todo_chain 相互独立,但会造成同样的按 wake 空转。

    复现环境:lane codex-side-bypass / goal loopx-meta,实现为当前 origin/main(d8e7af141)的已安装 CLI。

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions