Repository navigation
[Bug] long_todo_chain replan 在合法 vision-patch 写回后 4 分钟即重新武装,且投影指令无法清除该义务 #4667
Description
Activity
追加证据(同日 15:42Z–16:00Z,同一 lane,第三次/第四次重规划)
在同一 lane 上继续观察,机制比首帖更明确,而且投影给出的所有补救动作都被 delta 校验拒绝:
1. obligation id 每次求值都会重新生成
时间 (UTC) obligation id 触发它的 lane 动作 15:42 replan-c471cccdf3d0a276— 15:49 replan-ca90e6ddfe013b3515:44 更新一条 lane todo 的证据 15:56 replan-93faac7149153ea315:45–15:46 vision patch 写回 + 记账 + 归档 15:58 replan-ef73bf962100e1da15:51 新增/取代 successor 15:59 replan-574027881d7ee84d15:58 又一次 successor/取代 也就是说:lane 每做一次写回(哪怕正是为了清除该义务),obligation 就换一个新 id;而
todo add --replan-obligation-id会校验"当前 open obligation",所以上一刻读到的 id 在下一刻就可能失效(--replan-obligation-id does not match the current open obligation: expected replan-ef73bf962100e1da)。2. successor 回执要求"当前 obligation id + 最新一行",两者互相冲突
loopx/control_plane/goals/goal_frontier/ack_policy.py::replan_successor_transition_ack的匹配条件是:该 todo 为 open/advancement/本 lane 认领、replan_obligation_id等于当前 obligation、具备action_kind+target_key的 typed 绑定,并且long_todo_chain_transition_is_fresh(该 successor 的updated_at≥ 当前 frontier 中最新一行)。而第 1 点说明"当前 obligation"本身随每次写回变化,于是按上一个 id 建的 successor 在下次求值时已经作废;同时"必须是最新一行"意味着同一轮里任何后续的 todo 变更(例如 guidance 要求的 prune/supersede)都会让刚建好的绑定失效——我们实测到一次todo supersede --next-agent-todo生成的自动 successor 就作废了前一刻的绑定。3. refresh-state 的 ACK 不带 semantic_delta,永远无法绑定 trigger 检查点
refresh-state --autonomous-replan-recorded写出的autonomous_replan_ack只有delta_contract(repair 语义),没有semantic_delta;而_acknowledged_replan_obligation_id读的正是semantic_delta.obligation_id,normalize_projected_autonomous_replan_ack也要求semantic_delta.accepted is True。结果是三次连续 observation 的rearmed_after_obligation_id都停在同一个历史 id(replan-3363450b0714266d),说明这些 ACK 从未被当作本轮抑制依据。4. delta 校验对三种"按投影照做"的交付都给出拒绝
交付 拒绝理由 vision patch + blocker(15:45) 记录成功但 semantic_delta缺失,trigger 不抑制新建 typed successor(15:51、15:58、15:59) successor_or_supersede: no completed todo links a scoped open advancement successor/repeated stall requires a newly linked successor with a different direction记录 blocked(16:00) blocker: repeated stall fingerprint cannot close replan; create a new runnable direction or record exploration_exhausted with coverage evidence值得注意的是第三个:
autonomous_replan_obligation.todo_actions建议的动作正是"写一个具体 todo 或 vision delta",replan_action_packet.allowed_terminal又允许blocked,但两条路都被判为"不能关闭 replan"。影响(本次会话实测)
15:42Z–16:00Z 之间,lane 每 3–7 分钟被要求重规划一次,期间只完成了一次真实评审交付(#4663);其余轮次都在与校验器对齐,并各自消耗一次 quota 结算。这正是"replan 死锁/抖动"的用户可见表现。
建议(按修复成本从低到高)
- 让 obligation id 在触发指纹不变时保持稳定(trigger + frontier owned identity 不变就不换 id)。这样 successor 的
replan_obligation_id绑定才有意义,replan_successor_transition_ack才可能命中——这同时覆盖了"同一 lane 自己的写回不该让义务失效"。 - 让 refresh-state 的 ACK 也产出
semantic_delta(含trigger_checkpoints,按frontier_owned_identity绑定),使 vision-patch / blocker 路径能像 successor 路径一样被classifyAck识别;否则该 flag 对long_todo_chain实际上是 no-op。 - 让
todo_actions直接给出可执行的 typed successor(含action_kind、target_key、task_repository、--replan-obligation-id),使"照投影执行"就能收敛。 - 若
blocked属于允许的终态,则repeated stall fingerprint cannot close replan这条规则应与allowed_terminal对齐,或明确说明 blocked 只适用于非 replan 场景。
复现
同一 lane 连续三次
quota should-run:每次都会得到新的 obligation id;在两次之间按投影做一次refresh-state --autonomous-replan-recorded(或建一个绑定到当前 id 的 successor),观察第三次的 obligation id 与rearmed_after_obligation_id是否变化。- 让 obligation id 在触发指纹不变时保持稳定(trigger + frontier owned identity 不变就不换 id)。这样 successor 的
Re-measured at 16:22Z with the same wake sequence; the original report still holds, but the remaining cause is narrower than "the ACK is never written".
What is now ruled out. #4554 (
fix(replan): keep checkpoint replan ACKs visible past the run window) is merged, so an ACK is no longer dropped because 20 material runs elapsed. The writeback itself does record: the 16:09Z and 16:22Z rounds both producedautonomous_replan_ack.recorded=truewithdelta_kinds=["successor_or_supersede"], and both slots settled normally.What the churn actually rides on. Three consecutive wakes (
16:04Z,16:09Z,16:22Z) on this lane:wake frontier_revision frontier_owned_identity obligation_id 16:04Z todo_frontier_revision_v0:0690216829a7807d19c22d92null replan-328da7a9b816c40016:09Z todo_frontier_revision_v0:35e2dc8d61ddeb3ee1ccd7cenull replan-29553c67939e9fd716:22Z todo_frontier_revision_v0:fd5eb5a44e7ee00516c1cc92null replan-4bce1b3f1c569e2ffrontier_owned_identityis null every time while this lane holds claimed advancement rows, so the ACK fence falls back to the shared selectable digest.loopx/control_plane/todos/frontier_revision.tsexplains why:checkpoint()digests the selectable set (rows that are unclaimed or claimed by this agent) but derivesfrontier_owned_identityonly from rowsclaim === agent;readIndex()then serves an agent without an owned entrymatches[0] ?? index.unclaimed.classifyAck()suppresses the trigger only on an exactfrontier_revisionmatch or a non-null matching owned identity, so for this lane the fence is the digest over every unclaimed advancement row in the store.Consequence: any lane editing or claiming an unclaimed advancement row moves this lane's
frontier_revision, which changes the obligation id (it is hashed fromfrontier_identity, which includesfrontier_revision), which both re-mints the obligation and invalidates the ACK the lane just recorded. The suppression path is therefore unreachable for a lane with no owned rows: satisfying the replan requirement (a todo delta) moves the same digest that fences it.rearmed_after_obligation_idstaying pinned atreplan-3363450b0714266dacross all of these is consistent — no ACK ever survives.Narrower statement of the defect. It is not that the ACK is invisible or unwritten; it is that the fence identity is not lane-scoped for an agent with no owned rows, so the edge-trigger contract degenerates into "re-arm on any store activity", and the well-known "excess above the threshold should be spaced, not re-required" behaviour holds only while nothing else touches the store.
Candidate fix (bounded, for review). Make the fence identity lane-scoped the same way the measured count already is — i.e. derive an owned identity over the rows this agent's selection is computed from, not strictly
claim === agent, so another lane's identical-shape churn stops re-arming this lane while a material change to this lane's own chain still does. That needs the row codec to carry the facts the predicate needs (this layer currently sees onlyid/claim/excluded/updated/serialized/advancement), plus a fixture built from the three observations above and a negative twin proving a material own-row change still re-arms.Filed as an agent Todo on lane
codex-side-bypass(todo_7deaf5eae8d5, superseding a stale duplicate), so the work is tracked rather than re-derived each wake.Correction/refinement after the next wake (16:37Z), because it changes what the fix is for.
The previous comment argued the suppression path is unreachable and no ACK survives. That is still true for the revision fence, but it is not the whole story: the churn also has a sanctioned closeout I had not exercised. On the 16:30Z wake I recorded the owner decision as a real
blockertodo and wrote back with--progress-result-class blocked --repair-delta-kind blocker(the delta validator accepteddelta_kinds=["blocker"],recorded=true). On the next wake the payload contained zerolong_todo_chainoccurrences andquota should-runreturneddecision=runwith no obligation at all — and its own spend policy names "concrete blocker writeback" as a valid spend basis. Sotodo add --task-class blocker+ a blocked refresh-state is a working way out, which is the remedy the earlier report said was rejected.What that leaves as the actual defect is narrower and worth stating precisely, because it decides option A/B/C:
- The obligation identity is hashed from
frontier_identity, which contains the volatilefrontier_revision. Observed revisions across four wakes:0690…,35e2…,fd5e…,cc2d…— a different id every time (replan-328da7a9,replan-29553c67,replan-4bce1b3f,replan-34bd1a0a), withrearmed_after_obligation_idpinned atreplan-3363450b0714266dthroughout. - The fence is the digest over the shared selectable pool with
frontier_owned_identitynull (derived only fromclaim === agent;readIndexserves an owner-less agentindex.unclaimed, whose entry was computed withagent=null). - Therefore an agent that wants to keep working (rather than declare a blocker) cannot close the replan by producing a typed successor: satisfying the delta moves the same digest that fences it. That is the deadlock the fix has to address; the blocker path is a legitimate stop, not the spacing behaviour.
So option B (make the fence and the counted quantity the same lane-scoped thing) remains my recommendation for the "keep working" path, with A (material/interval re-arm) as the alternative if the store-scoped trigger is intended. Evidence for the four revisions and the ids is in this issue's previous comment; the lane's tracked todo is
todo_7deaf5eae8d5.- The obligation identity is hashed from
可用于收口的路径已复现:一个 typed successor +
successor_or_supersededelta补一条能稳定清掉义务的做法,供决定 A/B/C 时参考(lane
codex-side-bypass,2026-09-17T21:18Z):触发:本 wake 的 guard 给出
decision=autonomous_replan_required,obligationreplan-9a9546983527fdd6,triggerlong_todo_chain("current agent lane has a long selectable todo chain"),replan_settlement_contract.settlement_binding = {kind: todo, id: todo_9f68874dd840},semantic_obligation.settlement_bound=false,replan_action_packet.uncovered_frontier.required_any_of含new_runnable_successor。收口动作(本 turn 内完成,一次成功):
- 先落一个typed successor(
--task-class advancement_task --action-kind <token> --target-key <token> --claimed-by <lane>),指向真实待办而不是占位; - 再执行
refresh-state,逐字使用本 turn 的 settlement binding(这里是--todo-id todo_9f68874dd840 --turn-instance-id <本 wake>),并带--delivery-outcome outcome_progress --repair-delta-kind successor_or_supersede --autonomous-replan-recorded。
结果:
autonomous_replan_ack.recorded=true、delta_kinds=["successor_or_supersede"]、auto_evidence直接列出刚落的 successor todo id;随后同一个quota should-run返回decision=run、obligation=None、replan_required=false。为什么值得写进这里的决策
- 这条路径证明"只要 lane 真有 successor 可落",投影义务是可以被语义清掉的,不需要走 blocker、也不需要
--replan-obligation-id(实测--replan-obligation-id与--todo-id互斥,而 settlement binding 恰好给的是--todo-id,照抄投影会直接报错)。 - 但前提是本 turn 内完成;跨 turn 无效。上一 wake 我已经落过同一个 successor,但当时 turn 的 writeback 已经 committed,再想用它 ACK 会得到
committed_writeback_payload_conflict: ... do not begin a new Turn or repeat spend to repair it,只能等下一个 wake 重来——这正是本 issue 首帖描述的"每做一次写回就换一个 obligation"的同一机制。 - 因此若要修"投影指令无法清除义务",最小改动仍然是把这条可执行前提(落 typed successor + 用本 turn 的
--todo-idbinding +successor_or_supersede)写进投影,而不是新造一条路径。
- 先落一个typed successor(
追加复现(同日 20:07Z–20:16Z,同一 lane)——这次重新武装来自 vision-checkpoint 家族,不是
long_todo_chain,但"投影指令清不掉义务"这一点完全一样观测
- 本 wake 的 guard 起初是
decision=run(绑定todo_9f68874dd840),无 replan 义务。 - 按投影做一次常规
refresh-state(--classification pr_review_queue_readback --delivery-batch-scale single_surface --delivery-outcome outcome_progress --progress-scope agent_lane --progress-result-class advanced),随后quota spend-slot记账(140→141)。 - 记账后立刻再跑
quota should-run:decision=autonomous_replan_required,新 obligationreplan-af89f76ea5246ba5,rearmed_after_obligation_id=replan-008a3a0b30cfa1e0,trigger_count=4,triggers 为vision_acceptance_gap+vision_checkpoint_missing+ 两个 outcome-checkpoint 类;replan_action_packet.required_any_of=fresh_vision_path_outcome/new_runnable_successor/new_concrete_blocker/coverage_backed_exploration_exhausted/coverage_backed_no_followup。 - 触发文本自己说明了原因:
refresh-state closed a material segment without a per-agent vision decision; triggers=material_delivery_outcome。也就是说,同一次 refresh-state 只要没带 per-agent vision decision,就会被判成"material segment 未闭合"并立刻重新武装义务。
投影指令在这个状态下必失败
cli_channel.next_cli_actions[0]给的是--classification bounded_replan_progress ... --repair-delta-kind goal_vision_patch --autonomous-replan-recorded。照抄执行:- 带
--replan-obligation-id:--replan-obligation-id cannot be combined with --todo-id;而replan_settlement_contract.settlement_binding恰好给的就是--todo-id,两个投影字段互斥。 - 换成
--todo-id:committed_writeback_payload_conflict: committed writeback is unchanged; do not begin a new Turn or repeat spend to repair it. Retry the original delivery fields with only the missing vision decision。
可用解是:逐字复用本 turn 原先的 delivery 字段(
--classification pr_review_queue_readback,而不是投影给的bounded_replan_progress),只额外加--agent-vision-json,且不能再加--repair-delta-kind/--autonomous-replan-recorded。这样vision_checkpoint.decision才从missing_required变为patched,随后should-run回到decision=run、replan_required=false、无 obligation。据此建议的三处小修
- 投影的
next_cli_actions不应把bounded_replan_progress当作"修复已提交 writeback"的 classification——它与执行侧committed_writeback_payload_conflict契约直接矛盾,照做必失败。 - 错误信息要求的 "original delivery fields" 没有出现在投影里;应在修复指令中给出原字段,而不是让 agent 试错。
- 更根本的是第 4 点:material segment 的闭合要求要求同一次写回自带 per-agent vision decision,而常规
next_cli_actions模板不带--agent-vision-json/--vision-*,于是"按投影完成一次合法交付"本身就会制造下一个 replan 义务。这与long_todo_chain相互独立,但会造成同样的按 wake 空转。
复现环境:lane
codex-side-bypass/ goalloopx-meta,实现为当前origin/main(d8e7af141)的已安装 CLI。- 本 wake 的 guard 起初是
现象
lane 的
long_todo_chainreplan 义务在一次合法交付之后 4 分钟就重新武装并换了新的 obligation id,而该 lane 在此之前刚刚按 guard 投影的指令完成了 bounded replan 写回。结果是"每次唤醒都要求重规划"的循环。时间线(同一 lane
codex-side-bypass, goalloopx-meta)replan-c471cccdf3d0a276,triggerlong_todo_chain(selectable 21 > 阈值 20),frontier_revisiontodo_frontier_revision_v0:506463d15b19380633011591todo_4144c9deec0f)refresh-state --classification bounded_replan_progress --autonomous-replan-recorded --repair-delta-kind goal_vision_patch --repair-delta-kind blocker --agent-vision-json <packet>:写入 vision patch(含 acceptance_summary 与goal_path_delta_v0, outcome=replan)并记录autonomous_replan_ack(recorded: true)replan-ca90e6ddfe013b35,同一 trigger,frontier_revision 变为todo_frontier_revision_v0:7ba751eda7da58a9f6601e2a,rearmed_after_obligation_id: replan-3363450b0714266d(一个更早的 id,不是刚刚交付的c471ccc…)机制
loopx/control_plane/todos/frontier_revision.ts::classifyAck要求delta.trigger_checkpoints[]中存在frontier_revision等于当前 revision,或frontier_owned_identity等于当前 owned identity。loopx/control_plane/goals/goal_frontier/ack_policy.py会为它写入replan_semantic_delta_v0(outcomesnew_runnable_successor、obligation id、successor todo id、trigger_checkpoints)。而replan_successor_semantic_binding要求 successor 同时具备action_kind和target_key(或 explore refs)。todo_actions[0].text= "run a bounded long-chain vision replan: compare evidence with the active vision, group or prune the todo chain, and select the next high-value runnable slice" —— 直接todo add这段文字不会带上action_kind/target_key,因此不能被认定为 typed successor;cli_channel.next_cli_actions[0]只给refresh-state ... --autonomous-replan-recorded(外加 progress/vision 字段),照做会得到recorded: true但没有trigger_checkpoints的 ACK,于是 trigger 继续有效。也就是说:按照 guard 投影的指令字面执行,无法清除
long_todo_chain义务;只有额外知道"必须落一个带 action_kind+target_key 的 successor todo"才能清除。这一点目前只存在于实现里,没有出现在投影指令中。影响
lane 每 3–4 分钟被要求重规划一次,真正的推进工作被挤出;同时每次重规划都会写入 run/vision 记录并占用一次 quota 结算。同一 lane 在本次会话的另一条记录里已被用户提醒过该模式。
建议
三者取其一即可,且都属小改动:
todo_actions直接给出可执行的 typed successor(带action_kind、target_key、task_repository),使照做即可清除义务;todo_actions/cli_channel中明确写出"需要落一个带 action_kind+target_key 的 successor todo"这一前提;frontier_owned_identity绑定,避免自己的写回改动 revision 后立刻自失效)。复现
在一个 selectable open todo > 20 的 lane 上:执行
quota should-run得到 obligation A → 仅按投影指令做refresh-state --autonomous-replan-recorded(vision patch 或 blocker delta)→ 等待下一次quota should-run:会出现新的 obligation id,且rearmed_after_obligation_id指向更早的 obligation。本 lane 的绕行做法(已执行):用
todo add建立带action_kind=review_pull_requests+target_key=github-pr-review:huangruiteng/loopx:queue-round-<ts>且--replan-obligation-id replan-ca90e6ddfe013b35的 successor,并 supersede 掉同内容的旧 successor,使链条不增长。