Skip to content

fix(ci): reconcile exact-target recovery and settled readback - #5567

Open
songoow wants to merge 6 commits into
loopx-project:mainfrom
songoow:codex/ci-recurring-failures-20261004
Open

songoow wants to merge 6 commits into
loopx-project:mainfrom
songoow:codex/ci-recurring-failures-20261004

Conversation

@songoow

@songoow songoow commented Oct 4, 2026 •

Copy link
Copy Markdown
Collaborator

Problem and outcome

The previous shard-2 run failed 14 cases: nine exact-Goal context return cases, three settled-monitor recovery cases, one source fingerprint oracle, and one failed-Turn session identity fixture. This PR now integrates their shared repairs into one candidate instead of requiring an unmerged dependency to make those cases pass.

Changes

  • Authorize an already validated exact Goal/Agent target through the existing TypeScript source-recipient policy. Keep sender, selected/blocked target, original instance and registration checks. Reuse one private grant observation helper for catalog and exact-target callers.
  • Preserve receipt-bound historical monitor identity in selected_todo and agent_lane_next_action after settlement, while retaining skip/no-run/no-spend authority. Update the protocol and existing replay assertions to explain this readback.
  • Align fixtures with current contracts: include release identity in the independent source fingerprint; challenge actual Agent identity drift; exercise four real grant modes and restoration; run local-Goal success fixtures inside their registered workspace.
  • Pin both CLI output comparison references to the triggering event's base SHA, including merge groups. Update the bilingual quality guide and regenerate only changed registry-I/O census locations.

The exact-target, settled-monitor and fixture repairs reuse work from #5533 by @Duang777. This candidate does not incorporate that PR's process-supervisor changes. Its remaining independent fixes still need their own acceptance.

Validation

Source: latest main 1af7dbd43 plus this PR, candidate fcfb24c65.

  • All 14 original failing cases passed together with two workers.
  • 404 related Python cases passed across the completed related-module run (314 unaffected-by-fixture cases) and the final 90-case Turn-driver run. The first expanded run had 12 out-of-root workspace fixture failures after the new main admission rule; the fixture repair and complete driver rerun are recorded, not hidden by a selected green retry.
  • 124 TypeScript source-grant, exact-Goal lifecycle and quota-readback cases passed.
  • Real legacy/File/SQLite CLI monitor recovery, no second spend, lost-response return, revoked-source no-send, restore-once and App snapshot identity paths are included. No frontend settings or interaction change is required: existing readback uses the original instance and unchanged execution-authority fields, exercised by the App snapshot tests.
  • Standard premerge passed all 19 selected checks plus direct checks, with one recorded inherited advisory: module_metric_budget:loopx/extensions/lark/goal_topic_runtime.py. The same failure was reproduced in an untouched 1af7dbd43 checkout; neither that module nor its ceiling is changed here.
  • CLI base/head budget against fixed 1af7dbd43, semantic inventory/census, configured mypy (19 source files), CI Ruff scope, diff and public-boundary checks passed. No budget allowance was raised.

Future-facing pass: share the source-grant adapter while leaving policy decisions in the existing typed owner; reuse the existing selected-Todo projector rather than another rule implementation. Runtime/API changes require maintainer review on the new head.

Historical scope

The audit covered all 100 merged PRs by the submitting author: 262 failed pull-request runs across 67 PRs, 494 failed job logs, and 418 pytest node IDs recurring across at least two PRs. This fixes the identified shared causes above; it does not claim every historical failure is a single bug or that all CI is green. Fifteen no-job runs expired awaiting maintainer approval, and 19 logged DCO failures lacked sign-offs; their checks remain enforced.

Signed-off-by: song <liusongstep@gmail.com>
Signed-off-by: song <liusongstep@gmail.com>
@songoow
songoow requested a review from huangruiteng as a code owner October 4, 2026 11:21
…failures-20261004

Signed-off-by: song <liusongstep@gmail.com>
Reuse the reviewed shared source-grant and settled-readback fixes from PR loopx-project#5533 by Duang777. Keep recipient policy in the TypeScript owner and expose only receipt-bound historical identity after settlement.

Signed-off-by: song <liusongstep@gmail.com>
Exercise Agent identity drift rather than legal cross-Todo session reuse, include release identity in the independent fingerprint oracle, and execute local Goal success fixtures inside their registered workspace. Retain settled no-run/no-spend assertions.

Signed-off-by: song <liusongstep@gmail.com>
Signed-off-by: song <liusongstep@gmail.com>
@songoow songoow changed the title fix(ci): pin comparison baselines and repair grant recovery qualification fix(ci): reconcile exact-target recovery and settled readback Oct 4, 2026
@songoow

songoow commented Oct 4, 2026

Copy link
Copy Markdown
Collaborator Author

Repair validation on exact head fcfb24c65472805d3a03d9d9fc200f1927edb3d7:

The 14 failures from the previous shard-2 run now pass together: 9 exact-target authorization cases, 3 settled-monitor recovery cases, the runtime fingerprint case and failed-Turn session identity case. Their fixes are integrated into this candidate, reusing the relevant shared work from #5533 by @Duang777.

Broader validation: 314 other related Python cases passed, then the complete 90-case Turn-driver module passed after preparing its success workspaces inside the registered Goal. The earlier 12 out-of-root fixture failures are preserved in the validation history. 124 typed policy/lifecycle/readback cases also passed. Revoked source no-send, restore-once, original instance readback, lost-response recovery, real File/SQLite monitor flows and no second spend are covered.

Standard premerge: 19 selected checks plus direct checks passed. The inherited advisory module_metric_budget:loopx/extensions/lark/goal_topic_runtime.py also reproduces in the untouched main 1af7dbd43; no ceiling was increased. CLI budget on the fixed base, semantic census, CI Ruff, configured mypy and public-boundary scans passed. Future-facing pass shares the existing grant observation and selected-Todo projector; typed policy authority remains with its existing owner.

The new head needs fresh CI and maintainer review. These local results neither certify the remaining #5533 supervisor changes nor constitute approval to merge this runtime change.

@loopx-agent loopx-agent left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewer: model_agent; model=gpt-6.1-sol; provider=OpenAI; declaration_source=runtime_reported; reasoning_effort=xhigh.

精确版本:5567@fcfb24c65472805d3a03d9d9fc200f1927edb3d7;实际 fork 基础版本:1af7dbd。首次独立全 PR 评审,未查询、轮询或等待 CI。

动机

由管家向已注册的精确 Goal/Agent 交付工作、随后读取结果的使用者,以及重放已结算监控 Turn 的执行者。

例如合法的来源已获准向 builder 交付:旧版因再次走不支持严格注册表的通用目录而拒绝,新版走既有来源策略后成功;监控结算后,新版为仍可重建的工作项恢复历史投影,但完成后的工作项仍只在原回执中可读。

实际 base/head 对照确认合法精确目标恢复、越权目标继续拒绝;monitor 的原结算身份和无执行、无重复扣额保持。新增文档的无条件 selected_todo 读回承诺在 completed 反例上未兑现。

本轮不把历史 CI 统计、测试通过或单个投影等同于完整产品验收,不接管来源账户、执行权限、benchmark 或其他 agent,也不授权合并该运行时 PR。

改动思路

这个修复的大方向合理:精确目标先由 Goal lifetime 与注册成员校验,再复用既有 TypeScript 来源策略;Python 只观察来源与存储,不另建权限判断源。普通目录与已经核验的精确目标共享 _source_context_grant,分别提供完整目录或单个合法候选,不因为同一个 GitHub 身份而混淆请求者或权限。返回仍绑定原来源、原 instance、原会话,并按当前 sender/selected/blocked policy 重新准入。

监控读回复用既有 compact projector,保留 no-run/no-spend 字段。这里必须区分永久结算事实与当前工作 lane:前者由已提交回执拥有,后者可以在完成时消失。新增文档把两个工作投影的可用性写得比生产分支更强,因此当前整体结论不能批准。

具体改动

完整 14 文件 +190/−67:四个生产文件整合来源 grant adapter、精确 return 的 Agent 注册检查与 settled 历史投影;六个 Python 测试文件修正 sender 撤权、source release fingerprint、Agent/session scope 与合法 workspace fixture;workflow 同时固定 budget 的 base/main selector 为 event SHA;质量指南中英同步说明固定比较基线;monitor 协议新增历史投影说明;registry I/O manifest 仅调整四处行号。没有新 CLI 选项、grant 数据库或 quota/scheduler 状态。

先读取基础修订 1af7dbd43a1629b0356ecd5ee28a426ad45d1d7a 的 docs/reference/protocols/quota-monitor-observation-receipt-v0.md 与 docs/architecture/rfcs/goal-instance-identity-and-orphan-recovery-v0.md。原 monitor 契约的 Original-settlement-identity、Settled-no-execution、Original-historical-receipt 三条均通过:原回执仍可读,完成不会重开 Turn,无新操作或第二次扣额。instance RFC 的来源权威、精确 lifetime 与 existing grants 约束由原 owner 保留。新增文档的无条件 selected_todo 承诺单独评估,不能用它替自己制造通过标准。

关键代码讲解

  1. source_grant_observation.py:36 的 _source_context_grant:把目录或单个目标交给同一个 resolveSourceRecipients,保留 verified sender、source digest、selected/all_registered 与 blocked 规则。新增 exact adapter 不读取任意目标注册事实;调用方必须先持有合法 Goal scope。
  2. manager_context/__init__.py:99 的 deliver:在 active/registered/lifetime/lifecycle 校验之后才调用 exact target authority,随后按原 request lock 写 inbox/route 并读回。实际反例中 selected-other、blocked Goal、移除 sender、未注册 Agent 都拒绝且没有 entry/registry 修改;两个合法来源在 head 成功,base 错误拒绝。
  3. manager_context/roundtrip.py:388 的 _exact_return_context:先核验原 result/route/instance,再验证注册 Agent 与当前 grant,最后由现有 return admission/verification owner 处理重试。相关真实 store restart、原 HTTP snapshot、撤权后不发送、恢复后只返回一次均经过本轮回归;外部 provider sender 在测试中是模拟边界,不能声称真实 Lark 网络发送已验收。
  4. settlement_precedence.py:121 的 apply_settled_monitor_precedence:只在当前 work lane 已识别为 settled monitor 时恢复匹配的旧 action 与 compact selected 投影。Monitor 完成后该 lane 不可重建,127–129 行提前返回;最终 CLI 仍安全 settled,却没有新增文档保证的 selected 字段。

对主干的风险

[P2] 新增历史投影契约应与完成后的实际读回一致。 位置:docs/reference/protocols/quota-monitor-observation-receipt-v0.md:80。本轮用同一 synthetic fixture 和真实 CLI 在基础版本、精确 head 都执行 monitor admission → poll/settlement → 正常 completion → 同 Turn guard/duplicate poll。head 结果是:原 heartbeat_receipt.settlement_identity 正确,should_run=false、must_attempt_work=false、无 executable CLI actions、journal 不变、duplicate replay 不追加;但 selected_todo 与 agent_lane_next_action 都缺失,当前 lane 也不含原 monitor。open 与 28 项其他 Agent 工作的 crowded 情况则成功恢复投影。

旧版也没有 completed selected 字段,因此没有把它说成新引入的扣额或执行回归;阻塞的是本 PR 新发布、超出实际行为的共享契约。消费者按新文字实现历史读回会遇到正常生命周期下的缺字段。最小修复是把文档明确为:原回执始终可读,两个工作投影仅在对应 monitor lane 可重建时出现,并同步中文与回归断言;若必须无条件提供 selected_todo,则由已提交回执派生它,补 completed/superseded/archived 实际 CLI 覆盖,保持无执行、无第二次扣额。

独立验证:五组相关模块 183 passed;完整 Turn driver 90 passed;source grant/lifecycle/quota readback 19 TS passed;control-plane typecheck、Ruff、配置内 19 文件 mypy、diff hygiene 均通过。真实 legacy/File/SQLite monitor 回归及 source-session/ChatHTTPServer 路径包含在其中。配对探针 6 个来源 scope、3 个 monitor 生命周期场景:相同 fixture SHA-256 f7487f3bb7183441ffecaba26756c3776c2e5c8ef7bc876ed7d206634a646323,base/head 观察分别 50aaab28ef0c4cbdd942f45ed87538c78431cfa1b3be6c36f607ec3bd5affa7a / d1bfa6898e0cd0474220ceae8f4bf1b71c3aff1ffb3f82fe7ebfe03ba09dccd9。九个场景中 completed 新字段存在性 oracle 失败;三个 monitor 的无执行、无扣额 oracle 全部通过。没有删除或弱化失败断言。

验证设置历史保留:初版探针误读 deliver 返回字段、误用通用 update 完成;首次 TS 命令错误假定安装 tsx。改为真实返回字段、专用 completion 入口及仓库 native Node strip-types 命令后重跑;这些不是生产缺陷,也不抹去当前 completed 反例。全仓库套件、native Windows、实际外部 provider、安装 App/模型采纳与长期收益未测。

语义与 CI 对齐

复用现有 typed source scope、GoalRef、ReceiptBoundMonitorPhase 和 effective action,没有新增共享 vocabulary 或 substring 分类。先跑 changed-from advisory(0 个支持语法候选,动态构造不在其证明范围),再跑 semantic smoke;均通过。代码默认行为变化已由 PR、fixture 与协议披露,未宣称 opt-in/default-off;历史投影不授予权限。当前语义问题是新增文档与实际生命周期不一致,semantic inventory 通过不能证明此承诺。

workflow 的两个比较 ref 固定到 pull-request base、merge-group base 或当前事件提交,避免队列等待时 main 漂移;未获取历史运行日志,也未把作者历史 CI 统计当独立当前证据。实际固定 base/main 为上述不可变基础修订执行 CLI output differential budget,通过;registry I/O census 检查为 current(278 sites)。没有改任何 hard limit。公共差异仅包含可复用产品/开发契约;既有 untracked lockfile 不在 PR。

我的整体评价

REQUEST_CHANGES。 来源授权修复及 monitor 无重复副作用的价值已得到独立验证,结构与工作量也合适。long_horizon 在本次恢复/重放上保持;user_experience 因新文档指向完成后缺失的历史字段仍未证明一致。最小修复可只把新契约范围写准确,而不扩展生命周期或新增存储。未来重构检查认为共用来源 adapter 与既有 projection 已足够,拒绝为此加入平行权限或历史状态 owner。

修改后在新精确 head 重跑 completed readback 与相关 monitor/return 回归,再评审完整 PR。本结论不受 CI 是否成功影响,也不表示原回执丢失或发生第二次扣额;尚无授权由本 reviewer 合并这个通用运行时 PR。

English verdict: REQUEST_CHANGES - 5567@fcfb24c65472805d3a03d9d9fc200f1927edb3d7; align the new unconditional historical selected_todo contract with completed-monitor CLI readback. Original settlement identity and no-run/no-spend are preserved; 273 Python and 19 TS tests pass, but the independent completed-field oracle fails. CI not consulted.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants