Conversation
Signed-off-by: song <liusongstep@gmail.com>
Signed-off-by: song <liusongstep@gmail.com>
…failures-20261004 Signed-off-by: song <liusongstep@gmail.com>
Reuse the reviewed shared source-grant and settled-readback fixes from PR loopx-project#5533 by Duang777. Keep recipient policy in the TypeScript owner and expose only receipt-bound historical identity after settlement. Signed-off-by: song <liusongstep@gmail.com>
Exercise Agent identity drift rather than legal cross-Todo session reuse, include release identity in the independent fingerprint oracle, and execute local Goal success fixtures inside their registered workspace. Retain settled no-run/no-spend assertions. Signed-off-by: song <liusongstep@gmail.com>
Signed-off-by: song <liusongstep@gmail.com>
|
Repair validation on exact head The 14 failures from the previous shard-2 run now pass together: 9 exact-target authorization cases, 3 settled-monitor recovery cases, the runtime fingerprint case and failed-Turn session identity case. Their fixes are integrated into this candidate, reusing the relevant shared work from #5533 by @Duang777. Broader validation: 314 other related Python cases passed, then the complete 90-case Turn-driver module passed after preparing its success workspaces inside the registered Goal. The earlier 12 out-of-root fixture failures are preserved in the validation history. 124 typed policy/lifecycle/readback cases also passed. Revoked source no-send, restore-once, original instance readback, lost-response recovery, real File/SQLite monitor flows and no second spend are covered. Standard premerge: 19 selected checks plus direct checks passed. The inherited advisory The new head needs fresh CI and maintainer review. These local results neither certify the remaining #5533 supervisor changes nor constitute approval to merge this runtime change. |
loopx-agent
left a comment
There was a problem hiding this comment.
Reviewer: model_agent; model=gpt-6.1-sol; provider=OpenAI; declaration_source=runtime_reported; reasoning_effort=xhigh.
精确版本:5567@fcfb24c65472805d3a03d9d9fc200f1927edb3d7;实际 fork 基础版本:1af7dbd。首次独立全 PR 评审,未查询、轮询或等待 CI。
动机
由管家向已注册的精确 Goal/Agent 交付工作、随后读取结果的使用者,以及重放已结算监控 Turn 的执行者。
例如合法的来源已获准向 builder 交付:旧版因再次走不支持严格注册表的通用目录而拒绝,新版走既有来源策略后成功;监控结算后,新版为仍可重建的工作项恢复历史投影,但完成后的工作项仍只在原回执中可读。
实际 base/head 对照确认合法精确目标恢复、越权目标继续拒绝;monitor 的原结算身份和无执行、无重复扣额保持。新增文档的无条件 selected_todo 读回承诺在 completed 反例上未兑现。
本轮不把历史 CI 统计、测试通过或单个投影等同于完整产品验收,不接管来源账户、执行权限、benchmark 或其他 agent,也不授权合并该运行时 PR。
改动思路
这个修复的大方向合理:精确目标先由 Goal lifetime 与注册成员校验,再复用既有 TypeScript 来源策略;Python 只观察来源与存储,不另建权限判断源。普通目录与已经核验的精确目标共享 _source_context_grant,分别提供完整目录或单个合法候选,不因为同一个 GitHub 身份而混淆请求者或权限。返回仍绑定原来源、原 instance、原会话,并按当前 sender/selected/blocked policy 重新准入。
监控读回复用既有 compact projector,保留 no-run/no-spend 字段。这里必须区分永久结算事实与当前工作 lane:前者由已提交回执拥有,后者可以在完成时消失。新增文档把两个工作投影的可用性写得比生产分支更强,因此当前整体结论不能批准。
具体改动
完整 14 文件 +190/−67:四个生产文件整合来源 grant adapter、精确 return 的 Agent 注册检查与 settled 历史投影;六个 Python 测试文件修正 sender 撤权、source release fingerprint、Agent/session scope 与合法 workspace fixture;workflow 同时固定 budget 的 base/main selector 为 event SHA;质量指南中英同步说明固定比较基线;monitor 协议新增历史投影说明;registry I/O manifest 仅调整四处行号。没有新 CLI 选项、grant 数据库或 quota/scheduler 状态。
先读取基础修订 1af7dbd43a1629b0356ecd5ee28a426ad45d1d7a 的 docs/reference/protocols/quota-monitor-observation-receipt-v0.md 与 docs/architecture/rfcs/goal-instance-identity-and-orphan-recovery-v0.md。原 monitor 契约的 Original-settlement-identity、Settled-no-execution、Original-historical-receipt 三条均通过:原回执仍可读,完成不会重开 Turn,无新操作或第二次扣额。instance RFC 的来源权威、精确 lifetime 与 existing grants 约束由原 owner 保留。新增文档的无条件 selected_todo 承诺单独评估,不能用它替自己制造通过标准。
关键代码讲解
source_grant_observation.py:36的_source_context_grant:把目录或单个目标交给同一个resolveSourceRecipients,保留 verified sender、source digest、selected/all_registered 与 blocked 规则。新增 exact adapter 不读取任意目标注册事实;调用方必须先持有合法 Goal scope。manager_context/__init__.py:99的deliver:在 active/registered/lifetime/lifecycle 校验之后才调用 exact target authority,随后按原 request lock 写 inbox/route 并读回。实际反例中 selected-other、blocked Goal、移除 sender、未注册 Agent 都拒绝且没有 entry/registry 修改;两个合法来源在 head 成功,base 错误拒绝。manager_context/roundtrip.py:388的_exact_return_context:先核验原 result/route/instance,再验证注册 Agent 与当前 grant,最后由现有 return admission/verification owner 处理重试。相关真实 store restart、原 HTTP snapshot、撤权后不发送、恢复后只返回一次均经过本轮回归;外部 provider sender 在测试中是模拟边界,不能声称真实 Lark 网络发送已验收。settlement_precedence.py:121的apply_settled_monitor_precedence:只在当前 work lane 已识别为 settled monitor 时恢复匹配的旧 action 与 compact selected 投影。Monitor 完成后该 lane 不可重建,127–129 行提前返回;最终 CLI 仍安全 settled,却没有新增文档保证的 selected 字段。
对主干的风险
[P2] 新增历史投影契约应与完成后的实际读回一致。 位置:docs/reference/protocols/quota-monitor-observation-receipt-v0.md:80。本轮用同一 synthetic fixture 和真实 CLI 在基础版本、精确 head 都执行 monitor admission → poll/settlement → 正常 completion → 同 Turn guard/duplicate poll。head 结果是:原 heartbeat_receipt.settlement_identity 正确,should_run=false、must_attempt_work=false、无 executable CLI actions、journal 不变、duplicate replay 不追加;但 selected_todo 与 agent_lane_next_action 都缺失,当前 lane 也不含原 monitor。open 与 28 项其他 Agent 工作的 crowded 情况则成功恢复投影。
旧版也没有 completed selected 字段,因此没有把它说成新引入的扣额或执行回归;阻塞的是本 PR 新发布、超出实际行为的共享契约。消费者按新文字实现历史读回会遇到正常生命周期下的缺字段。最小修复是把文档明确为:原回执始终可读,两个工作投影仅在对应 monitor lane 可重建时出现,并同步中文与回归断言;若必须无条件提供 selected_todo,则由已提交回执派生它,补 completed/superseded/archived 实际 CLI 覆盖,保持无执行、无第二次扣额。
独立验证:五组相关模块 183 passed;完整 Turn driver 90 passed;source grant/lifecycle/quota readback 19 TS passed;control-plane typecheck、Ruff、配置内 19 文件 mypy、diff hygiene 均通过。真实 legacy/File/SQLite monitor 回归及 source-session/ChatHTTPServer 路径包含在其中。配对探针 6 个来源 scope、3 个 monitor 生命周期场景:相同 fixture SHA-256 f7487f3bb7183441ffecaba26756c3776c2e5c8ef7bc876ed7d206634a646323,base/head 观察分别 50aaab28ef0c4cbdd942f45ed87538c78431cfa1b3be6c36f607ec3bd5affa7a / d1bfa6898e0cd0474220ceae8f4bf1b71c3aff1ffb3f82fe7ebfe03ba09dccd9。九个场景中 completed 新字段存在性 oracle 失败;三个 monitor 的无执行、无扣额 oracle 全部通过。没有删除或弱化失败断言。
验证设置历史保留:初版探针误读 deliver 返回字段、误用通用 update 完成;首次 TS 命令错误假定安装 tsx。改为真实返回字段、专用 completion 入口及仓库 native Node strip-types 命令后重跑;这些不是生产缺陷,也不抹去当前 completed 反例。全仓库套件、native Windows、实际外部 provider、安装 App/模型采纳与长期收益未测。
语义与 CI 对齐
复用现有 typed source scope、GoalRef、ReceiptBoundMonitorPhase 和 effective action,没有新增共享 vocabulary 或 substring 分类。先跑 changed-from advisory(0 个支持语法候选,动态构造不在其证明范围),再跑 semantic smoke;均通过。代码默认行为变化已由 PR、fixture 与协议披露,未宣称 opt-in/default-off;历史投影不授予权限。当前语义问题是新增文档与实际生命周期不一致,semantic inventory 通过不能证明此承诺。
workflow 的两个比较 ref 固定到 pull-request base、merge-group base 或当前事件提交,避免队列等待时 main 漂移;未获取历史运行日志,也未把作者历史 CI 统计当独立当前证据。实际固定 base/main 为上述不可变基础修订执行 CLI output differential budget,通过;registry I/O census 检查为 current(278 sites)。没有改任何 hard limit。公共差异仅包含可复用产品/开发契约;既有 untracked lockfile 不在 PR。
我的整体评价
REQUEST_CHANGES。 来源授权修复及 monitor 无重复副作用的价值已得到独立验证,结构与工作量也合适。long_horizon 在本次恢复/重放上保持;user_experience 因新文档指向完成后缺失的历史字段仍未证明一致。最小修复可只把新契约范围写准确,而不扩展生命周期或新增存储。未来重构检查认为共用来源 adapter 与既有 projection 已足够,拒绝为此加入平行权限或历史状态 owner。
修改后在新精确 head 重跑 completed readback 与相关 monitor/return 回归,再评审完整 PR。本结论不受 CI 是否成功影响,也不表示原回执丢失或发生第二次扣额;尚无授权由本 reviewer 合并这个通用运行时 PR。
English verdict: REQUEST_CHANGES - 5567@fcfb24c65472805d3a03d9d9fc200f1927edb3d7; align the new unconditional historical selected_todo contract with completed-monitor CLI readback. Original settlement identity and no-run/no-spend are preserved; 273 Python and 19 TS tests pass, but the independent completed-field oracle fails. CI not consulted.
Problem and outcome
The previous shard-2 run failed 14 cases: nine exact-Goal context return cases, three settled-monitor recovery cases, one source fingerprint oracle, and one failed-Turn session identity fixture. This PR now integrates their shared repairs into one candidate instead of requiring an unmerged dependency to make those cases pass.
Changes
selected_todoandagent_lane_next_actionafter settlement, while retaining skip/no-run/no-spend authority. Update the protocol and existing replay assertions to explain this readback.The exact-target, settled-monitor and fixture repairs reuse work from #5533 by @Duang777. This candidate does not incorporate that PR's process-supervisor changes. Its remaining independent fixes still need their own acceptance.
Validation
Source: latest main
1af7dbd43plus this PR, candidatefcfb24c65.module_metric_budget:loopx/extensions/lark/goal_topic_runtime.py. The same failure was reproduced in an untouched1af7dbd43checkout; neither that module nor its ceiling is changed here.1af7dbd43, semantic inventory/census, configured mypy (19 source files), CI Ruff scope, diff and public-boundary checks passed. No budget allowance was raised.Future-facing pass: share the source-grant adapter while leaving policy decisions in the existing typed owner; reuse the existing selected-Todo projector rather than another rule implementation. Runtime/API changes require maintainer review on the new head.
Historical scope
The audit covered all 100 merged PRs by the submitting author: 262 failed pull-request runs across 67 PRs, 494 failed job logs, and 418 pytest node IDs recurring across at least two PRs. This fixes the identified shared causes above; it does not claim every historical failure is a single bug or that all CI is green. Fifteen no-job runs expired awaiting maintainer approval, and 19 logged DCO failures lacked sign-offs; their checks remain enforced.