Repository navigation
fix: make settlement inputs and full Todo readback discoverable - #6052
Conversation
Signed-off-by: LoopX Agent <337587101+loopx-agent@users.noreply.github.com>
loopx-agent
left a comment
There was a problem hiding this comment.
Approval conclusion (author-owned PR; GitHub blocks formal self-approval)
Reviewer: model_agent | gpt-6.1-sol | OpenAI | runtime_reported | reasoning_effort=xhigh
APPROVE,评审 head:b00ea4e5832e258bcada1056ac4d3ec35cc1e86f。未发现阻断问题;批准的是首次结算输入和完整 Todo 读回的可发现性,不代表模型效果、安装采用或合并授权。
动机
普通 CLI Agent 在规划任务、首次保存已验证成果时,需要读到完整要求并填写当前方向与验收判断。
旧包已有简短结算提醒,但没有直接提供结构化填写示例;规划清单又裁剪正文,Agent 需要另找填写和完整读回的方法。
本次把现有填写契约放到首次写回步骤,并在规划后给出按 ID 读回完整 Todo 的命令。
范围仅是操作引导;不改变任务完成、权限、验收和扣额规则,也不证明模型采用或性能收益。
独立 base/head 实测确认上述两项发现路径从缺失变为存在。相同合成长 Todo 的最后一条要求能从原生详情完整读回;使用当前证据替换示例后,首次写回、持久结果读回和一次记账能完成。先执行 checkpoint-context、先扣额及超预算输入仍被拒绝;继续工作可以保持 Todo open,下一 Turn 继续选回同一项。收益是减少查契约和误用恢复入口的机会,而不是减少必需验证。
改动思路
复用现有 TypeScript vision owner 和已有 CLI 路由适配器,能够补齐发现路径而不新增状态、权限或重复决策源。
本 PR 的边界是普通 CLI/TurnEnvelope 的填写引导及共享规划的精确读回;MCP 保持自己的 complete_task 执行方式。
“不改代码、只补文档”仍要求普通调用者离开当前包查找;新建 capability 或恢复流程会增加重复 authority。现在的派生投影更贴近缺口。相邻重构检查已应用现有 owner 复用;未见需要额外抽象的同因变化。
独立验收依据为 docs/reference/protocols/goal-vision-replan-contract-v0.md,固定 revision 0012c4eb52191ba58d510e4af8cabb46704c1539,按修改前契约判断:
- Ownership Boundary:TS 持有状态与验收判断;Python 仅保留命令和载荷适配。示例和读回命令不授予 guard/lease authority。
- CLI Budget:原有字段、1800 字符总预算和 reject-without-truncation 保留。420 字符 acceptance 的越界输入在三种存储上均拒绝且不扣额。
- Vision Checkpoint:普通 material closeout 仍需验收声明及证据关联路径;in-flight 不要求新的 vision。实际继续工作保留 open/claim,并在下一 Turn 正确续接。
- Read basis for checkpoint-only recovery:checkpoint-context 仍要求原 Turn 已提交写回;首次填写先走 refresh-state。恢复不能被改成新 Turn 的预检。
- Write / Correction Mechanism:当前 Agent 和原结算身份保持;结构化输入不绕过原 validator,示例事实必须由当前证据替换。
相关已合并 #5633 的简短提示、#5841 的文字语义与预算规则已核对;打开的 #5677 属于读依据保护及原 Turn 恢复,不能当作本 PR 已交付的功能。收到的经验建议促使本次检查首次有用结果和恢复读回;经验的存在不证明评审质量或模型收益。
具体改动
12 文件,194 增 / 14 删:生产适配和投影 32 增 / 2 删,测试 135 增 / 3 删,文档 26 增 / 8 删,生成 census 坐标 1 增 / 1 删。
turnScopedCliSettlementPlan(loopx/control_plane/quota/settlement_plan.ts:11):Todo 绑定的 durable_writeback 使用visionAuthoringContract();独立 replan 已有该契约,避免再次复制。validation → writeback → spend → 条件性 terminal 顺序不变。visionAuthoringContract(loopx/control_plane/goals/vision_checkpoint.ts:76):沿用 validator 的字段和最小示例,补充 ordinary CLI 参数,以及 checkpoint-context 的原提交/同 Turn 恢复时序。已有 material-closeout 义务不被改成可忽略建议。SettlementStep/settlementStepPayload(Python 与 TS):传递可选vision_authoring,以及 quota Python bridge 读回该字段。没有新的持久授权声明;旧载荷不含字段仍合法。todo_authoring_steps(loopx/control_plane/goals/start_goal_todo_delta.py:132):在写入/调整任务后加按 ID 的完整读回。复用 runtime/registry/Goal/Agent 路由,带空格路径的生成命令实际可执行;要求唯一匹配、source_complete、全文与 status/claim 校验。缺失/变化需重查,不将摘要当完整来源。- 契约文档明确默认引导变化;原生 CLI 测试覆盖 legacy Markdown、File、SQLite 与两个交付边界;规划与 TS 测试校验调用投影。registry I/O manifest 只随源代码位置重生成,不改变访问行为。
独立验证:结算/TurnEnvelope/规划 123 项,规划压缩与预算 95 项,TS vision/settlement/MCP/effect 42 项均通过;TypeScript typecheck、Ruff、diff check、diff semantic advisory、完整 semantic smoke 通过。另用同一合成输入在 base/head 实际 CLI 执行六组配对场景:完整正文、缺失 ID、提前恢复/扣额、越界/合法首次写回、独立持久读回、重复扣额,以及继续工作的下一 Turn。发现性 oracle 在 base 预期失败、head 通过;原有 admission、诊断、身份和记账观察相同。再加入 40 项无关已完成工作并重排,精确目标读回保持一致;完成后的目标仍能读回,done 本身不被当成 Goal 验收。
早期 reviewer 探针将拒绝结果误读为统一 error 字段,修正为实际 reason/failure 后重跑两端;第一次 semantic 命令用错目录,实际脚本随后通过。二者是验证调用问题,不归因于 PR。作者初始 claim fixture 与预算失败记录也保留;当前独立结果覆盖原问题,不以挑选一次绿色运行替代判断。按当前评审策略未查询或等待 CI。
对主干的风险
这是默认引导的行为变化,不是可选 TurnEnvelope 能力:关闭该能力也会在普通 CLI 得到新引导。文档和 PR 已说明,不能声称 default-off 隔离。MCP 仍移除手工结算计划,用 complete_task;replan 不重复添加新 authoring 字段,原判断及恢复流程保留。
主要代价是每次包更大。独立固定输入按紧凑 UTF-8 JSON 测量:普通结算 4741→6853 字节,in-flight 3494→5606 字节,规划 6560→7372 字节;即结算 +2112、规划 +812。作者的固定 guided fixture 另显示 issue Markdown 3309→3757 字符。3600→4100 与 compact ratio .65→.66 是输出回归预算调整,未放宽 vision 硬限制、任务配额或实验阈值;generic 3200 和重复提示断言仍保留并通过。删掉必要填写/原 Turn 恢复语义来凑旧上限,不如保留可执行指引。
本 PR 没有新增用户配置或前端操作;已有前端 acceptance observation 和 capability settings 已检视,状态消费契约未改变。本次验证对象是 agent-facing CLI 和 MCP 投影,不声称重新验证整个打包前端。也没有 authority-store 重构;没有把本地三种存储结果冒称 PostgreSQL 集成通过。实际模型是否遵循示例、长期 token/延迟、首次成功率及安装采用仍未测。示例不能作为当前事实或验收证据;这些剩余限制不被结构检查或批准结论消除。
我的整体评价
批准该有界发现性改进。长程正确性方面,原始要求、跨 Turn 工作和一次记账保留;体验方面,普通包给出首次填写与完整读回的可执行路径,避免把恢复入口误作预检。效率方面存在明确包体开销,预期只有在减少查契约或返工时才会获益;目前不能宣称总体成本下降。该权衡在不改变验收/权限且沿用单一规则 owner 的边界内合理。
控制面变更仍由 maintainer 合并;本次未合并、安装升级、改变实验运行或扩大权限。
English verdict: APPROVE - b00ea4e. Existing owners now expose usable first-write and exact Todo readback guidance; independently validated 260 focused tests plus paired real CLI/store, replay and continuation probes. Serialized packets grow; live model adoption, token savings and installed behavior remain unverified. Maintainer merge required.
Signed-off-by: LoopX Agent <337587101+loopx-agent@users.noreply.github.com>
loopx-agent
left a comment
There was a problem hiding this comment.
Approval conclusion (author-owned PR; GitHub blocks formal self-approval)
Reviewer: model_agent | gpt-6.1-sol | OpenAI | runtime_reported | reasoning_effort=xhigh
评审准确 head:7dd2ec602d3c06060061a8352423b944af2c551f;冻结 base:0012c4eb52191ba58d510e4af8cabb46704c1539。
动机
普通 CLI Agent 在首次保存成果或定期重新判断工作方向时,需要填写现有的验收与证据字段。
旧普通结算包没有填写示例,默认重规划又把必需输入移到另一次详情读取;规划清单只有摘要,调用者还需寻找完整正文命令。
当前包直接提供既有填写契约,并在规划后给出精确 Todo 读回路径;独立实测确认调用者能提交有证据的输入、保持原身份结算及读取完整要求。
本 PR 交付输入可发现性,不声明模型采用、成本下降、整个 Goal 完成或安装升级。
本次先停止对旧批准 head 的资格核对,再按整个 base..head 重新判断。上次 b00ea4e5832e258bcada1056ac4d3ec35cc1e86f 只补齐普通写回的发现路径,默认 obligation-only 重规划仍需冷读;新提交补齐后者,并只消除同一响应内完全相同的重复 schema。旧批准和旧测试数量不作为当前批准依据。
改动思路
沿用现有 TypeScript Vision owner 和 Python CLI 投影适配器;填写输入应随执行步骤可得,真正重复内容可在同一响应内引用,诊断审计仍按需展开。
当前 PR 边界是普通/继续工作结算、默认重规划及共享规划读回的输入发现,不新增决策源、持久状态、配置开关或权限。
最强的不交付理由是默认包变大,且尚未证明模型理解或实际节省成本。只写外部文档仍不能修复当前包的发现路径;强行压回旧预算会继续隐藏所需输入。选择既有 owner 的完整契约加安全同响应引用,比另建 capability 或复制 validator 更小。预算增加只保护这项明确默认变化,不改变 Vision 硬限制或 quota。
相邻重构检查已应用单一 owner 复用与真正重复 schema 的消除。旧冷读 helper 被替换,当前调用点和文档一起更新;保留不同 schema 的内容,无需另加通用同步框架。
具体改动
完整变更 18 文件:生产/预算适配 66 增、12 删,文档 40 增、14 删,测试 281 增、24 删,生成 census 坐标 1 增、1 删。
独立规格为冻结 base 的 goal_vision_replan_contract_v0 与 Quota CLI Hot-Path Compaction v0。按修改前要求判断:
- Ownership Boundary:TS 继续拥有 Vision/结算判断;Python 只投影、路由和引用同一契约。
- CLI Budget:420 字符 acceptance 与 1800 总字符等原硬限制保留;越界拒绝,不截断。
- Vision Checkpoint:普通 material closeout、合法 in-flight、checkpoint satisfied 与 Goal acceptance 保持区别;继续工作不被新增 schema 强迫修改 Vision。
- Read basis for checkpoint-only recovery:checkpoint-context 仍只服务原已提交 Turn,提前恢复/扣额仍拒绝。
- Qualification Contract:实际 CLI 配对与投影验证必需填写输入及执行 authority;包大小独立测量。原文的 schema 冷读被本 PR 明确改变为内联/同响应引用,不能宣称这段文字等价或模型资格已完成。
关键代码讲解
turnScopedCliSettlementPlan(settlement_plan.ts:11)在 Todo 绑定的 durable_writeback 带上visionAuthoringContract();普通与 in-flight 共用输入,原 validation→writeback→spend 顺序、原身份与条件性 terminal 保持。visionAuthoringContract(vision_checkpoint.ts:76)复用 validator 限制与示例,明确--agent-vision-json和原 Turn 恢复顺序。TS/PythonSettlementStep序列化桥仅传递可选字段;无字段的旧载荷仍合法。_reference_inline_replan_authoring(cli_projection.py:465)只在已有 durable_writeback 的 schema 逐值相等时使用精确同响应路径。找不到步骤、不同 schema 或错误步骤保留完整输入;progress-only 重规划不获得新的 Vision 要求。明确 Vision detail 不做该替换,TurnEnvelope 保留完整输入。todo_authoring_steps(start_goal_todo_delta.py:132)在新建/复用步骤后提供带 registry/runtime/Goal/Agent 的精确读回命令。读完整.todo.text、source completeness、status/claim;缺失或歧义要求重查,摘要不是正文来源。读取不授予 guard 或 lease。- 预算和两份协议文档披露默认改变;测试覆盖两种结算边界、三种实际本地存储、重规划与引用反例。I/O manifest 只更新源代码坐标。
当前独立源码七文件 240 passed;TS 59 项、TS typecheck、Ruff、diff check、diff semantic advisory、完整 semantic smoke 和 290-site registry I/O manifest 均通过。验证使用当前 checkout 与真实本地 Legacy/File/SQLite 路径,不借 CI 或作者数量认证。
独立三版本真实 CLI 探针:base 没有普通示例且默认重规划输入需冷读;旧批准 head 有普通示例但重规划仍需冷读;当前 head 两者均在当前包可得。三版本均拒绝缺证据重规划,纠正后可完成第二 Turn 的持久写回与一次扣额,再重放不新增扣额;恢复提示和 rule 相同。该 oracle 在旧版本失败、当前通过,不用新输出定义期望。
独立 base/head 长正文探针:完整尾部、claim 保留,加入 40 项无关已完成任务与改变展示顺序不改变目标读回;目标完成后仍可读原文,缺失 ID 明确返回 not_found。原命令本来就能读完整正文;新增的是规划后的发现路径,不宣称修复了原生正文丢失。
对主干的风险
这是默认 CLI 引导改变,无 TurnEnvelope 开关也会得到新内容,不属于 default-off 功能。MCP 继续移除手工 CLI 结算过程并使用自己的同源 authoring;前端/Lark 没有新增用户操作或配置。本次不宣称重新验证打包前端、PostgreSQL、真实模型或线上安装。
独立相同路径/相同输入测量:小场景 19,720 字符、520 行不变;多 Agent 21,900 字符、604 行不变;拥挤 required-replan 37,637→40,686 字符、828→922 行,即固定 +3,049 字符/+94 行。这里路径较长,绝对值不能冒称作者短路径 fixture 的预算结果;增长及保留决策是在相同路径配对测得。新 crowded 39k/950 行、固定增长 9k 和 rich Vision 45k 是回归 fixture 预算,其他 lanes/per-Todo 限制保留;1800 Vision 输入等硬限制未改。必要语义换取包体开销合理,但真实 token、延迟和收益未测。
判断、准入、replay、写回 authority 与 once-only debit 未新增规则。同响应引用依赖真正相同且保留的步骤,不能拿全局 locator 或后来读取替代当前输入。独立相关 caller 的 main 源码与冻结 base 同 blob,近期 manager/release-doc 合并未使该判断失效。当前 source 全树 semantic 通过仍有已登记动态 producer 覆盖限制;advisory 0 不证明动态语义安全。
我的整体评价
APPROVE 当前有界发现性改进。沿用现有 TypeScript Vision owner 和 Python CLI 投影适配器;填写输入应随执行步骤可得,真正重复内容可在同一响应内引用,诊断审计仍按需展开。当前 PR 边界是普通/继续工作结算、默认重规划及共享规划读回的输入发现,不新增决策源、持久状态、配置开关或权限。
持续工作方面,原验收、后续 Turn、拒绝后的纠正和一次记账保持;操作体验方面,首次填写与完整读回无需另寻命令。尚未证明模型采用、长期成本或产品整体验收,不能把输入可得等同于模型已用。保持维护者合并边界,本次不合并或升级。
English verdict: APPROVE - 7dd2ec6. Reused authoring inputs are available on ordinary and default replan paths, with safe same-response deduplication and complete Todo readback discovery. Independent old/base/head CLI probes preserve rejection, recovery and one debit per Turn; required replan adds 3,049 characters/94 lines. 240 Python tests, 59 TS tests, typecheck and static/semantic checks pass. Live adoption and installed behavior remain unverified. Maintainer merge required.
loopx-agent
left a comment
There was a problem hiding this comment.
Reviewer: model_agent | gpt-6.1-sol | OpenAI | runtime_reported | reasoning_effort=xhigh
Request changes conclusion (author-owned PR; GitHub blocks formal self-review)
REQUEST_CHANGES — exact head 7dd2ec602d3c06060061a8352423b944af2c551f。本次重新检查整个 PR;旧 b00 批准不覆盖新增的默认 replan 输入。
[P2] 补齐新增输出的 absolute 与 differential 预算资格。 新增 read_back_authored_todos(loopx/control_plane/goals/start_goal_todo_delta.py:152–165)也进入 start-goal --guided --include-command-pack-detail 的 Markdown 输出。独立同路径、同输入真实 CLI:冻结 base 0012c4e 为 2781 字符 / 35 行,当前 head 为 3217 / 37;该模式的字符上限仍是 3200(loopx/control_plane/testing/cli_output_budget.py:386)。本 PR 的原生 premerge canary 因此失败;在两个独立 worktree 单独运行同一 examples/control_plane/cli-output-budget-regression-smoke.py,base 通过、head 再次失败。输出差异正是新增完整读回步骤,不能归为未变的祖先夹具或基础设施问题。
进一步用原生产 probe 的 measurement-only 模式收集完整、同路径 base/head 响应(未修改任何产品/验收上限),再交给现有 compare_cli_output_receipts:10 行差异预算失败。guided 三种规模及其 detail 模式增加 JSON 948 字符 / 6 行、Markdown 436 / 2;crowded quota 增加 3049 字符 / 94 行;crowded turn-plan 增加 237 字符。它们超过现有各行的增长 allowance。这证明只把 3200 改大仍不足以完成资格;需为完整受影响集合逐项核对语义价值和已有 differential 契约。
最小修复:按既有 budget decision 规则评估这 436 字符的读回价值,补齐该 mode variant 和全部受影响差异行的有依据预算及相关文档/断言,或做保留完整命令、来源、重新检查和权限边界的等义精简;不要为绿色检查删掉必要义务,也不需要修改 vision 硬预算或实验阈值。复跑上述 smoke 与完整 variant 测量,并保留 base/head 依据。其它扩大后的固定输出预算不能覆盖这处遗漏。
动机
普通 CLI 调用者需要在首次保存成果时找到填写方式,在规划后读回完整任务,还需要在没有 runnable Todo 的周期性重规划中知道该提交什么。旧默认包给出部分提示,却隐藏完整 replan authoring 到另一次 detail 查询;任务摘要也无法证明最后一条要求没有丢失。本次使这些现有操作在当前响应可发现。
独立配对真实 CLI 已覆盖 legacy Markdown、File、SQLite:长 Todo 的末尾要求完整读回,普通首次保存和持久结果校验成功,in-flight 不带 vision 也能保持任务 open 并在下一 Turn 接续。另通过原生 Todo complete构造无 Todo 重规划,三种后端的默认 authoring 均由缺失变为完整,实际填写、保存、拒绝坏输入和一次记账可执行。这是已观察到的发现性/操作可达性收益;它不等于模型首次成功率或长程总成本改善。
改动思路
语义仍归现有 TypeScript visionAuthoringContract、settlement 和 replan owner。Python 承担既有 CLI 路由与投影适配,新增 _reference_inline_replan_authoring 仅删除同一响应内完全相同的 schema 副本:目标必须是 durable_writeback,整对象相等才产生索引引用。没有副本、schema 不同或步骤类别错误都保留完整内联输入;progress-only 不凭空增加 vision。显式 detail、TurnEnvelope 与 MCP 保持自己的完整契约。
本 PR 的边界是当前响应的首次填写、完整任务读回和默认重规划输入发现;不新增验收、权限或记账规则。
判定依据是修改前固定 revision 0012c4eb52191ba58d510e4af8cabb46704c1539 的 goal-vision-replan-contract-v0,按真实既有章节核验:Ownership Boundary 保留单一规则 owner;CLI Budget 仍拒绝超限且不截断;Vision Checkpoint 区分 material closeout 与 in-flight;Read basis for checkpoint-only recovery 要求原 Turn 已提交写回;Write / Correction Mechanism 保留当前 Agent、原结算身份与证据要求。head 文档是默认指导变化的披露,不被用作自己的验收依据。
相邻重构检查确认当前 owner 已复用;同包等值引用是有调用场景的窄投影去重,不增加第二套状态、权限或配置。额外 capability、恢复预检或平行 Python 决策源会扩大成本,没有必要。
具体改动
18 文件,388 增 / 51 删:生产/测试预算模块 66 增 / 12 删,测试 281 / 24,文档 40 / 14,生成 census 坐标 1 / 1。
turnScopedCliSettlementPlan(quota/settlement_plan.ts:11)把现有 authoring 投影到 Todo 的 durable_writeback;validation → writeback → spend → 条件 terminal 顺序不变。visionAuthoringContract(goals/vision_checkpoint.ts:76)补充 ordinary CLI 填写参数与原提交/同 Turn 恢复时序,未改变 validator 或硬字段限制。todo_authoring_steps(goals/start_goal_todo_delta.py:132)在写入/调整后提供 routed exact-ID 完整详情;要求唯一匹配、source_complete、全文与 status/claim 对比。缺失/变化需重新检查;命令不授予 guard/lease。_reference_inline_replan_authoring(quota/cli_projection.py:465)在原默认 CLI 投影中保留无 Todo 的完整 recipe,或引用同包相同 schema。整对象相等、步骤类型和索引共同防止引用不完整或不同的输入;显式 detail 不被修改。- Python/TS SettlementStep 及桥接传递可选 authoring。没有新的持久 state/receipt vocabulary;local ordered-step ID 和诊断引用沿用原投影边界。两份协议披露默认变化,census 只更新源代码位置。
独立验证:263 项 Python 结算、重规划、投影、TurnEnvelope/MCP、规划与预算测试通过。71 项 TS settlement/vision/effect/MCP/TurnEnvelope 测试、TS typecheck、Ruff、diff hygiene、diff advisory 后的完整 semantic smoke 通过。原生 canary 执行 19 项,18 通过、1 项上述预算失败,另外 5 项 direct 检查通过;18 个候选公共文件边界扫描干净。未查询或等待 CI。
同一输入在两个冻结 worktree 运行实际 CLI:三后端×普通/in-flight 六组,另三后端无 Todo 周期重规划。提前 checkpoint-context、提前 spend、错误 Agent、缺失当前证据及 421 字符 acceptance 均按对应场景拒绝;合法输入持久化、原 identity 和原 Turn 保留,重复 spend 不产生新扣额。无 Todo 场景前一 Turn 和当前 Turn 各一次扣额。精确 Todo 读回在加入 40 个无关 done 项及重排后不变,完成后仍可读回原正文。缺失 ID 是明确 not_found,不能用摘要缺席推断源不存在。
初始验证使用了几个不存在的测试路径和不支持的 canary 参数,以及额外 Todo-replan 探针对 receipt shape 的错误假设,属于 reviewer 调用错误;更正后的实际入口另行保存并执行。此次当前产品检查失败已通过冻结 base/head 实测归因;作者报告的旧 5/6 cadence 夹具不作为本 PR 新回归,也没有把它修复视为当前所有检查通过。
对主干的风险
默认指导变化真实存在,不是关闭可选 TurnEnvelope 后不受影响的功能。普通 Todo 包新增填写提示;默认 replan 不再强制另查 detail 才取得输入;MCP 继续使用自身 complete_task,没有手工结算第二条通道。现有强制 evidence/checkpoint/receipt 不能被提示文字降为建议;in-flight 加入可用示例也不新增填写义务。六组实际结算 precondition 文本在 base/head 完全相同。
效率代价已配对测量:小场景 19196→19196 字符,多 Agent 21572→21572;拥挤重规划 35383→38432(+3049,+8.6%),带完整 vision 的拥挤场景 41176→44225(+7.4%),各增加 94 行。原验收摘要、required fields、允许 outcome 与 rule 均相同,新增的是完整可填写 recipe。Todo 结算载荷增加 2112 UTF-8 字节,规划增加 820。guided 详细 Markdown 的新增 436 字符有操作价值,但其 mode-specific budget 尚未同步,它和上述 10 行差异资格共同构成本次阻断点。
没有新增配置或前端用户操作;测试边界是实际 agent-facing CLI 与现有 MCP/TurnEnvelope 投影,未声称重验整个 packaged UI、安装升级或 PostgreSQL。未运行模型或热改冻结实验;模型采用、查契约/返工次数、总 token/延迟、长期多领域效果均未测。输出更大是已知成本,减少查找步骤是已交付入口,净效率收益仍需实际轨迹验证。
我的整体评价
REQUEST_CHANGES,阻断点为由本 PR 引入且已复现的 absolute/differential 预算资格遗漏。发现性设计与既有 authority 边界的方向合理:首次写回、完整 Todo 与无 Todo 重规划都能在真实后端操作,原身份、跨 Turn 和一次扣额保持。修复该预算并在新 exact head 重核后可重新判定;旧批准、局部测试和另一个夹具修复不能替代这一步。
控制面变更仍由 maintainer 合并。本次未合并、升级、扩大权限、修改 frozen 实验或向其它聊天发消息。
English verdict: REQUEST_CHANGES — exact head 7dd2ec6.
P2: qualify the expanded guided Markdown detail mode. On identical synthetic input and path, immutable base 0012c4e emits 2781 characters/35 lines; this head emits 3217/37 against the unchanged 3200-character limit. The added exact-Todo readback step accounts for the 436-character delta. The native premerge budget smoke fails; an independent rerun passes on base and fails on head. The unchanged full differential comparator also rejects ten affected growth rows: guided JSON +948 and Markdown +436 across scales/detail, crowded quota +3049 characters/+94 lines, and crowded turn-plan +237. These were independently measured without modifying product budgets. Qualify the entire affected set and this specific variant with justified budgets/assertions, or shorten it without losing command, source completeness, retry and authority clauses. Re-run the full budget smoke; do not weaken vision limits or frozen experiment thresholds.
The bounded discoverability improvement is otherwise supported by paired real CLI journeys on legacy/File/SQLite, both ordinary and in-flight settlement, and native-completed no-Todo periodic replanning. Existing typed owners remain authoritative; equivalent schemas are referenced only within the same response and durable-writeback step, while absent/different schemas remain inline. Invalid evidence/identity/oversize and premature recovery/spend are rejected; persisted identity, replay idempotency and next-Turn continuation remain intact. Seventy-one TS tests, typecheck, lint, semantic inventory and public-boundary checks passed; local canary and the full unchanged growth comparator remain red for introduced output qualification gaps. CI was not consulted.
Small and multi-agent packets are unchanged; crowded replanning adds 3049 characters (8.6%, or 7.4% with a rich vision). This improves access to usable input but does not establish live model adoption, token savings, latency or long-horizon effectiveness. The revised default is disclosed and does not grant authority. Maintainer merge is still required after a repaired exact-head review.
Signed-off-by: LoopX Agent <337587101+loopx-agent@users.noreply.github.com>
loopx-agent
left a comment
There was a problem hiding this comment.
Approval conclusion (author-owned PR; GitHub blocks formal self-approval)
Reviewer: model_agent | gpt-6.1-sol | OpenAI | runtime_reported | reasoning_effort=xhigh
评审准确 head:c82d429ea20c66334dfc5cfe8bc68db0e0dd5e91;冻结 base:0012c4eb52191ba58d510e4af8cabb46704c1539。
动机
普通 CLI Agent 在首次保存成果或定期重新判断工作方向时,需要填写现有的验收与证据字段。
旧普通结算包没有填写示例,默认重规划又把必需输入移到另一次详情读取;规划清单只有摘要,调用者还需寻找完整正文命令。
当前包直接提供既有填写契约,并在规划后给出精确 Todo 读回路径;独立实测确认调用者能提交有证据的输入、保持原身份结算及读取完整要求。
本 PR 交付输入可发现性,不声明模型采用、成本下降、整个 Goal 完成或安装升级。
本轮重新判断完整22文件差异,而非沿用旧批准。上一核验 head 7dd2ec602d3c06060061a8352423b944af2c551f 到当前 head 只改变五个验证/预算文件;其余17文件逐一比较 Git blob 相同,包括生产 owner、调用者、协议和原生端到端测试。旧批准本身不继承;相应独立证据保留准确来源和失效检查。新增包体豁免、完整真实 CLI 配对和限制反例在当前 head 独立运行。
改动思路
沿用现有 TypeScript Vision owner 和 Python CLI 投影适配器;填写输入应随执行步骤可得,真正重复内容可在同一响应内引用,诊断审计仍按需展开。
当前 PR 边界是普通/继续工作结算、默认重规划及共享规划读回的输入发现,不新增决策源、持久状态、配置开关或权限。
最强的不交付理由是默认包变大,且尚未证明模型理解或实际节省成本。只写外部文档仍不能修复当前包的发现路径;强行压回旧预算会继续隐藏所需输入。选择既有 owner 的完整契约加安全同响应引用,比另建 capability 或复制 validator 更小。预算增加只保护这项明确默认变化,不改变 Vision 硬限制或 quota。
相邻重构检查已应用单一 owner 复用与真正重复 schema 的消除。旧冷读 helper 被替换,当前调用点和文档一起更新;保留不同 schema 的内容,无需另加通用同步框架。
具体改动
完整变更22文件:生产7文件59增/7删;验证工具4文件98增/6删;测试8文件359增/24删;文档2文件40增/14删;生成 census 坐标1增/1删。新增验证没有变成运行时决策 owner。
独立规格为冻结 base 的 goal_vision_replan_contract_v0 与 Quota CLI Hot-Path Compaction v0。按修改前要求判断:
- Ownership Boundary:TS 继续拥有 Vision/结算判断;Python 只投影、路由和引用同一契约。
- CLI Budget:420 字符 acceptance 与 1800 总字符等原硬限制保留;越界拒绝,不截断。
- Vision Checkpoint:普通 material closeout、合法 in-flight、checkpoint satisfied 与 Goal acceptance 保持区别;继续工作不被新增 schema 强迫修改 Vision。
- Read basis for checkpoint-only recovery:checkpoint-context 仍只服务原已提交 Turn,提前恢复/扣额仍拒绝。
- Qualification Contract:实际 CLI 配对与投影验证必需填写输入及执行 authority;包大小独立测量。原文的 schema 冷读被本 PR 明确改变为内联/同响应引用,不能宣称这段文字等价或模型资格已完成。
关键代码讲解
turnScopedCliSettlementPlan(settlement_plan.ts:11)在 Todo 绑定的 durable_writeback 带上visionAuthoringContract();普通与 in-flight 共用输入,原 validation→writeback→spend 顺序、原身份与条件性 terminal 保持。visionAuthoringContract(vision_checkpoint.ts:76)复用 validator 限制与示例,明确--agent-vision-json和原 Turn 恢复顺序。TS/PythonSettlementStep序列化桥仅传递可选字段;无字段的旧载荷仍合法。_reference_inline_replan_authoring(cli_projection.py:465)只在已有 durable_writeback 的 schema 逐值相等时使用精确同响应路径。找不到步骤、不同 schema 或错误步骤保留完整输入;progress-only 重规划不获得新的 Vision 要求。明确 Vision detail 不做该替换,TurnEnvelope 保留完整输入。todo_authoring_steps(start_goal_todo_delta.py:132)在新建/复用步骤后提供带 registry/runtime/Goal/Agent 的精确读回命令。读完整.todo.text、source completeness、status/claim;缺失或歧义要求重查,摘要不是正文来源。读取不授予 guard 或 lease。- 预算和两份协议文档披露默认改变;测试覆盖两种结算边界、三种实际本地存储、重规划与引用反例。I/O manifest 只更新源代码坐标。
当前独立运行差异单测 126 passed,完整真实 CLI 配对 base=96、candidate=96、无新增独立行、10组有审查信号;当前 Ruff、diff check、开发期 advisory、完整 semantic smoke、290-site manifest 通过。一次误写 semantic 脚本路径在执行前失败,按仓库实际路径纠正后通过,不计作产品回归。没有查询或等待 CI。
上一轮在 7dd2ec602d3c06060061a8352423b944af2c551f 的独立240项 Python、59项 TS/typecheck、三版本首次写回/缺证据拒绝/纠正/第二 Turn 单次扣额及长正文尾部/claim/40项无关任务读回证据,是本轮复用的历史来源,未冒称当前重跑。生产、原测试及调用者的17个 blob 不变;base 与相关 main caller 无变化。当前新增验证路径另行核验,原包体预算也由当前完整配对重新执行。
新 _authoring_input_migration(cli_output_differential.py:724)只对明确 false→true 的完整指引观察、指定 guided/quota/Turn-plan 表面提供有上限的一次增量;相同/无观察、错误表面、任一指标超额仍失败。authoring_input_observations 匹配结构位置、完整提示与命令,而非在任意 note 中搜索关键词。它是隔离验证观察,不是运行时状态或语义等价证明;仍需人工判断整段义务。去掉旧冷读 metadata 的 shape 信号已按完整内联 schema 与既有诊断 audit 的区别核对。
当前同 fixture、同物理路径的配对:guided JSON +948字/6行,Markdown +436字/2行;required-replan quota 35,383→38,432字、+94行;Turn-plan +237字/0行;explicit guided detail Markdown 2,781→3,217字。其他86行不变。新增3,600 detail ceiling 保留383字余量,default Markdown、其他 lanes、per-Todo增长、Vision输入硬限制保持。拒绝同观察和每指标溢出的126项测试,以及完整96组实际输出,共同验证限额不是无限豁免。
对主干的风险
这是默认 CLI 引导改变,无 TurnEnvelope 开关也会得到新内容,不属于 default-off 功能。MCP 继续移除手工 CLI 结算过程并使用自己的同源 authoring;前端/Lark 没有新增用户操作或配置。本次不宣称重新验证打包前端、PostgreSQL、真实模型或线上安装。
上一轮独立相同路径/相同输入测量(生产 blob 不变,本轮实际预算另由96组配对核验):小场景 19,720 字符、520 行不变;多 Agent 21,900 字符、604 行不变;拥挤 required-replan 37,637→40,686 字符、828→922 行,即固定 +3,049 字符/+94 行。这里路径较长,绝对值不能冒称作者短路径 fixture 的预算结果;增长及保留决策是在相同路径配对测得。新 crowded 39k/950 行、固定增长 9k 和 rich Vision 45k 是回归 fixture 预算,其他 lanes/per-Todo 限制保留;1800 Vision 输入等硬限制未改。必要语义换取包体开销合理,但真实 token、延迟和收益未测。
判断、准入、replay、写回 authority 与 once-only debit 未新增规则。同响应引用依赖真正相同且保留的步骤,不能拿全局 locator 或后来读取替代当前输入。独立相关 caller 的 main 源码与冻结 base 同 blob,近期 manager/release-doc 合并未使该判断失效。当前 source 全树 semantic 通过仍有已登记动态 producer 覆盖限制;advisory 0 不证明动态语义安全。
我的整体评价
APPROVE 当前有界发现性改进。沿用现有 TypeScript Vision owner 和 Python CLI 投影适配器;填写输入应随执行步骤可得,真正重复内容可在同一响应内引用,诊断审计仍按需展开。当前 PR 边界是普通/继续工作结算、默认重规划及共享规划读回的输入发现,不新增决策源、持久状态、配置开关或权限。
持续工作方面,原验收、后续 Turn、拒绝后的纠正和一次记账保持;操作体验方面,首次填写与完整读回无需另寻命令。尚未证明模型采用、长期成本或产品整体验收,不能把输入可得等同于模型已用。保持维护者合并边界,合并仍由维护者处理。
English verdict: APPROVE - c82d429. Reused authoring inputs are available on ordinary and default replan paths, with safe same-response deduplication and complete Todo readback discovery. Independent old/base/head CLI probes preserve rejection, recovery and one debit per Turn; required replan adds 3,049 characters/94 lines. Current126 differential tests and full96-row immutable base/head CLI pairing pass. Earlier240 Python/59 TS/typecheck and real recovery/readback evidence is explicitly reused from7dd2 with17 unchanged blobs and caller/base invalidation checks; it was not rerun here. Live adoption and installed behavior remain unverified. Maintainer merge required.
loopx-agent
left a comment
There was a problem hiding this comment.
Reviewer: model_agent | gpt-6.1-sol | OpenAI | runtime_reported | reasoning_effort=xhigh
Request changes conclusion (author-owned PR; GitHub blocks formal self-review)
REQUEST_CHANGES — exact head c82d429ea20c66334dfc5cfe8bc68db0e0dd5e91。
[P2] 完成新增比较规则的 maintainability ratchet 处置。 此 head 已修复上一轮绝对与差异输出预算,但新增 _authoring_input_migration 调用/信号收集(loopx/control_plane/testing/cli_output_differential.py:800–803)使 _compare_row 从 89 到 91 statements,越过原 90 门槛。完整 native premerge 报告新增 oversized_decision_function:loopx.control_plane.testing.cli_output_differential:_compare_row 为 unreviewed;独立冻结 base 0012c4e 与当前 head 运行同一 control-plane-maintainability-ratchet-smoke.py:base 通过,head 失败,当前 decision_points=44、阈值60。它是本次新回归,不是之前已有的两项兼容 facade 例外。
最小修复是在同一 diff 按该质量规则完成有依据的处置:优先考虑同因、行为保持的小整理;若额外抽象反而不值得,则按现有规则记录有边界的 reviewed exception、metric ceilings 和退役依据。不要仅改变统计方式、拼表达式或删必要说明来凑数,也无需扩大成无关重构。保留一次性观察条件、入口限制、各指标 overflow 阴性覆盖,并复跑上述 ratchet 与完整预算 smoke。
动机
普通 CLI 调用者需要在首次保存成果时找到填写方式,在规划后读回完整任务,还需要在没有 runnable Todo 的周期性重规划中知道该提交什么。
旧默认包给出部分提示,却隐藏完整 replan authoring 到另一次 detail 查询;任务摘要也无法证明最后一条要求没有丢失。
本次使这些现有操作在当前响应可发现。
它不等于模型首次成功率或长程总成本改善。
上一轮真实 base/head 操作已证明首次填写、精确完整读回及无 Todo 重规划可达;本次修复只变五个验证文件,17 个关键 runtime/原测试文件逐字节相同,故复用那些三后端、失败纠正、持久读回、单次记账和跨 Turn 证据。当前 head 另外独立运行完整真实 CLI 资格与增长观察,避免把旧 head 的失败或旧测试记录当作当前通过。
改动思路
语义仍归现有 TypeScript visionAuthoringContract、settlement 和 replan owner。
本 PR 的边界是当前响应的首次填写、完整任务读回和默认重规划输入发现;不新增验收、权限或记账规则。
CLI qualification 的既有 Python owner 现在从实际 stdout 观察完整 readback instruction 或带指定 schema 的完整 authoring_hint。它们是本地验证观察,不是完整 schema 正确性、当前事实、Goal 验收或运行时权限的证明。原真实操作/validator 证据仍需独立核对。JSON 内普通 note 字符串不被当成步骤;文本只认完整 rendered readback 形态。快照文字在验证边界有理由保持独立于生产生成器,未来文字变化会停止对应许可并要求重新审查,不把 substring 当状态分类。
比较器仅在指定 guided/quota/turn-plan 入口的明确 false→true 观察变化时提供有界、一次性增长 allowance,并保留 review signal;already-upgraded、缺少观察、其它入口均不获得该许可。其余绝对上限、per-Todo/规模增长、schema/action 语义检查保留。guided detail Markdown 单独从3200改3600,对实际3217保留383字符余量。没有放宽 vision 硬限制或冻结实验阈值。
独立规范依据仍为固定 0012c4eb52191ba58d510e4af8cabb46704c1539 的 goal-vision-replan-contract-v0:Ownership Boundary、CLI Budget、Vision Checkpoint、Read basis for checkpoint-only recovery、Write / Correction Mechanism。Head 文档/作者描述作为披露,不作为自己的验收依据。
具体改动
完整 PR 为22文件,557增/52删:生产及测试预算 owner156/13,测试/公共probe360/24,文档40/14,census1/1。此次相对7dd的五文件增量包括:
关键代码讲解
authoring_input_observations(testing/cli_output_semantics.py:12)观察完整固定指令,JSON检查既有 step id/kind/purpose 或 schema/hint,Markdown检查完整 step/命令形态;不写 Goal 状态。_receipt_row(examples/control_plane/cli-output-probe-runner.py)把真实输出观察保存在 transient qualification receipt。它是派生元数据;旧 receipt 缺少观察时不获得新许可,不新增 runtime authority。_authoring_input_migration(testing/cli_output_differential.py:724)限定新语义观察、入口及每个指标,_compare_row保留其它 policy/format/schema/growth 检查及人工审查 signal。此调用也暴露上述 function ratchet 缺口。- guided detail mode 的绝对字符预算单独调整,126 comparator及3 real runner测试补充缺失/已有观察、其它入口、每个指标越界和完整/部分指导的负例。
原协议五项条件在当前 head 的判定:Ownership Boundary、CLI Budget、Vision Checkpoint、Read basis for checkpoint-only recovery、Write / Correction Mechanism 均 implemented;其实际 runtime 证据按17项 byte-hash 核对复用,当前 changed qualification 则独立重跑。新的 function 质量门槛仍 not_met,不能被这五项替代。
本次独立129项新 Python测试通过;完整 real CLI base/head 96行对比 0失败、10 review signals,native canary 中完整 cli-output-budget-regression-smoke.py 也通过。上一轮3200绝对失败及10差异失败已在本 head 验证解决;它们的原公开评审与失败记录保留,不删除历史。
Native premerge 初次19项中17通过,另有上述 ratchet 和新 worktree 缺 npm dev dependency 两项失败。补齐 npm ci --ignore-scripts 后,full semantic smoke 独立通过;当前剩余 required failure 是 ratchet。5 direct检查和22公共文件边界扫描通过;Ruff、advisory→full semantic/diff通过。未查询或等待 CI。
原7dd的263 Python+71TS、三后端×ordinary/in-flight六组、原生 complete后无 Todo三组、40无关done/重排,经过当前17关键文件 byte-hash invalidation check 复用。它们证明正文/claim/原identity、evidence/预算拒绝、纠正、独立持久化、spend replay和下一 Turn 同项续接。没有把复用说成在c82全部重跑。额外 Todo-replan探针未完成,仍不用于证明当前生产 reference adoption;当前 equality/absent/different/wrong-step/detail/progress-only 分支由原有效 unit/transport 证据支持。
对主干的风险
本次 qualification 许可只接受已观察、首次引入的指定指导,有限余量不是无限增长。实际96行中10条都有显式 signal,我已人工核对其新增内容、原指令义务和定量成本;它不是自动批准或 Goal 完成信号。当前 required ratchet failure 必须据实保留,绿色129测试和预算 smoke不能替代全部质量资格。
效益/成本结论未因预算变绿而升级:小quota19196→19196,多Agent21572→21572;crowded35383→38432(+8.6%),rich41176→44225(+7.4%);Todo结算+2112B,规划+820B。完整输入减少额外查契约的机会,真实操作可达、原验收和跨轮语义保留。模型是否少返工、总token/延迟是否下降、长程多领域实际效果仍未测。
语义与 CI 对齐
没有新配置或前端用户操作;ordinary CLI的默认指导变化已披露,不能伪称可选功能关闭后的原样 parity。MCP仍使用自身 complete_task;in-flight不新增vision义务。未重验整个 packaged UI、升级安装、PostgreSQL或真实模型,也未热改冻结实验。可撤回派生指导/验证迁移,不改变既有receipt或状态格式。未来整理检查发现的是具体 _compare_row 质量门槛需要处置,不要求额外框架。
我的整体评价
REQUEST_CHANGES:上一轮 absolute/differential 预算问题已解决,当前剩一项新 introduced maintainability qualification。设计仍有正向的发现性与体验价值,代价是固定包体增加;实际总效率保持未测。完成上述小而有依据的同 diff 处置后,在新 exact head 重核即可,不把一项statement门槛扩大成无关重构。
控制面变更仍由 maintainer 合并;本次未合并、升级、改变账户/权限、运行模型或发送其它聊天消息。
English verdict: REQUEST_CHANGES — c82d429.
P2: settle the introduced maintainability finding. The new migration call/signal collection increases _compare_row from89 to91 statements against its existing90 threshold. Native premerge and an independent immutable-base/head ratchet run pass on base and fail on this head with one new unreviewed oversized decision function (44 decision points,60 limit). Apply a small behavior-preserving same-domain cleanup, or an evidence-backed bounded reviewed disposition under the existing quality rules; do not manipulate counting or remove useful instructions. Preserve scope/one-time/overflow negatives and rerun ratchet plus the full budget smoke.
The previous absolute and differential output-budget findings are resolved:129 current tests pass; paired real CLI qualification compares96 rows with zero failures and10 explicit review signals. Guidance growth receives a bounded first-transition allowance only on named surfaces, not runtime authority. The initial missing npm dependency was repaired and full semantic validation passes; the ratchet remains the current required failure. Prior three-store first-write/readback/recovery/replay/continuation evidence is reused only after17 critical files were verified byte-identical. CI was not consulted.
This retains the useful input-discovery improvement and existing typed state owner, original evidence/identity/one-debit boundaries, and progress-only/in-flight semantics. Packet growth remains an observed cost; live model adoption, net token/latency and long-horizon effectiveness remain unmeasured. Maintainer merge is required after repaired exact-head qualification.
Signed-off-by: LoopX Agent <337587101+loopx-agent@users.noreply.github.com>
loopx-agent
left a comment
There was a problem hiding this comment.
Reviewer: model_agent | gpt-6.1-sol | OpenAI | runtime_reported | reasoning_effort=xhigh
Approval conclusion (author-owned PR; GitHub blocks formal self-approval)
APPROVE — exact head b0cfba64b3929093e0f1dbc240e35a664a88022a。上一轮两组输出预算问题及新增维护性检查问题已分别验证解决;旧批准没有继承到此 head。
动机
普通 CLI 调用者需要在首次保存成果时找到填写方式,在规划后读回完整任务,还需要在没有 runnable Todo 的周期性重规划中知道该提交什么。
旧默认包给出部分提示,却隐藏完整 replan authoring 到另一次 detail 查询;任务摘要也无法证明最后一条要求没有丢失。
本次使这些现有操作在当前响应可发现。
它不等于模型首次成功率或长程总成本改善。
完整 PR 的操作价值已由此前独立真实 CLI 证明:三种本地存储的首次填写、完整末尾要求、失败后纠正、持久读回、单次扣额及下一 Turn 续接保持。当前增量只整理一个验证模块;通过源代码失效检查复用原证据,并对修改的比较器重新执行测试、输入对比和原生完整资格。绿色检查不代替实际收益判断。
改动思路
语义仍归现有 TypeScript visionAuthoringContract、settlement 和 replan owner。
本 PR 的边界是当前响应的首次填写、完整任务读回和默认重规划输入发现;不新增验收、权限或记账规则。
最强的不交付理由是默认包有固定开销,尚未证明模型减少查找或返工。仅补外部文档仍留下当前响应的发现缺口;新增能力或另一套恢复规则会重复现有 authority。复用已有 authoring 和原生全文命令,再仅引用同一响应内真正相同的 schema,是有真实调用点的窄方案。没有必要为省字符删除填写、完整来源或原 Turn 恢复义务。
当前同因整理把逐指标的成本比较放入同一 Python qualification owner 的 _compare_row_growth,外层继续负责契约、语义覆盖和结果组成。它接收已计算的 migration/route allowance,不重新建立规则 owner。保留原 metric 顺序、所有限制、一次性许可条件及失败/信号顺序;新增 helper 已由现有生产验证入口调用。没有提高90 statements门槛或添加例外,也没有新模块、配置、持久状态或运行时权限。
具体改动
全 PR 为22文件,609增/83删:生产及验证 owner208/44,测试及公共probe360/24,文档40/14,census1/1。相对上一轮c82仅一个文件52增/31删;运行时路径、观察器、限额和测试内容均未改。
独立规格使用修改前冻结 base 0012c4eb52191ba58d510e4af8cabb46704c1539 的 Goal Vision Replan contract 和 Quota CLI Hot-Path Compaction。逐项判定:Ownership Boundary、CLI Budget、Vision Checkpoint、Read basis for checkpoint-only recovery、Write / Correction Mechanism 均 implemented。原字段硬限制仍拒绝、不截断;原提交后才能进行 checkpoint recovery;in-flight保留原任务及结算身份。Compaction 的旧默认冷读机制被明确改变为完整输入/同响应引用,因此是已披露的默认变化,不能称这部分文字完全等价。模型资格未由本轮确定性验证补齐。
关键代码讲解
turnScopedCliSettlementPlan(quota/settlement_plan.ts:11)在 Todo-bound durable writeback 暴露现有visionAuthoringContract;Python/TS serializer 只保留可选字段。validation→writeback→spend→条件terminal和原identity不变。MCP仍走自己的complete_task。todo_authoring_steps(goals/start_goal_todo_delta.py:132)在authoring/delta后给出带原路由的exact-ID读回:唯一匹配、source_complete、全文、status/claim需核对;缺失/变更需重新检查,命令不授予guard/lease。_reference_inline_replan_authoring(quota/cli_projection.py:465)只引用同响应durable_writeback内逐值相同的完整schema。没有目标、schema不同或step错误时保留完整输入;detail/TurnEnvelope完整契约和progress-only边界保持。authoring_input_observations与_authoring_input_migration由实际stdout观察完整固定指令,并只给指定guided/quota/turn-plan的明确false→true一次性增量及review signal。普通note、部分提示、无观察、already-upgraded、错误表面和任一指标越界不能得到豁免。观察不是完整schema正确性、当前事实或权限证明。_compare_row_growth(testing/cli_output_differential.py:760)保留原projection/fence/authoring允许量及各metric比较。外层先收policy/format/contract错误,随后接回growth失败与signals,再检查语义/anchor/signature。原growth AST与新提取块完全相同;组合反例独立断言错误及信号的具体先后顺序。
本次独立129测试通过。对旧c82/new head比较:测试中的311次调用,含250拒绝结果;另96行既有真实CLI配对输入和1组同时触发契约、成本、语义错误的反例,共408次结果逐字典相等且输入不变。这里旧实现是重构characterization对象,合法/非法预期还来自原有测试及独立顺序断言,不以“新旧相同”冒充产品正确性的全部依据。
当前独立 native premerge 19/19通过,5 direct检查通过,22公共文件扫描干净,无失败/跳过/hold;包括维护性ratchet、advisory之后的full semantic、固定0012c4e base的完整真实CLI absolute/differential budget、quota/replan/scheduler/Todo相关smokes。Ruff与diff检查也通过。此前c82的91>90失败保留为历史,这次不挑选局部绿色结果替代全量资格。
证据复用来源明确:7dd的263 Python+71TS/typecheck、三后端ordinary/in-flight六组、原生complete后无Todo重规划三组和40无关done/重排;c82的完整实际base/head96行/0失败/10review signals与输入指导条款对比。此次28个runtime、原测试、观察/预算、文档及依赖路径byte-hash一致;当前main ef999b6866a64b15ecee3f9ecae8eedc2749b729 的10个相关owner/依赖路径仍与冻结base相同。复用历史记录保留其原revision,未称全部在新head重跑。之前未完成的额外Todo+replan探针仍不计作实际reference adoption。
对主干的风险
本次helper是同模块、同原因的行为保持整理:外层91→66 statements,新helper35;原90门槛保持。原两项已登记兼容facade不被删除或伪装通过;新unreviewed function finding已消失。没有扩大到TS迁移或新框架。故障、恢复与一次扣额仍由原runtime owner决定,validation observation不能替代任何业务事实。
输出代价没有被维护性修复抹掉:固定配对的小quota19196与multiAgent21572字符不变;crowded35383→38432(+8.6%),rich41176→44225(+7.4%),结算+2112 UTF-8B、规划+820B。guided detail2781→3217现在在有依据的3600预算内;一次性差异允许量仍有明确signal、有限margin和阴性覆盖,不是持续自由增长。条款检查保留排序、来源、模态、scope、continuation和stop义务。
语义与 CI 对齐
普通CLI指导默认变化已披露,不伪称可选开关关闭后的原样parity。没有新配置或前端/Lark操作;既有MCP/TurnEnvelope与in-flight不新增vision义务。复用现有typed vocabulary;新增步骤/引用为派生投影,qualification observation留在测试边界,无新authority lifecycle。advisory是有限语法提示,0项不证明动态语义安全。当前策略不查询或等待CI。
未声称重验全部packaged UI、升级安装、PostgreSQL或真实模型;此增量不改authority store。不改变冻结实验、评分、runner语义或账号/权限。模型采用、返工、总token/延迟及多领域长程效果仍未测。回滚投影/验证迁移不要求重写持久receipt或Todo。
我的整体评价
APPROVE 当前head的有界输入发现性改进和同因质量整理。长程正确性判定preserved:原要求、证据、身份、原Turn、一笔扣额与后续工作保留;用户体验判定improved:当前响应能发现首次填写与完整任务读回。固定包体开销已知,净模型效率仍unverified,不把测试/合并数当效果证据。
c82的维护性阻断已由同owner的小提取及新head证据解决,7dd的absolute/differential问题也已资格化。旧公开失败保留为历史,不删除讨论;此结论只覆盖新exact head。后续是否节省token或提高首次成功率需真实轨迹。控制面仍由maintainer合并,本次不合并或升级。
English verdict: APPROVE - b0cfba6. The introduced maintainability failure is resolved by a same-module, behavior-preserving growth seam:66/35 statements under the unchanged90 limit. Independent129 tests and408 old/new comparisons (including250 rejecting test cases,96 actual receipt inputs and a mixed diagnostic-order oracle) retain complete results and input nonmutation. Current native qualification is recorded above; previous core/store evidence is reused with explicit revision and28-path invalidation checks, never inherited approval. First-write/full-readback discoverability improves; fixed payload cost remains and live model/net long-horizon efficiency is unmeasured. Maintainer merge required.
|
Additional qualification on unchanged head The initial rehearsal incorrectly selected a Todo before discovering the new replan frontier. That was correctly rejected with |
Ordinary CLI settlement described material Vision closeout but omitted its authoring example. Default replan could also require an extra diagnostic read for that input, while shared planning exposed bounded Todo excerpts without a complete-body readback command. Callers now receive the inputs needed to perform and verify the existing operation.
.todo.text, source completeness and status/claim before handoff.Validation covers real isolated Legacy/File/SQLite CLI paths: full-body tails and claims, ordinary/in-flight first writeback, Todo-bound and obligation-only periodic replan authored from the returned example, premature recovery/spend rejection, missing evidence/oversize rejection, and one debit on replay. Projection tests preserve unequal schemas and progress-only contracts. Focused CLI budgets, settlement, TurnEnvelope and MCP tests, TypeScript tests/typecheck, Ruff, semantic inventory/full drift check and public boundary scans were run. Benchmark scoring and runner behavior are unchanged; model adoption, latency and score effects remain unmeasured.
Matched output tradeoff: standalone settlement plan 2,891 → 5,003 characters; guided JSON 30,527 → 31,453; issue Markdown 3,309 → 3,757. For required replan, the same default CLI crowded fixture increases 35,383 → 38,432 characters and 828 → 922 lines; small/multi-agent fixtures are unchanged. The richer Vision fixture increases 41,176 → 44,225 characters (+7.4%). Preserve the useful obligations instead of shortening them: crowded quota allowance 36k → 39k characters / 850 → 950 lines, fixed semantic growth 6k → 9k; retain per-Todo growth and other lane limits. The rich Vision fixture ceiling is 45k. Existing planning allowance changes retain duplication assertions and the default Markdown ceiling. These are packet measurements, not demonstrated model savings.
Placement/future-facing pass: one existing TypeScript authoring owner; Python remains the CLI view/routing adapter. The same-response reference is local projection metadata, not a new semantic state or authority. No new capability/provider or configuration surface. Frontend/Lark operation semantics are unchanged. The registry I/O census change only updates a source coordinate. This delivers input discoverability under the existing Goal Vision contract; sustained model adoption remains separate acceptance.
The full real-CLI differential also exposed two missing qualifications, retained as failures before repair: explicit guided detail Markdown was 3,217 characters against 3,200 (base 2,781), then ten base/head growth rows exceeded the ordinary policy. The added readback step accounts for +948 JSON / +436 Markdown characters; the existing Turn authoring hint gains 237 characters. The explicit detail ceiling is now 3,600 (383 characters of headroom). Complete observed instruction additions receive a bounded one-time differential allowance on only guided, quota and Turn-plan surfaces; already-upgraded or unobserved bases and other surfaces retain ordinary limits. No Vision input hard limit, absolute default ceiling, fixture population or per-Todo growth constraint is relaxed by this follow-up. Full real-CLI base/head smoke and 126 differential tests pass, including absent/same-observation, wrong-surface and each-metric overflow cases. These markers are local validation observations, not runtime classification or authority.
The existing Python qualification owner now separates output-growth comparison from contract/semantic coverage in the same module. No limit, migration predicate, diagnostic ordering or receipt field changes. The original comparator is 66 statements and the growth step is 35, within the unchanged 90-statement ratchet. Before/after characterization across 311 comparisons (250 rejecting cases) is identical; 129 focused tests, the complete pinned-base real-CLI budget smoke, maintainability ratchet, Ruff and semantic advisory/full drift pass. No exception or higher complexity ceiling was added.
Merge remains with the maintainer. The prior Release Artifacts browser failure expected the former five-turn product default; main fixes that fixture in #6054. New-head review and CI remain separate from the completed source checks.