Skip to content

test(chat): align cadence and exact Todo fixtures - #5980

Merged
huangruiteng merged 1 commit into
mainfrom
codex/review-b-daily-chat-fixtures
Oct 9, 2026
Merged

huangruiteng merged 1 commit into
mainfrom
codex/review-b-daily-chat-fixtures

Conversation

@loopx-agent

Copy link
Copy Markdown
Collaborator

Two Chat regression fixtures still encoded older contracts: the absent cadence catalog expected completed Todos, and the manager exact-Todo mock returned a list. On the current main baseline both tests fail; this change uses the already documented five-effective-Turn default and returns the canonical single todo shape for exact reads.

The existing continuation, conditions, decision-scope and truncation assertions remain. The exact-read fixture now verifies the selected Goal/Todo and retains scope assertions; two negative cases distinguish a missing record from an unavailable authority without exposing its error. Production code, configuration and permissions are unchanged. The separate SSH evidence fixture repair is already proposed in #5961 and is reused rather than duplicated here.

Validation: both failing nodes reproduced on immutable base d33cc133e5948a30d3e5535e0339fd1ed34c2497; 26 tests passed across both changed modules and the real exact-Todo CLI suite. A separate isolated File/SQLite manager read exercised full owner continuation, conditions/scopes, non-owner privacy, missing and invalid selectors, with source bytes unchanged. Ruff, compile and diff checks passed; the native canary ran four direct checks with no selected risk checks or manual holds. These results do not claim the entire repository test suite is healthy.

Acceptance: existing docs/quota-allocation.md cadence default and loopx/control_plane/todos/list_readback.py exact-read contract; no new decision owner or vocabulary. Future-facing pass: keep the selector-aware double local to its consumer test; no new shared fixture abstraction is needed. The negative cases guard existing public read semantics.

Author: model_agent, OpenAI GPT-6.1 Sol. Independent review is still required; no merge is included.

Signed-off-by: LoopX Agent <337587101+loopx-agent@users.noreply.github.com>

@loopx-agent loopx-agent left a comment

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewer: model_agent; gpt-6.1-sol; OpenAI; runtime_reported; reasoning_effort=xhigh.

Approval conclusion (author-owned PR; GitHub blocks formal self-approval)
Exact head: d971dc6; base: d33cc13.

动机

维护者运行 Chat 回归测试时,会遇到已落后的周期单位和精确 Todo 返回形状,无法可靠区分测试夹具错误与产品回归。
配置页读取默认重规划周期、管理器读取长 Todo 时,旧测试误报失败;本 PR 让精确选择返回单个 Todo,并继续核对长文本尾部、恢复条件和私有信息边界。
已验证 26 项相关测试通过;基础版本两处旧夹具失败,而两版本的隔离 File/SQLite 真实读取均保留完整尾部和非 owner 隐私边界。
本 PR 不改变运行时、周期配置、Todo 存储、用户操作步骤、权限或源码中已存在的精确读取行为。

改动思路

复用已有 cadence owner 和精确 Todo projection,修正当前 Chat 消费者的测试输入。实际读取先给有界 overview,随后按 selector 读完整 detail;测试必须分别表达这两个调用。没有必要添加新运行时 abstraction。

具体改动

  • tests/test_chat_machine_configuration_api.py:默认单位改为 effective_turns,仍为 5。
  • tests/test_chat_manager_details.py:假读取器按 selector 返回 list 或单个 todo;保留长 continuation 尾部、恢复条件和 scope 断言,并添加 missing/unavailable 的错误与隐私区分。两文件共 +39/-2。
    规范映射:docs/quota-allocation.md;revision d33cc13。settled effective Turn cadence、single exact Todo payload、bounded durable regression seam 均 implemented;另对照 canonical context projection 与 docs/reference/protocols/decision-scope-v0.md。已有 canonical/CLI 测试覆盖存储侧,本 PR 补的是 Chat 消费者边界。完整队列及 #5961 对照未发现同形低价值批量拆分。

对主干的风险

Head 的三个相关测试文件 26 passed;同一基础版本命令 22 passed / 2 failed,两失败均为旧夹具。独立隔离 File/SQLite + 实际 Node projection 在 B/H 都通过完整尾部、conditions/scopes、非 owner 隐私、missing/unavailable 和无写入检查。两次猜错文件名未执行测试,以及第一次 synthetic fixture 缺少 canonical schema_version,均保留为评审设置失败;后续正确来源的执行替代它们。diff-check、两文件 Ruff 通过;语义 advisory 不代表全树等价证明。未运行全仓、安装后 UI、live Lark 或远端 CI;测试维护不改变这些入口。未来重构检查:selector-aware local seam 已足够,无需新共享框架。

我的整体评价

APPROVE,无阻塞发现。这是恢复已有消费者回归信号的可独立维护改动,未以作者声明或测试数量代替 native readback。GitHub 当前调用账号与作者同为 loopx-agent,正式自批准受平台限制;本正文记录精确版本的批准结论,平台 COMMENTED 不等同于 GitHub APPROVED。未授权本轮合并。

English verdict: APPROVE — d971dc6; stale Chat fixtures repaired, 26 focused tests pass and independent File/SQLite base/head readbacks preserve full-tail/privacy behavior; no blocking findings.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants