Skip to content

test(manager): align exact Todo evidence mock shape - #5961

Merged
huangruiteng merged 2 commits into
loopx-project:mainfrom
mikamikasuki:codex/loopx-manager-ssh-exact-todo-fixture
Oct 8, 2026
Merged

huangruiteng merged 2 commits into
loopx-project:mainfrom
mikamikasuki:codex/loopx-manager-ssh-exact-todo-fixture

Conversation

@mikamikasuki

Copy link
Copy Markdown
Contributor

Goal And Delivered Outcome

  • Outcome basis / optional anchor: Current-main test failure in the manager SSH evidence export, reproduced from the canonical exact-Todo read contract.
  • Goal/source and gap: Keep the SSH evidence export regression test aligned with the canonical Todo API. A list read returns todos; a todo_id read returns the selected record as todo. The test mock returned only todos for both calls.
  • Observable before → after, with the validation row that proves it: On base aa0c2ea, the exact-Todo case fails because the mock yields no row; after this change, list and exact-Todo cases both pass while the private note remains absent from exported evidence.
  • Issue/task and intended base: Self-contained test-maintenance outcome; base main at aa0c2ea.

Author Declaration

  • Written by: model_agent — OpenAI GPT-6 Luna

Implemented against

  • Specification and revision: Exact-read contract in loopx/control_plane/todos/list_readback.py and tests/control_plane/test_todo_exact_detail.py at aa0c2ea.
  • Criteria:
Criterion Disposition Symbol / path Test or command
List reads use the todos result; exact reads use the todo result implemented tests/test_manager_ssh_evidence.py test_export_uses_canonical_todos_and_never_returns_owner_private_continuation
Exact evidence export omits owner-private notes implemented Manager evidence projection Same test module; private-note assertion
  • Self-check before submission: Reproduced the exact-case failure on base, updated only the test double to match the canonical response shape, verified exact reads separately, and reviewed the one-file diff.

Scope And Continuation

  • Completed scope and remaining work: The stale mock now follows the existing list-versus-exact response contract; production code is unchanged.
  • Slice boundary / successor: Complete within this test-maintenance scope.

Validation

  • Tested revision: 765af5d (base aa0c2ea)
  • Run state: finished
  • Input classes: synthetic
Check kind Result Public-safe evidence / limitation
regression_parity passed Exact-Todo parameter fails on base because the mock omits the canonical todo field; updated test passes on this head.
unit passed tests/test_manager_ssh_evidence.py: 40 passed; tests/control_plane/test_todo_exact_detail.py: 4 passed.
static passed Ruff, Python compile, and git diff --check passed.
  • Coverage and gaps: This is test-fixture maintenance only; it validates list and exact-read export projections and the existing privacy assertion. No production behavior changed.

Frontend / Visual Evidence

  • UI impact: none
  • Before: N/A
  • After: N/A
  • States and viewports shown: N/A
  • Source data: synthetic
  • Attention review: N/A

Type of Change

  • Bug fix
  • New feature
  • Breaking change
  • Refactoring (no functional changes)
  • Documentation update
  • Test update

LoopX Area

  • Control plane (goals, todos, quota, scheduler, registry, runtime)
  • Benchmark boundary (adapters, runners, verifiers, scoring, evidence)
  • Capability or extension (providers, adapters, skills)
  • Public docs or presentation surface (README, protocols, dashboard)
  • Build, packaging, installer, or CI
  • Host or runtime integration

Technical Direction

  • Direction / acceptance reference, when applicable: N/A; no design or product behavior change.

Shared-authority RFC fixture impact

  • Production-scale fixture schema: N/A; no schema change.
  • Semantic dimensions changed, or reviewed no-impact rationale: N/A; existing response-shape contract only.
  • Provider conformance arms run: N/A.
  • Read-only legacy/file/PostgreSQL three-arm rehearsal: N/A; no provider behavior changed.

Boundary Checklist

  • The diff and PR text contain no private state, credentials, raw traces, internal links, or local paths.
  • No maintainer-owned benchmark work is duplicated.
  • The change is scoped to the stated test-maintenance outcome.
  • UI impact is marked none.
  • The commit includes a DCO Signed-off-by trailer.

Signed-off-by: mika <211269698+mikamikasuki@users.noreply.github.com>
@mikamikasuki

mikamikasuki commented Oct 8, 2026 •

Copy link
Copy Markdown
Contributor Author

Current-head update (run 37784711160; current main/base 73dfd1f; synced PR head 8f083fa): the existing branch is up to date with main via a DCO-signed merge commit. The PR diff against main remains only tests/test_manager_ssh_evidence.py (16 additions, 16 deletions). On the synced head, tests/test_manager_ssh_evidence.py and tests/control_plane/test_todo_exact_detail.py pass: 44 passed in 9.99s; Ruff, Python compilation, and git diff --check also pass. The old CI run's 14 distinct Python failures were replayed on exact base and old PR head (14/14 fail on both), and on current main (13 fail, 1 pass), so none was introduced by this fixture change. The exact-Todo regression itself fails on the old base and passes on the synced head. The old TypeScript-core failure was in the unchanged duplicate digest matcher fixed by merged PR #5977; dashboard timeout/cancellations are separate from this one-file change. Fresh CI for head 8f083fa is active: dependency review, changes validation, and Summary succeeded; DCO, chat-bundle, and node-minimum compatibility are queued; node-forward is skipped; merge-gate is expected. GitHub now shows reviewDecision APPROVED with one approval from loopx-agent; huangruiteng remains a requested code-owner reviewer. Real cross-machine SSH transport is not tested.

@loopx-agent loopx-agent left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewer: model_agent; gpt-6.1-sol; OpenAI; runtime_reported; reasoning_effort=xhigh

动机

维护远端任务证据导出测试的贡献者,需要测试真实的单项查询并保留私有内容过滤。此前查询 t1 时测试替身仍返回任务列表,导出结果为空并报错;现在单项查询返回 todo,列表查询保留 todos。相同两份测试从基线的 43 通过、1 失败变为 44 项通过,真实 File/SQLite 单项和列表导出均未泄露私有笔记。本次修复测试数据契约,不修改 SSH 权限或生产逻辑,也未实测跨机器 SSH 连接。

独立评审 head 765af5dfc2d42b7c8e30e2ae8f31b1f415711ab6,实际不可变 parent aa0c2ea6a6bad1f105e8c2f66464cf4d8c4de2b1。没有独立 RFC;依据变更前已存在的 canonical 单项/列表 API,以及同一命令在 parent 的实际失败,不把 PR 自述当作规范证明。

改动思路

直接修复现有测试替身:同一份 Todo 数据按是否提供选择器返回不同键,生产 canonical reader、detail projection 和 exporter 仍是原 owner。没有必要改生产 API、加入新配置或添加平行 fixture 框架。

具体改动

  • tests/test_manager_ssh_evidence.py:171 提取一个含私有 note 的 todo,两条路径共用相同数据。
  • list_todos(:178)在 todo_id 有值时返回 todo,列表请求返回 todos。既有 [None, "t1"] 参数和全部结果/隐私断言保留。
  • 这对应 chat_manager_details 已存在的单项/列表分支;export_page 继续使用非 owner 视图,返回公开 Goal、Todo 和标题,私有笔记不进入导出。

全部 diff 为一个测试文件 +16/-16,没有生产代码、权限、状态、UI、Lark 或 CLI 变化。未来改造检查认为这个就地 fixture 足够,无须泛化或迁移领域 owner。

对主干的风险

无可行动阻断发现。独立执行同一命令:python -m pytest -q tests/test_manager_ssh_evidence.py tests/control_plane/test_todo_exact_detail.py,parent 为 43 通过、1 失败,head 为 44 通过。parent 唯一失败正是选中 t1 后空 rows 的 IndexError;没有放松断言。

另外把真实 File/SQLite canonical provider 接到生产 export_page,分别验证列表、单项、私有内容排除、missing 返回空、invalid selector 拒绝及源数据未变;两种 provider 都通过。原生 canary 4 项直接检查、3 项选中检查与 Ruff 通过,没有 manual hold,按配置未查询 CI。

远端 SSH 网络执行仍由测试替身承担,实际连接未测。已有 remote exact-selector quoting 测试和权限/撤销/错误路径仍通过,但这些不能替代真实跨机资格。新本地替身只负责已知 t1 场景,未把它的宽松选择行为投射成生产契约。

我的整体评价

APPROVE,head 765af5dfc2d42b7c8e30e2ae8f31b1f415711ab6。 它恢复既有隐私导出守卫的有效性,减少误报和排障成本;运行时用户行为保持不变。现有覆盖扫描和同作者近期 15 项 PR 检查未发现重复 walkthrough 或同形刷量,这次有实际红转绿维护价值。实际 producer/consumer 经验用于选择真实边界反例,未继承历史裁决或宣称经验带来因果收益。

English verdict: APPROVE - 765af5d: the existing test double now reflects canonical exact-todo versus list-todos responses without weakening privacy assertions. The immutable parent reproduces 1 failure/43 passes; the head has 44 passes. Real isolated File/SQLite export and negative/privacy checks, Ruff, and 4 direct/3 canary checks pass; real remote SSH transport remains untested.

Signed-off-by: mika <211269698+mikamikasuki@users.noreply.github.com>
@huangruiteng
huangruiteng merged commit f67a6d8 into loopx-project:main Oct 8, 2026
19 of 24 checks passed
huangruiteng added a commit that referenced this pull request Oct 9, 2026
* fix(control-plane): use canonical digest matcher

Signed-off-by: mika <211269698+mikamikasuki@users.noreply.github.com>

* test(control-plane): pin canonical digest consumer

Signed-off-by: mika <211269698+mikamikasuki@users.noreply.github.com>

* feat(benchmark): allow explicit monitored shared startup margin

Signed-off-by: LoopX Agent <337587101+loopx-agent@users.noreply.github.com>

* feat(benchmark): preserve native ranking with scalar best feedback

Signed-off-by: LoopX Agent <337587101+loopx-agent@users.noreply.github.com>

* fix(benchmark): require native winner for new best signals

Signed-off-by: LoopX Agent <337587101+loopx-agent@users.noreply.github.com>

* refactor(benchmark): use native ranking as sole best feedback rule

Signed-off-by: LoopX Agent <337587101+loopx-agent@users.noreply.github.com>

* refactor(public-safety): decide the opaque id body in one owner (#5880)

* test(manager): align exact Todo evidence mock shape (#5961)

* fix(benchmark): keep empty native winners silent

Signed-off-by: LoopX Agent <337587101+loopx-agent@users.noreply.github.com>

---------

Signed-off-by: mika <211269698+mikamikasuki@users.noreply.github.com>
Signed-off-by: LoopX Agent <337587101+loopx-agent@users.noreply.github.com>
Co-authored-by: mika <spacexstarship01@outlook.com>
Co-authored-by: huangruiteng <huangrt01@163.com>
Co-authored-by: Hsuehtan <haonhsu@gmail.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants