Skip to content

test(authority): refresh todo content revision after edit - #5991

Merged
huangruiteng merged 1 commit into
loopx-project:mainfrom
mikamikasuki:codex/loopx-todo-content-revision-test
Oct 9, 2026
Merged

huangruiteng merged 1 commit into
loopx-project:mainfrom
mikamikasuki:codex/loopx-todo-content-revision-test

Conversation

@mikamikasuki

Copy link
Copy Markdown
Contributor

Goal And Delivered Outcome

  • Outcome basis: A reproducible failure in the canonical Todo authority test on main.
  • Goal/source and gap: After a Todo text edit, the test compared the updated record with its original snapshot and required the original content_revision to remain unchanged.
  • Observable before to after: On base 3ed5d6b, the focused test fails because the updated text produces digest sha256:3b61f6402feb68632608a56817e5dd4ffdf0b041d60b0d1d40a1d9f0a4305036 while the assertion expects the old digest. The test now derives the expected revision from the edited text and verifies that it differs from the prior revision.
  • Issue/task and intended base: No separate issue is required for this self-contained test correction. Target main at 3ed5d6b.

Author Declaration

  • Written by: model_agent — OpenAI GPT-6 Luna Medium

Implemented against

  • Specification and revision: loopx/control_plane/todos/summary_item.py::todo_text_content_revision at 3ed5d6b; Todo freshness behavior is also described in PR #5911.
  • Criteria:
Criterion Disposition Symbol / path Test or command
An edited Todo's content revision is derived from its normalized text and differs from the prior revision implemented tests/control_plane/test_local_coordination_authority.py test_real_canonical_provider_preserves_complete_complex_todo_semantics
  • Self-check before submission: Reproduced the stale assertion on the exact base, confirmed the digest function uses normalized Todo text, searched for duplicate work, and reviewed the one-file test-only diff.

Scope And Continuation

  • Completed scope and remaining work: Corrected the expected content revision after the test edits provider-owned Todo text; production behavior is unchanged.
  • Slice boundary / successor: Complete within this scope.

Validation

  • Tested revision: 3b48964
  • Run state: finished
  • Input classes: synthetic
Check kind Result Public-safe evidence
regression_parity passed Exact-base focused test failed on the stale digest expectation; the corrected assertion passes on this head.
unit passed Focused canonical-provider Todo update test: 1 passed.
static passed Ruff and git diff --check passed for the changed test.
  • Coverage and gaps: The focused case checks the canonical provider readback after a Todo text update and preserves existing sibling-field assertions. No production code changed.

Frontend / Visual Evidence

  • UI impact: none
  • Before: N/A
  • After: N/A
  • States and viewports shown: N/A
  • Source data: synthetic
  • Attention review: N/A

Type of Change

  • Test update

LoopX Area

  • Control plane (goals, todos, quota, scheduler, registry, runtime)

Technical Direction

  • Direction / acceptance reference, when applicable: N/A; this is a test-only correction to existing Todo revision behavior.

Shared-authority RFC fixture impact

  • Production-scale fixture schema: N/A
  • Semantic dimensions changed, or reviewed no-impact rationale: N/A; no authority schema changed.
  • Provider conformance arms run: N/A
  • Read-only legacy/file/PostgreSQL three-arm rehearsal: N/A

Boundary Checklist

  • No private state, credentials, raw traces, internal links, or local machine paths are included.
  • No maintainer-owned benchmark work is duplicated.
  • The change is scoped to this test correction.
  • UI impact is marked none.
  • Every commit includes a DCO Signed-off-by trailer.

Signed-off-by: mika <211269698+mikamikasuki@users.noreply.github.com>

@loopx-agent loopx-agent left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewer: model_agent; model=gpt-6.1-sol; provider=OpenAI; runtime_reported; reasoning_effort=xhigh.

Reviewed exact head 3b4896427c13158e2d775de14669aae406ac4ea2 against immutable base 3ed5d6bc88ffaaa05f2b1bc6a98ed5d921379fd1. APPROVE,无阻塞发现。

动机

维护者修改 Todo 兼容入口后,需要用既有真实 provider 回归检查完整记录是否保留。
旧测试把编辑前的正文摘要带入编辑后的预期记录,正确编辑也会失败;当前测试重新计算新正文摘要,并要求它与旧摘要不同。
同一不可变基线复现旧摘要断言失败;当前精确 head 的完整测试文件 38 项通过,复杂记录与相邻任务的原有完整相等断言仍保留。
本 PR 只修正既有测试,不改变生产 Todo、权限、存储格式或 CLI/UI/Lark 行为,也不验收所有 provider 或平台。

完整读取已合并的 #5911 及基线全文摘要 helper 后判断:这是维护已有正文新鲜度契约的测试预期,不是新增 canonical 持久化字段。没有定义此断言修复的独立书面规格,因此不虚构 RFC 验收编号。

改动思路

继续使用同一个真实 canonical File fixture、实际 CLI 和完整记录相等断言。预期摘要取自声明的新正文,随后明确要求新旧摘要不同;直接从实际读回取预期值或删掉摘要比较都会削弱回归。复用现有规范化摘要 owner,未增加第二个生产决策源或单独 smoke。

具体改动

整个 PR 只有 tests/control_plane/test_local_coordination_authority.py,+9/-1:导入 todo_text_content_revision,定义 edited_text 和预期新摘要,检查非空,在完整 edited-record 预期中覆盖摘要,再断言它不同于 claimed_item 的旧摘要。原 claim、dry-run/apply、原 Markdown 移除后真实 canonical 读回、复杂 metadata、相邻 complex/successor 完整相等及 timestamp 检查全部保留。

独立对照:基线运行同一 focused function,1 failed/37 deselected,17 个字段相同,只有旧摘要与新摘要不匹配。当前 head 完整测试文件 38 passed。另用独立 hashlib 对声明新正文执行规范化并计算 SHA256,得到 3b61f6402feb68632608a56817e5dd4ffdf0b041d60b0d1d40a1d9f0a4305036,与 helper 一致且不同于旧值;预期没有从 provider 实际输出倒推。

对主干的风险

主要风险是修测试时悄悄放宽完整记录保留条件。全量 diff 和实际 fixture 执行表明,此改动只修正本应变化的派生正文摘要,仍比较其他字段及相邻任务,新旧不等检查也会拒绝陈旧摘要。Ruff、diff check、DCO 通过;未查询或等待 CI。

现有覆盖扫描区分上游 selection/concurrency 摘要测试与这个下游完整记录兼容 fixture;同作者的 hook/successor/report/process 测试 PR 维护不同边界,本 PR 没有新增同形 smoke。验证使用合成数据上的真实 File provider/CLI,不冒充 SQLite、PostgreSQL、所有平台或安装 App 验收;生产 TypeScript/权限/存储规则和用户入口没有改动。

我的整体评价

APPROVE。基线反例证明旧预期错误,当前完整文件和独立摘要 oracle 支持这个有界维护结果,没有弱化原有数据保留断言。未来重构检查已考虑 fixture 与摘要 owner:复用当前 helper 和既有测试足够,无需追加框架。评审结论不授予 merge authority。

English verdict: APPROVE — exact head 3b48964; fixes the stale expected text revision without weakening full canonical-record or sibling preservation. Immutable-base failure, 38 current-head tests, an independent normalized-text SHA256 oracle and focused static/DCO checks support this test-only maintenance change.

@mikamikasuki

Copy link
Copy Markdown
Contributor Author

CI attribution for run 37847658548 at head 3b48964: JUnit artifact python-junit-2 reports three failing tests. I reran those exact selectors on this PR's base 3ed5d6b; all three fail there as well (3 failed in 1.60s). They are in test files already changed by open PRs #5986 (test_todo_complete_settlement_capabilities.py), #5993 (test_doubao_model_behavior_actor.py), and #5980 (test_chat_machine_configuration_api.py). This PR changes only test_local_coordination_authority.py, and its focused regression passed locally. The remaining test shards are still running, so this is an interim attribution rather than a terminal CI result.

@mikamikasuki mikamikasuki mentioned this pull request Oct 9, 2026
8 of 12 tasks
@huangruiteng
huangruiteng merged commit f070471 into loopx-project:main Oct 9, 2026
23 of 29 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants