Skip to content

test(explore): align replan qualification with inline context - #5962

Merged
huangruiteng merged 2 commits into
loopx-project:mainfrom
mikamikasuki:codex/loopx-explore-context-qualification
Oct 9, 2026
Merged

huangruiteng merged 2 commits into
loopx-project:mainfrom
mikamikasuki:codex/loopx-explore-context-qualification

Conversation

@mikamikasuki

@mikamikasuki mikamikasuki commented Oct 8, 2026 •

Copy link
Copy Markdown
Contributor

Goal And Delivered Outcome

  • Outcome basis / optional anchor: Self-contained maintenance outcome; no separate issue is required for this scope.
  • Goal/source and gap: The replan semantic-action qualification treated Explore turn context as a standalone top-level required_reads command. Quota packets deliver the turn-start hook result inline through interaction_contract.agent_channel.work_context.sources; the stale helper failed before the modeled action and its skip/replay/scope cases no longer matched the packet contract. The Explore README described the previous delivery shape.
  • Observable before → after: At the PR's original base 5f51559dc93b43d023a9c915c52ec1dc409e3e65, the replan qualification file had 4 failures and 19 passes. The actor fixture now verifies the scoped, read-only Explore context in the quota packet before the next action and does not perform a duplicate Explore read. The README describes the inline contract.
  • Issue/task and intended base: Direct maintenance PR against main; current base is 73dfd1f66dadbd7c08b25a94813c4d51ec99d1fe.

Author Declaration

  • Written by: model_agent — OpenAI GPT-6 Luna

Some coding work was AI-assisted. I personally reviewed the changes, verified the relevant behavior and tests, and take responsibility for the submission.

Implemented against

  • Specification and revision: The quota context-delivery contract in loopx/control_plane/work_items/context_readback.py and tests/capabilities/test_explore_turn_context.py::test_real_quota_delivers_replayable_context_without_admitting_unhealthy_goal.
  • Criteria:
Criterion Disposition Symbol / path Test
The admitted Explore source is available in the quota packet before the next agent action. implemented interaction_contract.agent_channel.work_context.sources tests/control_plane/test_replan_semantic_action_behavior.py
An inline Explore source is not duplicated as a separate required read. implemented work_context.sources and required_reads Replan actor and Explore turn-context tests
Explicit Explore required reads, when present, use the nested interaction contract. implemented _required_explore_read_command test_required_explore_read_uses_nested_interaction_contract
  • Self-check: Reproduced the stale fixture failure at the original base, traced the packet projection and contract-test history, searched for matching issues/PRs, and reviewed the final three-file diff. No production runtime behavior changes.

Scope And Continuation

  • Completed scope: Align the replan actor qualification and Explore README with the inline context contract.
  • Slice boundary: Complete within this scope; no successor is needed.

Validation

  • Tested revision: PR head caedbafcba3621c12f7bfe0a2ba218838803b3b0, synced with canonical main 73dfd1f66dadbd7c08b25a94813c4d51ec99d1fe.
  • Run state: Local validation finished. Current-head hosted CI is running; the full repository suite did not complete.
  • Input classes: synthetic.
Check kind Result Evidence / limitation
Unit passed Five focused suites covering replan semantics, Explore turn context, selection reentry, periodic lookback, and effective-turn replanning: 82 passed in 98.06s.
Static passed Ruff on changed Python files, Python compilation, worktree diff check, and base-to-head diff check.
Pre-merge passed loopx canary premerge --from-git-diff: 19 selected checks passed, zero failures, zero manual holds.
Hosted CI in progress New head run 37832856777 and associated checks started at 2026-10-08 19:33 UTC. GitHub Summary succeeded; DCO, dependency review, Frontstage Pages, PostgreSQL integration, and Release Artifacts were queued at last readback; merge-gate is expected.
Full suite incomplete An earlier python -m pytest -q -n 2 run was interrupted around 25% after failures appeared in untouched tests/architecture/test_semantic_* tests. No complete full-suite result is claimed.
  • CI attribution: The earlier Frontstage Pages failure on head a342ba63cebbc908b1766c5771a2cdfb999d7260 was in the unchanged bilingual-link smoke and reproduced on its exact base. The semantic link-check correction is tracked separately in #5968; it is outside this PR's diff.
  • Coverage and gaps: Scripted transport and synthetic Goal fixtures verify the quota packet shape and bounded tool sequence; the public Explore quota test independently verifies inline source and replay command. No live provider or external backend was used.

Frontend / Visual Evidence

  • UI impact: none
  • Before/after: N/A
  • Source data: synthetic

Type of Change

  • Bug fix
  • New feature
  • Breaking change
  • Refactoring (no functional changes)
  • Documentation update
  • Test update

LoopX Area

  • Control plane (goals, todos, gates, claims, evidence, quota, runtime)
  • Benchmark boundary (adapters, runners, verifiers, scoring, evidence)
  • Capability or extension (providers, adapters, skills)
  • Public docs or presentation surface (README, protocols, dashboard)
  • Build, packaging, installer, or CI
  • Host or runtime integration

Technical Direction

N/A; this change aligns a qualification fixture and capability documentation with the existing packet contract.

Shared-authority RFC fixture impact

  • Production-scale fixture schema: N/A.
  • Semantic dimensions changed: None; production schemas and authority behavior are unchanged.
  • Provider conformance arms and three-backend rehearsal: N/A.

Boundary Checklist

  • No private state, credentials, raw traces, internal links, or local machine paths in the diff or PR text.
  • No duplicate maintainer-owned benchmark work.
  • Scope is limited to the current Explore context contract.
  • Visual evidence completed or marked UI impact none.
  • Every commit includes a DCO Signed-off-by trailer.

@mikamikasuki

Copy link
Copy Markdown
Contributor Author

Follow-up for the replan-recovery fixture: its prior setup inserted surface-only run-history rows without matching quota admission receipts or settlement writebacks, so it could not legitimately trigger the documented settled-effective-Turn periodic review. On canonical main 9d7680a, the original focused case failed with long_todo_chain as the only trigger; the same assertion failed on the pre-fix PR head. The test now uses that valid replan trigger while keeping periodic cadence covered by the dedicated settlement/effective-turn tests. Validation after the fix: 27 passed across test_selection_replan_reentry.py, test_autonomous_replan_periodic_lookback.py, and test_effective_turn_replan.py on both PR head and current main; Ruff and git diff --check passed. Updated head: a342ba63cebbc908b1766c5771a2cdfb999d7260.

@mikamikasuki

Copy link
Copy Markdown
Contributor Author

CI update for head a342ba6: Frontstage Pages build run 37811578012 fails in the unchanged examples/blog-bilingual-index-smoke.mjs at edgebench-feedback-and-memory (line 102). This failure reproduces on the exact PR base 5f51559; #5962 does not touch that script or the article. It is an exact-href false negative: the existing relative link resolves to the same paired-edition URL. The narrowly scoped correction is in PR #5968. I have left #5962 unchanged for this unrelated failure; the updated Python checks are still running.

@mikamikasuki

Copy link
Copy Markdown
Contributor Author

The failing typescript-core (3/3) check in run 37832857599 is in the unchanged tests/control_plane_ts/host_process.test.ts, outside this PR's diff. CI observed the 4-second renewal fixture hit the proved deadline before the test's expected renewal rejection. I opened the focused test-only correction in #5989; the old fixture fails under a controlled 3-second renewal delay, while the revised fixture passes the same delayed case and its Node 22.22.3 focused suite. #5962 still needs a fresh successful hosted run after that fix lands; I am not counting the current failed check as passed.

Signed-off-by: mika <211269698+mikamikasuki@users.noreply.github.com>
Signed-off-by: mika <211269698+mikamikasuki@users.noreply.github.com>
@mikamikasuki
mikamikasuki force-pushed the codex/loopx-explore-context-qualification branch from caedbaf to 86f3a79 Compare October 9, 2026 01:30
@mikamikasuki

Copy link
Copy Markdown
Contributor Author

Validation update for head 86f3a79, rebased on canonical main b27c4c7.

The remaining shard failure is part of this PR's fixture scope: the exact [repeat] case failed in three clean runs on current main because the fixture expected a top-level required_reads entry while the quota packet carries Explore context inline under interaction_contract.agent_channel.work_context.sources. The updated qualification checks that inline context and rejects a duplicate read.

Validation on the rebased head: both affected modules passed (31 tests); Ruff, Python compilation, and git diff --check origin/main...HEAD passed. Both PR commits retain DCO sign-off.

Hosted checks have restarted for this head; the browser currently shows Summary passing, six jobs queued, and merge-gate expected. Code-owner review is still pending. The full repository suite has not completed.

@mikamikasuki mikamikasuki mentioned this pull request Oct 9, 2026
8 of 12 tasks

@loopx-agent loopx-agent left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewer: model_agent — gpt-6.1-sol / OpenAI; runtime_reported; reasoning_effort=xhigh.

动机

APPROVE,当前完整 head 没有发现阻塞问题。

维护重规划资格测试的开发者,在 quota 已把 Explore 上下文随结果返回时,会被旧测试误导为还缺少一次工具读取。
旧 fixture 从已移除的顶层 required_reads 取 Explore 命令,组合路径在真正行动前失败;新版先核验包内 scoped、read-only 上下文,再沿原 frontier 选择与回写。
新 head 的真实 CLI 资格、作用域与重放验证通过;文档与现有单一内容载体一致,恢复选择继续按原 Turn 结算一次。
这次交付是既有资格测试和文档维护,不改变生产准入、配额、重规划或 provider 权限;没有证明真实模型采纳、远端服务或全库测试通过。

精确 head:86f3a792e46108f43ac7bcbfe0cece06d1afd74c;不可变 base:b27c4c7290b9fea5deac1afa54694f3d5c263b93。作者正文的旧 head、旧 base、三文件和测试数没有作为当前事实采用;本轮核验两个提交、全部四文件和实际 CLI。

改动思路

复用现有 TypeScript interaction owner 的来源投影和结算规则;Python 仅修正资格工具的读取位置,避免重新实现生产决策或增加第二份上下文。
当前 PR 完成真实 CLI 资格与现有内联契约对齐,保留未完成必读项、作用域、重放和周期结算约束;不新增产品操作或通用协议。

quota 先通过原 hook/source adapter 取得上下文,TypeScript owner 把已提供来源放在 work_context、未完成读取保留在嵌套 required_reads。测试消费当前包,再执行原 frontier/source 和 typed successor/writeback,实际准入、来源失败、重规划、身份及配额结算仍由原 Core 负责。把顶层旧命令重新塞回生产包或删掉 pending 检查都会制造第二契约,因此本 PR 只调整现有资格 seam。

具体改动

关键代码讲解

  • _required_explore_read_command(loopx/control_plane/testing/replan_semantic_action_behavior.py:537):先读嵌套 agent_channel.required_reads,再筛选 Explore;嵌套列表不存在才保留原测试 helper fallback。具体 source、before_work、Goal、Agent 仍逐项校验。空的当前嵌套列表优先于旧顶层指针。
  • _inline_explore_context_action(tests/control_plane/test_replan_semantic_action_behavior.py:92):检查真实 quota 已附带的唯一来源、准确身份、read_only 与 replay command,拒绝重复 pending Explore,然后回到原 frontier。它是 scripted fixture 观察,不能证明真实模型已理解上下文。
  • selection recovery 测试(tests/control_plane/test_selection_replan_reentry.py:18):移除没有准入/结算回执的人工 surface-only 行,用真实 long_todo_chain 验证恢复;原 Turn 身份、fresh admission、reentry 与只扣一次的断言继续保留。

规范依据:docs/reference/required-work-context.md + docs/reference/protocols/goal-vision-replan-contract-v0.md,固定版本 b27c4c7290b9fea5deac1afa54694f3d5c263b93。Required work context 为 implemented:已提供来源不重复读取,未完成项仍精确且有顺序,来源失败须恢复重准入,内容不给权限。Default review cadence 为 implemented:周期按 verified settled effective Turns 计数;selection 的 long-chain 触发与专门周期证明分开。修改的 Explore README 对照这些既有规范,不用自己的新文字证明自己正确。

对主干的风险

五个 source 模块独立执行 head 82 passed,base 5 failed / 79 passed。base 四项失败来自旧顶层读取/旧 skip-repeat-scope 假设,另一个失败是未结算历史不能触发 periodic_review_due。新版删除的 obsolete 用例没有被当作“相同测试自然变绿”:相同独立 8 项 nested oracle 在 head 8 passed、base 7 failed / 1 passed,验证合法 mixed pending list、重复项、错误 Goal、错误 Agent、错误 source、错误 order、legacy-only 兼容与 authoritative empty 优先。

相关 source suites 经真实 loopx CLI、生产 TypeScript owner 和隔离文件状态运行;Explore replay 读取与内联 content 一致,unhealthy Goal 仍 should_run=false/delivery_allowed=false。组合路径创建 scoped successor 并重新准入,选择恢复按原 Turn 结算一次;专门 effective-turn/periodic suites 仍拒绝不合格计数。模型 transport 是 scripted,不是一次真实付费模型验收。

Ruff、diff whitespace、两个提交 DCO、变更 advisory 及两端完整 semantic smoke 通过。风险式 premerge 的 5 项直接检查与 19 项选择检查全部通过,零失败、零 warning、零 manual holds。没有新增共享词汇、权限或默认开关;未配置 Explore 时原普通路径保持。README 改的是载体描述,读取当前 sources before work、freshness、scope、bounded detail 和建议非强制的条款保留,没有把义务缩成可忽略提示。未查询或等待 CI;全库 suite、live model/provider、安装后的应用未测。

我的整体评价

交付判断 goal_achieved 仅指这次资格与文档维护:开发者不再遇到旧载体造成的假失败,user_experience 为 improved;长期执行所需的来源失败、重放、scope、真实结算与周期 owner 保持,long_horizon 为 preserved。没有把局部维护称作 Explore 全能力或父级 Goal 完成。

复用现有 TypeScript interaction owner 的来源投影和结算规则;Python 仅修正资格工具的读取位置,避免重新实现生产决策或增加第二份上下文。
当前 PR 完成真实 CLI 资格与现有内联契约对齐,保留未完成必读项、作用域、重放和周期结算约束;不新增产品操作或通用协议。

Future-facing pass 保留生产 TS 单一 owner,Python 仅适配已有测试 seam;这个有界修复无需新增状态机或扩大重构。嵌套 empty 与 legacy fallback 的取舍已通过反例核验。Approval 不代表合并授权、全部旧架构 producer 安全或真实模型采纳。

English verdict: APPROVE — exact head 86f3a79. The qualification and documentation now follow the existing inline work-context contract without duplicating fulfilled reads or fabricating settled-turn cadence. Eight identical nested-read counterexamples and 82 actual CLI source cases pass at head; production context, admission and settlement owners remain unchanged. Live model adoption and the complete repository suite are unverified. No remote CI was consulted; merge readiness and authority remain separate.

@huangruiteng
huangruiteng merged commit 5f01436 into loopx-project:main Oct 9, 2026
25 of 32 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants