diff --git a/docs/architecture/rfcs/hierarchical-agent-stride-control-v0.md b/docs/architecture/rfcs/hierarchical-agent-stride-control-v0.md index 3f2a4e559e..e290d5629d 100644 --- a/docs/architecture/rfcs/hierarchical-agent-stride-control-v0.md +++ b/docs/architecture/rfcs/hierarchical-agent-stride-control-v0.md @@ -404,6 +404,84 @@ protected-operation policy remain authoritative for permission. Supervisors, dashboards, and scheduler projections may recommend intervention but cannot silently convert a proposal into authority. +### 7.5 Long-running work and decision checkpoints (proposed refinement) + +This 2026-10-04 refinement is a proposal for M3/M5, not shipped adaptive +scheduling. It separates an **experiment boundary**, a **model decision +checkpoint**, and an **accountable Turn boundary**. One experiment can span +several Turns; every admitted accountable Turn still validates, writes back, +and settles its own identity. Waiting for the experiment does not defer an +owed debit, keep a Turn open indefinitely, or make earlier work free. + +The nearest owning capability/provider must define, before launch: + +- the question, candidate revision, comparison and bounded postcondition; +- the authorized job identity, resource/time limits and recovery owner; +- which observable result can change the next decision, and how to validate + that result; a training checkpoint, evaluation or build is meaningful only + under this declared contract; +- the last consumed result revision, next observation time/deadline and + stop/cancel behavior; elapsed time and log growth alone are not evidence. + +These are capability/provider data. The kernel consumes identity, readiness, +freshness and existing effect/settlement facts; it does not interpret epochs, +benchmark scores, proof counts or process log prose. No generic job scheduler, +worker launcher or new built-in capability is needed for this proposal. + +The proposed handoff is: + +| Phase | Owner and required behavior | +|---|---| +| Start bounded work | A normal selected Todo/Turn admits the effect. The provider persists job/candidate identity and an operational receipt. Settle that Turn truthfully: accepted launch proves launch, not the final experiment outcome. | +| Observe while waiting | The existing host/runtime or capability observer checks the exact job without a reasoning-model call. Reuse one due monitor or existing observation channel; do not create a Todo per poll. Record observer CPU/IO cost even when no agent slot is spent. | +| Return a decision checkpoint | The provider validates a new result against the declared postcondition, binds its revision/artifact digest and emits a compact observation. Failure, cancellation and deadline expiry must also produce actionable observations. | +| Admit a decision | The existing scheduler/Turn path rechecks Goal lifetime, ownership, quota, capabilities and authority. A ready observation is input, not execution permission. Other runnable work can proceed while this experiment waits. | +| Consume and continue | In the accountable consuming Turn, commit result adoption and the next step through the normal writeback before acknowledging consumption. Replayed/duplicate delivery recovers that receipt without duplicating launch or adoption. Settlement replay is idempotent for the same Turn identity; a newly admitted recovery Turn still settles normally. | + +Current Todo resume conditions include `monitor_changed` and `resume_at`; an +implementation should use the existing typed monitor generation and timeout +fallback where sufficient. **The missing integration is the provider-result +binding and material-generation update**, not a new arbitrary `job_done` token. +Only a validated decision checkpoint may advance that generation. The monitor +must name the exact job/result source and bound Todo; unrelated log changes +cannot release it. A safety observation deadline catches a lost notification +or crashed process; it is not a periodic request to reinterpret unchanged work. +Unsupported hosts retain current execution and show unavailable readiness, +rather than silently claiming model-wake suppression. + +Identity must survive host/Agent restart: Goal lifetime, owning Agent/Todo, +provider/run id, attempt or job generation, candidate revision, result revision +and evidence/artifact refs. Keep raw logs and native job handles with the +provider; expose only bounded authorized references. A PID alone is not a +recoverable identity. Provider absence or an ambiguous external effect remains +unknown under existing recovery rules; it cannot justify an automatic relaunch. +The consumed cursor advances only with validated durable adoption, so a crash +after result delivery re-offers the result. Stopped/revoked Goals and expired +claims fence late results; replacement attempts cannot consume old results. + +Persist the incumbent artifact reference and its qualifying result separately +from the running candidate. The provider updates it only after a comparable +valid result satisfies its declared constraints. Recovery verifies reference +availability and content before reuse; a missing candidate artifact does not +silently replace the incumbent or authorize another launch. + +The first experiment is a **shadow observation** over one real capability +runtime: report unchanged polls, decision checkpoints, result-to-decision +latency, native outcome and total model/observer cost. Then qualify opt-in +model-admission suppression only for that supported provider/host combination. +Do not change the generic heartbeat prompt or all scheduler defaults. Turning +the policy off restores fixed admission while retaining job/evidence history; +it neither kills work implicitly nor changes acceptance or accounting. + +Acceptance requires real controlled process execution and recovery, not just +serialized events: duplicate and out-of-order results, changed candidates, +restart, missed notification, deadline, cancellation, unknown effect, competing +claim and owner stop. A packaged frontend journey must show the bound work, +waiting reason, last/next observation, result and stale/unavailable recovery +through existing task/capability surfaces. CLI readback alone is a partial +slice. Scientific usefulness follows the independent qualification in +[research decisions](research-exploration-control-plane-v0.md#115-evidence-conditioned-experiment-decisions-proposed-refinement). + ## 8. Measurement Model The first implementation should measure before it controls. @@ -652,6 +730,8 @@ policy framework. transitions; - validate recommendations against independent acceptance rules; - retain Todo and replan as the only execution authorities. +- qualify the proposed §7.5 provider checkpoint/monitor binding in shadow; + distinguish model decisions from experiment completion and Turn settlement. ### M4: Authority-stride shadow recommendations @@ -665,6 +745,8 @@ policy framework. - retain hard ceilings and rollback; - compare repeated results against the pinned fixed profile; - publish limitations and failure modes with any claimed improvement. +- for long-running work, qualify §7.5 lost-result recovery and default-off + parity before suppressing model admission; observer cost remains measured. ## 13. Validation Criteria diff --git a/docs/architecture/rfcs/hierarchical-agent-stride-control-v0.zh-CN.md b/docs/architecture/rfcs/hierarchical-agent-stride-control-v0.zh-CN.md index 8486b28878..15997e658b 100644 --- a/docs/architecture/rfcs/hierarchical-agent-stride-control-v0.zh-CN.md +++ b/docs/architecture/rfcs/hierarchical-agent-stride-control-v0.zh-CN.md @@ -364,6 +364,75 @@ Goal 与 vision state 继续拥有 acceptance 权限。User gate 与 protected-o policy 继续拥有 permission 权限。Supervisor、dashboard 与 scheduler projection 可以建议干预,但不能静默把 proposal 变成 authority。 +### 7.5 长作业与决策 checkpoint(修订提案) + +这项 2026-10-04 修订属于 M3/M5 提案,不代表 adaptive scheduling 已交付。 +需要分开**实验边界**、**模型决策 checkpoint**和**可记账 Turn 边界**。 +一个实验可以跨多个 Turn;每个已获准且需要记账的 Turn 仍以自己的 identity +完成 validation、writeback 和 settlement。等待实验不会延后欠付 debit、让 +Turn 无限保持打开,或把此前工作变成免费。 + +最近的 capability/provider owner 应在启动前声明: + +- 问题、candidate revision、比较对象和有界 postcondition; +- 已授权的 job identity、资源/时间上限和 recovery owner; +- 什么可观察结果会影响下一决策,以及如何验证;训练 checkpoint、评测或 + 构建只有在这份已声明的合同下才有决策意义; +- 已消费 result revision、下次观测时间/deadline、stop/cancel 行为;仅经过 + 一段时间或日志增长不是证据。 + +这些是 capability/provider 数据。Kernel 消费 identity、readiness、freshness +及现有 effect/settlement 事实,不解释 epoch、benchmark 分数、证明数量或进程 +日志 prose。本提案不需要新增通用 job scheduler、worker launcher 或 built-in +capability。 + +拟议 handoff 如下: + +| 阶段 | Owner 与必需行为 | +|---|---| +| 启动有界作业 | 普通已选 Todo/Turn 准入 effect。Provider 持久化 job/candidate identity 与 operational receipt。该 Turn 如实结算:启动成功只证明启动,不证明实验最终结果。 | +| 等待中的观测 | 现有 host/runtime 或 capability observer 无需 reasoning-model 调用即可检查精确 job。复用一个 due monitor 或已有 observation channel,不为每次 poll 创建 Todo。即使不消费 agent slot,也记录 observer CPU/IO 成本。 | +| 返回决策 checkpoint | Provider 按已声明的 postcondition 验证新结果,绑定 revision/artifact digest 并返回紧凑 observation。失败、取消和 deadline 到期同样必须形成可操作 observation。 | +| 准入决策 | 现有 scheduler/Turn 路径重新检查 Goal lifetime、ownership、quota、capability 与 authority。Ready observation 是输入,不是执行许可。等待此实验时仍可推进其他 runnable 工作。 | +| 消费并继续 | 在负责消费的可记账 Turn 中,通过普通 writeback 提交 result adoption 与下一步骤,再确认消费。重放/重复交付恢复原 receipt,不重复 launch 或 adoption。同一 Turn identity 的结算重放幂等;新获准 recovery Turn 仍正常结算。 | + +当前 Todo resume condition 已有 `monitor_changed` 与 `resume_at`。若足够, +实现应复用 typed monitor generation 与 timeout fallback。 +**缺失的是 provider-result binding 和 material-generation 更新集成**,不是 +新增任意 `job_done` token。只有通过验证的决策 checkpoint 才能推进 generation。 +Monitor 必须绑定精确 job/result source 与 Todo;无关日志变化不能释放等待。 +Safety observation deadline 用来发现丢失通知或崩溃进程,不是周期性要求模型 +重新解释未变工作。不支持的 host 保持现有执行并显示 readiness unavailable, +不能暗称已抑制模型唤醒。 + +Identity 必须跨 host/Agent restart 保留:Goal lifetime、owning Agent/Todo、 +provider/run id、attempt 或 job generation、candidate revision、result revision +与 evidence/artifact ref。Raw log/native job handle 留在 provider,仅投影有界 +且授权的引用。单个 PID 不是可恢复 identity。Provider 缺失或外部 effect 状态 +不明确时,按现有 recovery 规则保留 unknown,不能据此自动 relaunch。只有 +validated durable adoption 才推进 consumed cursor,因此交付后崩溃会重新提供 +结果。Stopped/revoked Goal 和过期 claim 拦截 late result;替换 attempt 不能 +消费旧结果。 + +Incumbent artifact reference 及其 qualifying result 与当前运行 candidate 分开 +持久化。Provider 仅在可比有效结果满足已声明 constraint 后更新 incumbent。 +Recovery 在复用前验证引用可用性与内容;candidate artifact 缺失不能静默 +替换 incumbent,也不授权再次 launch。 + +先在一个真实 capability runtime 做 **shadow observation**:报告 unchanged poll、 +decision checkpoint、result-to-decision latency、原生结果和模型/observer 总成本。 +随后仅对该 provider/host 组合资格验证 opt-in model-admission suppression。 +不改变通用 heartbeat prompt 或所有 scheduler 默认行为。关闭策略恢复 fixed +admission,保留 job/evidence history,不隐式杀作业,不改变 acceptance 或 accounting。 + +验收必须覆盖真实受控进程执行与 recovery,不只是 event 序列化:重复与乱序 +结果、candidate 改变、restart、丢失通知、deadline、取消、unknown effect、竞争 +claim 和 owner stop。Packaged frontend 旅程通过已有 task/capability surface 显示 +绑定工作、等待原因、上次/下次观测、结果及 stale/unavailable 恢复;仅 CLI +readback 属于 partial slice。科学效用遵循 +[研究决策](research-exploration-control-plane-v0.zh-CN.md#115-由证据驱动的实验决策修订提案) +的独立资格验证。 + ## 8. 测量模型 第一版实现应先测量,再控制。 @@ -597,6 +666,9 @@ executor、scheduler 或 generic policy framework。 - 用独立 acceptance rule 验证建议; - Todo 与 replan 仍是唯一 execution authority。 +- 在 shadow 中验证 §7.5 的 provider checkpoint/monitor binding,区分模型 + 决策、实验完成与 Turn settlement。 + ### M4:Authority-stride shadow recommendation - 区分 report 与 authority-changing intervention; @@ -610,6 +682,9 @@ executor、scheduler 或 generic policy framework。 - 对 pinned fixed profile 做 repeated comparison; - 任何 improvement claim 都同时公开 limitation 与 failure mode。 +- 长作业抑制 model admission 前,验证 §7.5 的 lost-result recovery 与 + default-off parity;observer 成本继续纳入测量。 + ## 13. 验证标准 只有证明以下各项,研究计划才算合格: diff --git a/docs/architecture/rfcs/long-horizon-harness-benchmark-research-program-v0.md b/docs/architecture/rfcs/long-horizon-harness-benchmark-research-program-v0.md index 29f54f3b9b..671ec2e877 100644 --- a/docs/architecture/rfcs/long-horizon-harness-benchmark-research-program-v0.md +++ b/docs/architecture/rfcs/long-horizon-harness-benchmark-research-program-v0.md @@ -655,6 +655,16 @@ layer at a time: The study should estimate model- and work-class-specific response curves, not one global tool-call or Todo-count threshold. Wider is not automatically better. +The proposed [long-running-work refinement](hierarchical-agent-stride-control-v0.md#75-long-running-work-and-decision-checkpoints-proposed-refinement) +tests **decision timing**, not deferred Turn accounting. First capture one +provider's validated result/checkpoint and bound monitor in shadow; only a +qualified opt-in arm may suppress unchanged model admission. Keep native +evaluation, job resources, deadlines and every owed Turn settlement identical. +Measure observer CPU/IO, model tokens/calls, result-to-decision latency, missed +results, duplicate execution/debits and native outcome per total budget. A +missing callback must recover through the declared observation deadline. +Neither a longer training run nor fewer bookkeeping calls alone qualifies it. + ### 8.2 Evidence, stall detection, and semantic replan The core hypothesis is that a durable coverage ledger plus semantic progress @@ -668,6 +678,18 @@ whether direction changes create progress. DeepSWE validates whether the same mechanism improves repository outcomes without reward-specific shortcuts. ALE tests transfer to heterogeneous professional workflows. +The proposed [experiment decision path](research-exploration-control-plane-v0.md#115-evidence-conditioned-experiment-decisions-proposed-refinement) +qualifies result adoption separately from cadence. Hold session continuity, +planning, model and total budget fixed; compare delivered evidence alone with +evidence-linked continuation/successor adoption, then combine mechanisms only +after their independent comparisons. Include valid negative evidence, +inconclusive/conflicting measurements, failed prerequisites, justified repeats, +uncovered same-family probes and changed inputs reopening a retired scope. +Assess whether the executed next action follows the scoped result, alongside +native outcome and cost. Citation/schema compliance alone is insufficient; +forced pivots that suppress useful experiments count as failures. These are +proposed studies, not current C2/C4 evidence or permission to launch jobs. + ### 8.3 Research exploration and composition The [Research Exploration Control Plane RFC](./research-exploration-control-plane-v0.md) diff --git a/docs/architecture/rfcs/long-horizon-harness-benchmark-research-program-v0.zh-CN.md b/docs/architecture/rfcs/long-horizon-harness-benchmark-research-program-v0.zh-CN.md index 957987a6b3..b031e19e61 100644 --- a/docs/architecture/rfcs/long-horizon-harness-benchmark-research-program-v0.zh-CN.md +++ b/docs/architecture/rfcs/long-horizon-harness-benchmark-research-program-v0.zh-CN.md @@ -552,6 +552,14 @@ non-material event mix;它不能证明 benchmark integrity,也不能复用 研究应估计 model/work-class-specific response curve,而不是一个全局 tool-call 或 Todo-count threshold。更宽并不天然更好。 +拟议[长作业修订](hierarchical-agent-stride-control-v0.zh-CN.md#75-长作业与决策-checkpoint修订提案) +测试**决策时机**,不延后 Turn 记账。先 shadow capture 一个 provider 的 validated +result/checkpoint 和绑定 monitor;只有通过资格验证的 opt-in arm 才能抑制 +unchanged model admission。原生评测、job 资源、deadline 与所有欠付 Turn +settlement 保持一致。测量 observer CPU/IO、模型 token/call、result-to-decision +latency、漏结果、重复执行/debit 与总预算下原生结果。Callback 丢失必须按已声明 +observation deadline 恢复。训练更久或 bookkeeping call 更少,单独都不够资格。 + ### 8.2 Evidence、空转检测与 semantic replan 核心假设是 durable coverage ledger 与 semantic progress observation 能阻止重复 maintenance @@ -563,6 +571,15 @@ LHTB 是首要动态实验场,因为 partial reward 与 checkpoint 可以显 DeepSWE 验证相同机制能否改善 repository outcome,且不依赖 reward-specific shortcut。 ALE 验证它能否迁移到异构专业工作流。 +拟议[实验决策路径](research-exploration-control-plane-v0.zh-CN.md#115-由证据驱动的实验决策修订提案) +把 result adoption 与 cadence 分别验证。固定 session continuity、planning、model +和总预算,对比仅交付 evidence 与 evidence-linked continuation/successor adoption, +独立比较成立后再组合机制。包含有效负证据、inconclusive/conflicting measurement、 +失败前置条件、有依据复验、未覆盖的同 family probe 与 input 改变后重开 scope。 +除原生结果和成本外,判断实际执行的 next action 是否遵循 scoped result。 +Citation/schema compliance 单独不够;压制有用实验的强制 pivot 算失败。 +这些是拟议 study,不是当前 C2/C4 evidence,也不授予 job launch 权限。 + ### 8.3 研究探索与组合 [研究型探索控制面 RFC](./research-exploration-control-plane-v0.zh-CN.md)定义 typed research diff --git a/docs/architecture/rfcs/loopx-overall-roadmap-v0.md b/docs/architecture/rfcs/loopx-overall-roadmap-v0.md index bac4e0a661..4908390f4a 100644 --- a/docs/architecture/rfcs/loopx-overall-roadmap-v0.md +++ b/docs/architecture/rfcs/loopx-overall-roadmap-v0.md @@ -329,8 +329,8 @@ subsystem was not performed. Section 8 records the focused audit. | [Human Attention Wishlist v0](human-attention-wishlist-v0.md) | S5/S11 | Accepted; Held | P3: reopen only on repeated second real need; sidecar cannot alter gates/quota/scheduling | | [Human-confirmed domain operations (v0)](human-confirmed-domain-operations-v0.md) | S8/S9 + R2/R3 | Accepted; canonical operation seam exists, managed native transport under owner review | Qualify source-context vs admitted-executor separation, exact human approval→one-shot consumption→domain evidence→original-route return; owned Turn/delegation need not await Desktop authentication. Live approval/effect/wakeup remain unqualified; finance provider stays separate | | [Provider-side authorization at effect acceptance (v0)](provider-effect-acceptance-v0.md) | S8/S9, supporting S2/S4 | Accepted design only; no runtime integration or qualified provider | M1: controlled provider and deterministic revoke/crash/replay conformance; strict production wiring remains gated by exact Goal lifetime, receipt retention and independent provider qualification | -| [Research Exploration Control Plane v0](research-exploration-control-plane-v0.md) | S11/S3 | Accepted; partial M2 composition/successor | P1: independently verify observation/write-time gate/closure basis; defer inferred triggers and model selection | -| [Hierarchical Agent Stride Control v0](hierarchical-agent-stride-control-v0.md) | S11/S7 | Accepted; M1 read-only observation | P2: matched shadow stride experiment with costs/events; no direct production cadence change | +| [Research Exploration Control Plane v0](research-exploration-control-plane-v0.md) | S11/S3 | Accepted; partial M2 composition/successor; §11.5 result/adoption refinement proposed | P1: one real provider result -> scoped interpretation -> adopted step/successor, exact lineage/shared write gate and packaged readback; qualify shadow before opt-in, retain composition gaps; defer inference and M4 portfolio ranking | +| [Hierarchical Agent Stride Control v0](hierarchical-agent-stride-control-v0.md) | S11/S7 | Accepted; M1 read-only observation; §7.5 checkpoint/monitor binding proposed | P2: matched shadow decision timing with observer/model costs and lost-result recovery; preserve per-Turn settlement, no default cadence change | | [Goal-scoped Capability Portfolio v0](goal-scoped-capability-portfolio-v0.md) | S1/S3/S6/S8/S11 | Accepted; read-only settings/context inspection slice | P1: correction→fresh-session decision across existing owners, then demand-driven composition and measured method evolution; no universal memory database or second opt-in | | [Post-Outcome Memory Utility Attribution v0](post-outcome-memory-utility-attribution-v0.md) | S6/S11 | Accepted; Stage 1 verified-outcome binding | P1 read-only reducer→P2 pilot: distinguish recalled/applied/utility, no automatic ranking change | | [Obelisk Session Evidence Provider v0](obelisk-session-evidence-provider-v0.md) | S6/S8 | Accepted; optional read-only evaluation | P2: explicit gap recall with resolvable provenance/access, off-path parity; no work authority | diff --git a/docs/architecture/rfcs/loopx-overall-roadmap-v0.zh-CN.md b/docs/architecture/rfcs/loopx-overall-roadmap-v0.zh-CN.md index cd3a62ef1e..d512d26a59 100644 --- a/docs/architecture/rfcs/loopx-overall-roadmap-v0.zh-CN.md +++ b/docs/architecture/rfcs/loopx-overall-roadmap-v0.zh-CN.md @@ -266,8 +266,8 @@ Muse 设计页在浏览器超时,其文章通过网页检索读取。本次调 | [Human Attention Wishlist v0](human-attention-wishlist-v0.zh-CN.md) | S5/S11 | 已接受;Held | P3:第二个重复真实需求出现才重开;sidecar 不改变 gate/quota/调度 | | [Human-confirmed domain operations (v0)](human-confirmed-domain-operations-v0.zh-CN.md) | S8/S9 + R2/R3 | 已接受;规范操作接缝存在,受管原生传输待 owner review | 验收来源上下文与已准入执行者分离、精确用户批准→单次消费→垂域证据→原路返回;自有 Turn/delegation 不等待 Desktop 认证。真实批准/效果/唤醒仍未资格化;金融 provider 保持独立 | | [Provider 在效果接受点执行授权(v0)](provider-effect-acceptance-v0.zh-CN.md) | S8/S9,S2/S4 支撑 | 已接受设计;尚未接入 runtime,也未准入 provider | M1:controlled provider 与 deterministic revoke/crash/replay conformance;strict production 接线仍需精确 Goal 生命周期、receipt retention 与独立 provider 资格 | -| [Research Exploration Control Plane v0](research-exploration-control-plane-v0.zh-CN.md) | S11/S3 | 已接受;M2 composition/successor 局部实现 | P1:observation/write-time gate/closure basis 独立验证;自选模型和推断触发继续 defer | -| [Hierarchical Agent Stride Control v0](hierarchical-agent-stride-control-v0.zh-CN.md) | S11/S7 | 已接受;M1 只读观测 | P2:matched shadow stride 实验,定义代价与事件;不直接改变生产节奏 | +| [Research Exploration Control Plane v0](research-exploration-control-plane-v0.zh-CN.md) | S11/S3 | 已接受;M2 composition/successor 局部实现;§11.5 result/adoption 修订为提案 | P1:一条真实 provider result -> 有范围 interpretation -> adopted step/successor,精确 lineage/shared write gate 与 packaged readback;shadow 资格先于 opt-in,保留 composition gap;推断与 M4 portfolio 排序继续 defer | +| [Hierarchical Agent Stride Control v0](hierarchical-agent-stride-control-v0.zh-CN.md) | S11/S7 | 已接受;M1 只读观测;§7.5 checkpoint/monitor binding 为提案 | P2:matched shadow 决策时机、observer/model 成本与 lost-result recovery;保留每 Turn settlement,不改变默认 cadence | | [Goal-scoped Capability Portfolio v0](goal-scoped-capability-portfolio-v0.zh-CN.md) | S1/S3/S6/S8/S11 | 已接受;只读配置/上下文检查切片 | P1:复用既有 owner 验收纠正→新会话决策,再做按需组合与可测方法演化;不建万能记忆库或第二个启用开关 | | [Post-Outcome Memory Utility Attribution v0](post-outcome-memory-utility-attribution-v0.zh-CN.md) | S6/S11 | 已接受;Stage 1 verified-outcome 绑定 | P1 只读 reducer→P2 pilot:区分 recalled/applied/utility,归因不自动改 ranking | | [Obelisk Session Evidence Provider v0](obelisk-session-evidence-provider-v0.zh-CN.md) | S6/S8 | 已接受;可选只读评估 | P2:显式 gap recall、来源与权限可解析、关闭不影响主流程;不当 work authority | diff --git a/docs/architecture/rfcs/research-exploration-control-plane-v0.md b/docs/architecture/rfcs/research-exploration-control-plane-v0.md index 2c5bfa54ac..b7c32355fe 100644 --- a/docs/architecture/rfcs/research-exploration-control-plane-v0.md +++ b/docs/architecture/rfcs/research-exploration-control-plane-v0.md @@ -154,6 +154,7 @@ milestone status. | No exact obligation/Todo/result lineage | A research receipt is not proof of an authorized Todo transition or accepted Goal closure. | | Cold shadow not adopted by hot status/frontier | The existing #3173 projection remains behavior-compatible; canonical research obligations still need M3 integration. | | No dismissal or deferral contract | Evidence invalidation is visible, but typed candidate retirement and resumption remain unimplemented. | +| No experiment-result adoption gate | Negative evidence can be recorded without proving that the next bound decision used its scope, validity or uncertainty. Proposed single-experiment integration is defined in §11.5; current cold receipts do not enforce it. | | Live qualification incomplete | Deterministic and real CLI/file-log tests establish state semantics, not model selection quality or scientific truth; no live Lark sync is qualified by projection tests. | | No promotion evidence for inferred combinations | Shared constraints are not known to be precise enough to trigger obligations. | @@ -600,6 +601,119 @@ turn discovering which evidence command satisfies a protocol ceremony. A delivery receipt proves that context arrived; the observation proves whether the next action used it. +### 11.5 Evidence-conditioned experiment decisions (proposed refinement) + +This 2026-10-04 proposal refines M3 before M4 portfolio selection. It does not +claim a shipped experiment controller or change ordinary replan defaults. +[Decision checkpoints](hierarchical-agent-stride-control-v0.md#75-long-running-work-and-decision-checkpoints-proposed-refinement) +determine when a result needs attention; this section defines what a research +decision must retain. Explore is the nearest capability owner. Native +measurement, statistical interpretation and job/artifact handling stay with +the domain or benchmark provider. Generic kernel rules do not interpret scores +or infer research truth from Todo churn. + +#### Bind a question and the scope of its evidence + +For the first real caller, compose an explicit versioned Explore attachment; +do not add unknown fields to research observation v0. Its conceptual data are: + +| Data | Owning meaning | +|---|---| +| Hypothesis and declared mechanism family | Stable Explore node/revision and explicit relation to a family node. Similar names/prose do not prove equivalence; a family is a scoped mechanism claim, not a kernel-inferred taxonomy. | +| Probe and applicability | Planned intervention, candidate/input fingerprints, comparison, environment/data/evaluator revision, scope, resource/stop bounds. The provider declares what would support, contradict or leave the claim unresolved. | +| Result and validity | Exact provider result revision/evidence, comparison and uncertainty/guardrail assessment. Failed, missing or incomparable measurements remain distinct from valid negative evidence. | +| Decision and adoption | Exact input/result fingerprints, existing `continue`/`no_change`/`replan` path outcome, evidence refs and bound task step or grounded successor. A repeat needs a new discriminating question or an explicit replication/uncertainty reason and budget. | + +This is a design inventory, not an accepted wire schema or new global enum. +Explore owns the future TS codec/reducer; compute/capture providers adapt it +without a parallel decision owner. Reuse nodes, findings, `supports`/`refutes`, +closure and input-fingerprint semantics. Compose a versioned decision +attachment with the existing Goal path delta and Todo-bound recommendation +receipts, rather than add another Next Action or settlement store. An ML +adapter can map its existing hypothesis ledger's weakened/retired states; +that vocabulary does not become a generic progress result class. + +| Evidence | Legal decision basis | +|---|---| +| Build or measurement failed | Repair the prerequisite or record the blocker; do not assert scientific refutation. | +| One valid candidate is worse | Preserve the incumbent and bind evidence to this candidate/probe scope. Continue another probe, replicate or replan with reasons; do not automatically retire the family. | +| Effect is within uncertainty or measurements conflict | Preserve unresolved/contradictory evidence; choose a bounded discriminating probe or justified replication. A small score decline is not drift. | +| Declared coverage contradicts a hypothesis | Close/revise that exact claim with its closure basis. A successor cites the exclusion and what changes; a same-mechanism probe addressing an uncovered condition remains legal. | +| Coverage supports family retirement | Retire only the declared family/scope with attributable evidence. New input, evaluator or applicability can reopen investigation; retirement never settles Goal acceptance by itself. | + +Rejecting one parameter setting can leave a family open; demonstrating failure +to meet a declared constraint under covered conditions can justify a different +family. The kernel validates attribution and transitions, not causal truth or +experiment quality. Renaming identifiers or allocating a new evidence id for +an unchanged measurement is not novelty. No fixed failure count forces a pivot. + +#### Make result adoption observable at the existing decision boundary + +1. Record/read back the validated result on the exact experiment node with + provider provenance and current input fingerprints. A launch receipt alone + is not research evidence. +2. Under an explicit per-Goal opt-in, project one compact pending decision for + its bound experiment/Todo: question, incumbent, applicable exclusions, + uncertainty/invalidity, input/result refs and remaining budget. Building the + projection needs neither a full log nor a new reasoning-model call. +3. The normal admitted Agent decides. Continuing is legal if it states what + new information the next probe should obtain and why the exclusion does not + already answer it. A pivot names the changed question/mechanism and its + evidence through the existing grounded successor path. +4. Explore's typed reducer qualifies the decision attachment. The shared write + gate joins it to the exact Todo/replan/Turn and current input revisions, + committing it with the normal task-step or successor receipt. Generic + replan keeps its existing legal outcomes. A record/read, stale citation or + unrelated Todo cannot discharge this opt-in decision duty. +5. Re-offer pending results after a crash; acknowledge only durable adoption. + Replay recovers existing decision receipts; same-Turn settlement remains + idempotent, while a newly admitted recovery Turn settles normally. Changed + inputs or results invalidate pending decisions; historical facts remain + available. + +The required relation is **result -> interpretation -> adopted action**. +Binding is not proof of quality. Operational validity, new research information, +native outcome improvement and Goal acceptance are separate observations. A +useful negative result can be delivery without a score gain; a series of +negative experiments must still expose changed knowledge and remaining frontier. + +#### Smallest usable path and qualification + +Start with the existing **benchmark-toolkit admitted-run monitor/closeout** +seam, composed with an Explore experiment node. Its +[`benchmark_runtime_observation_v0` and continuity reducers](../../../loopx/capabilities/benchmark_toolkit/README.md) +already separate runtime liveness, attributable terminal results and launch +generation. They classify provider facts; they neither run experiments nor +establish result validity. A controlled provider must execute and recover one +bounded build/evaluation trial, validate its result, deliver the pending +decision, adopt an evidence-linked continuation/successor and read it back. +Initially use a complete trial result; intermediate training checkpoints need +their own provider postcondition and qualification. A preview-only domain +pack is not an execution backend. This bounded M3 +prerequisite need not wait for inferred composition, a new multi-agent +scheduler or M4 ranking, and does not close the remaining M3 composition work. + +First shadow-capture actual decisions against negative, inconclusive, failed +and invalidated inputs without changing scheduling. Enforcement is a separate, +reversible per-Goal opt-in in the existing Explore configuration owner. +Default-off shared surfaces retain current behavior. Disabling stops new +decision duties, preserves evidence and completes owed settlement; it does +not delete job/result history or waive existing replan/authority obligations. + +The enforcing slice includes the existing frontend task/Explore journey: +question, candidate/incumbent, waiting/result state, scoped negative evidence, +proposed next action and reason; an authorized decision and its readback, +including stale-result recovery. Optional Lark consumes the same projection +and authority. A CLI-only prerequisite remains partial until the packaged +interaction is qualified. + +Qualify real provider process and canonical state, revision/lease fences, +duplicate/lost delivery, stop/revocation, missing and contradictory results, +unjustified replay/renaming, over-retirement, changed inputs reopening a scope, +valid same-family probes and default-off parity. Then independently qualify +model adoption and scientific utility under matched total budgets. Lower +control cost must not come from hiding useful negative evidence. + ## 12. Write-Time Enforcement The write gate must reuse the same current goal-frontier and Explore gap @@ -804,7 +918,7 @@ control-plane failures. | M0 | RFC, current-state inventory, and explicit ownership decision | Maintainer review; no runtime behavior | Accepted design | | M1 | Characterization fixtures plus typed research observation and closure contract in Explore | Deterministic normalization, privacy, compatibility, and negative tests | Implemented evidence/CLI slice; live research qualification remains separate | | M2 | Explicit-only composition candidate, canonical gap projection, and read-only status shadow | No pairwise inference; bounded packet; projection parity | Partial: #3173 legacy quota/successor; canonical binary cold shadow in CLI/Lark projection; hot status adoption and live Lark qualification remain | -| M3 | Goal-frontier obligation, exact Todo/experiment lineage, and shared write-time gate | State/replay matrix and premerge canary pass | Not started | +| M3 | Goal-frontier obligation, exact Todo/experiment lineage, shared write-time gate; proposed §11.5 single-experiment result/adoption prerequisite | State/replay matrix, real provider/state readback, default-off parity and packaged journey before enforcement | Not started; §11.5 is a design proposal, not delivered runtime | | M4 | Bounded multi-candidate cards, `composition_selection_v0`, real model-tool behavior qualification, and repeated live shadow | Model autonomously selects a legal semantic action from the delivered candidate set; selection quality is no worse than the declared fallback; compact receipts only | Not started | | M5 | Shared-constraint candidate ranking in shadow mode | Precision and cost evidence; no automatic trigger | Not started | | M6 | Optional inferred trigger | Explicit maintainer decision and measured promotion thresholds | Deferred | @@ -829,7 +943,11 @@ composition-frontier/result-layer checks. It covers canonical reverse pairs, terminal coverage, attribution, replay, invalidated inputs, stale experiment lineage and real CLI/file-log readback. Projection tests do not establish live remote effects or model-selected research behavior. Delivery is tracked in -[#5214](https://github.com/loopx-project/loopx/issues/5214); the issue remains open. +[#5214](https://github.com/loopx-project/loopx/issues/5214), now closed. Its closure +does not supply missing live/M3 qualification. Remaining research direction is +tracked by [#4391](https://github.com/loopx-project/loopx/issues/4391) and +[#3243](https://github.com/loopx-project/loopx/issues/3243); preserve the explicit +milestone gaps here rather than create a parallel task tree. M3 is the first behavior-changing slice. It should be a separate PR so the obligation and write gate can be reviewed and reverted independently from the @@ -918,6 +1036,7 @@ This is a living RFC, not an append-only diary. |---|---| | 2026-08-13 | Adopt Explore as the canonical research-topology owner; choose explicit-only composition candidates for v0; represent joint work as an experiment node; defer shared-constraint inference to shadow qualification. | | 2026-08-13 | Separate eligibility from ranking: the control plane owns a legal bounded candidate set, while the model autonomously prioritizes among multiple eligible candidates. A selection receipt proves a scheduling choice, not research truth. Defer the protocol to M4 rather than adding it to the #3173 runtime slice. | +| 2026-10-04 | Propose §11.5 as a bounded M3 prerequisite: exact result -> scoped interpretation -> adopted task step/successor. Reuse Explore and task/replan receipts; keep Turn settlement, native scoring and scientific qualification independent. This RFC update delivers no automatic policy. | ## 20. Acceptance Criteria for the RFC diff --git a/docs/architecture/rfcs/research-exploration-control-plane-v0.zh-CN.md b/docs/architecture/rfcs/research-exploration-control-plane-v0.zh-CN.md index a497f608fa..56a2dc1b28 100644 --- a/docs/architecture/rfcs/research-exploration-control-plane-v0.zh-CN.md +++ b/docs/architecture/rfcs/research-exploration-control-plane-v0.zh-CN.md @@ -135,6 +135,7 @@ hypothesis”或“新 probe family”。它无法持久表达:A 和 B 都已 | 没有精确 obligation/Todo/result lineage | Research receipt 不证明已授权 Todo transition 或已接受 Goal closure。 | | Cold shadow 未接入 hot status/frontier | 既有 #3173 投影保持行为兼容;canonical research obligation 仍需 M3 集成。 | | 没有 dismissal/deferral contract | Evidence invalidation 可见,但类型化 candidate retirement/resumption 尚未实现。 | +| 没有 experiment-result adoption gate | 负证据可被记录,但不能证明下一绑定决策使用了其 scope、validity 与 uncertainty。§11.5 提议 single-experiment 集成;当前 cold receipt 不强制执行该提案。 | | Live qualification 不完整 | Deterministic 与真实 CLI/file-log 测试证明状态语义,不证明 model selection 质量或科学结论;projection 测试不构成 live Lark sync 资格。 | | inferred combination 没有 promotion evidence | 共享 constraint 的精度还不足以直接触发 obligation。 | @@ -553,6 +554,101 @@ Host 应投影: 满足协议仪式。Delivery receipt 证明 context 到达;observation 证明下一步是否 真正使用了 context。 +### 11.5 由证据驱动的实验决策(修订提案) + +这项 2026-10-04 提案细化 M3,先于 M4 portfolio selection;不代表实验控制器 +已交付,不改变普通 replan 默认行为。 +[决策 checkpoint](hierarchical-agent-stride-control-v0.zh-CN.md#75-长作业与决策-checkpoint修订提案) +决定何时关注结果,本节定义研究决策应保留什么。最近的 capability owner 是 +Explore;原生测量、统计解释和 job/artifact 处理留在 domain/benchmark provider。 +Generic kernel 不解释分数,也不从 Todo 更新推断研究真相。 + +#### 绑定问题与证据适用范围 + +第一个真实 caller 组合显式 versioned Explore attachment,不向 research +observation v0 偷加 unknown field。概念数据如下: + +| 数据 | 所属语义 | +|---|---| +| Hypothesis 与已声明 mechanism family | 稳定 Explore node/revision,显式关联 family node。相似名字/prose 不证明等价;family 是有范围的机制假设,不是 kernel 推断的分类。 | +| Probe 与 applicability | 计划 intervention、candidate/input fingerprint、比较对象、环境/数据/evaluator revision、scope、资源/停止边界。Provider 声明什么支持、反驳或不能确定该假设。 | +| Result 与 validity | 精确 provider result revision/evidence、比较和 uncertainty/guardrail assessment。失败、缺失、不可比的测量与有效负证据区分。 | +| Decision 与 adoption | 精确 input/result fingerprint,现有 `continue`/`no_change`/`replan` path outcome、evidence ref 与绑定 task step 或 grounded successor。重复 probe 要有新的判别问题,或明确的 replication/uncertainty 理由和预算。 | + +这是设计清单,不是已接受 wire schema 或新 global enum。未来 TS codec/reducer +归 Explore;compute/capture provider 适配,不新增平行决策源。复用 node、finding、 +`supports`/`refutes`、closure 和 input-fingerprint 语义。Versioned decision +attachment 与现有 Goal path delta、Todo-bound recommendation receipt 组合, +不新增 Next Action 或 settlement store。ML adapter 可映射已有 hypothesis ledger +的 weakened/retired 状态,但这套词表不成为 generic progress result class。 + +| 证据 | 合法决策依据 | +|---|---| +| 构建或测量失败 | 修复前置条件或记录 blocker,不声称科学假设被反驳。 | +| 一个有效 candidate 更差 | 保留 incumbent,证据绑定该 candidate/probe scope。带理由选择其他 probe、复验或 replan,不自动退休整个 family。 | +| 效果处于 uncertainty 内或测量冲突 | 保留 unresolved/contradictory evidence,选择有界判别 probe 或有依据复验。小幅掉分不是 drift。 | +| 已声明 coverage 反驳 hypothesis | 以 closure basis 关闭/修订精确假设;successor 引用排除结论与改变之处。针对未覆盖条件的同机制 probe 仍合法。 | +| Coverage 支持退休 family | 只退休有 attributable evidence 的已声明 family/scope。新的 input、evaluator 或 applicability 可重开调查;退休不直接 settle Goal acceptance。 | + +否定一个参数仍可保留 family;在覆盖条件内证明不能满足已声明 constraint, +才可能支持换 family。Kernel 验证 attribution 与 transition,不认证因果真相或 +实验质量。仅重命名 identifier 或为未变测量分配新 evidence id 不是 novelty。 +不用固定失败次数强制 pivot。 + +#### 在现有决策边界让 adoption 可观察 + +1. 在精确 experiment node 记录/读回 validated result,保留 provider provenance + 与当前 input fingerprint。Launch receipt 单独不是 research evidence。 +2. 显式 per-Goal opt-in 时,为绑定 experiment/Todo 投影紧凑 pending decision: + 问题、incumbent、适用 exclusion、uncertainty/invalidity、input/result ref 和 + 剩余预算。构造投影不需要完整日志或新增 reasoning-model 调用。 +3. 由正常获准 Agent 决策。继续合法,但须说明下一 probe 要获得什么新信息, + 以及 exclusion 为什么尚未回答该问题。Pivot 指出改变的问题/机制与证据, + 使用已有 grounded successor 路径。 +4. Explore typed reducer 验证 decision attachment。Shared write gate 关联 + 精确 Todo/replan/Turn 与当前 input revision,随普通 task-step/successor + receipt 提交。Generic replan 保留已有合法 outcome。Record/read、stale + citation 或无关 Todo 不能解除 opt-in decision duty。 +5. Crash 后重交 pending result,仅 durable adoption 后确认消费。Replay 恢复 + 原 decision receipt;同一 Turn 的结算重放幂等,新获准 recovery Turn 仍正常 + 结算。Input/result 改变使 pending decision 失效,历史事实继续保留。 + +要求的是 **result -> interpretation -> adopted action**。绑定不证明质量。 +Operational validity、新研究信息、原生结果提升和 Goal acceptance 是不同 +observation。有效负实验可以是合法交付而没有涨分;连续负实验仍须暴露知识 +变化与剩余 frontier。 + +#### 最小有用路径与资格验证 + +先接现有 **benchmark-toolkit admitted-run monitor/closeout** seam,与 Explore +experiment node 组合。其 +[`benchmark_runtime_observation_v0` 与 continuity reducer](../../../loopx/capabilities/benchmark_toolkit/README.md) +已有 runtime liveness、可归属 terminal result 和 launch generation 的区分。 +它们分类 provider fact,不运行实验,也不认证 result validity。受控 provider +须实际执行/恢复一次有界 build/evaluation trial、验证结果、交付 pending decision、 +adopt 带证据 continuation/successor 并读回。先使用完整 trial result;训练中间 +checkpoint 要另有 provider postcondition 与资格验证。Preview-only domain pack +不是 execution backend。 +有界 M3 prerequisite 不必等待 inferred composition、新 multi-agent scheduler +或 M4 排序,也不关闭 M3 剩余 composition 工作。 + +先 shadow capture:按 negative、inconclusive、failed、invalidated input 评估 +实际决策,不改变 scheduling。Enforcement 是现有 Explore configuration owner +中的独立、可回滚 per-Goal opt-in;默认关闭时共享 surface 保持当前行为。 +关闭停止派生新 decision duty,保留 evidence 并完成 owed settlement,不删除 +job/result history,不豁免已有 replan/authority obligation。 + +Enforcing slice 包含现有 frontend task/Explore 旅程:问题、candidate/incumbent、 +waiting/result state、有范围负证据、拟议 next action 与理由;已授权决策和 +readback,包含 stale-result recovery。Optional Lark 消费同源 projection/authority。 +CLI-only prerequisite 在 packaged interaction 验证前仍是 partial。 + +资格验证覆盖真实 provider process 与 canonical state、revision/lease fence、 +重复/丢失交付、stop/revocation、缺失/冲突结果、无依据 replay/改名、过度退休、 +input 改变后重开 scope、合法同 family probe 与 default-off parity。随后在 +matched 总预算下独立验证模型 adoption 和科学效用;降低控制成本不能靠隐藏 +有效负证据。 + ## 12. 写时强制 Write gate 必须复用 quota 所用的同一 current goal-frontier 与 Explore gap @@ -742,7 +838,7 @@ rule,以及 model variance 与 control-plane failure 的分离。 | M0 | RFC、current-state inventory 与显式 ownership decision | Maintainer review;无 runtime behavior | 已接受的设计 | | M1 | Characterization fixture,以及 Explore 中的 typed research observation 与 closure contract | Deterministic normalization、privacy、compatibility 与 negative test | Evidence/CLI 切片已实现;真实研究 qualification 独立保留 | | M2 | Explicit-only composition candidate、canonical gap projection 与 read-only status shadow | 不做 pairwise inference;packet 有界;projection parity | 部分实现:#3173 legacy quota/successor;CLI/Lark projection 的 canonical binary cold shadow;hot status adoption 与 live Lark qualification 仍未完成 | -| M3 | Goal-frontier obligation、精确 Todo/experiment lineage 与共享 write-time gate | State/replay matrix 与 premerge canary 通过 | 未开始 | +| M3 | Goal-frontier obligation、精确 Todo/experiment lineage、共享 write-time gate;§11.5 提议 single-experiment result/adoption prerequisite | State/replay matrix、真实 provider/state readback、default-off parity;enforcement 前验证 packaged journey | 未开始;§11.5 是设计提案,不是已交付 runtime | | M4 | 有界 multi-candidate card、`composition_selection_v0`、真实 model-tool behavior qualification 与重复 live shadow | 模型从交付 candidate set 中自主选择合法 semantic action;选择质量不劣于 declared fallback;只保留 compact receipt | 未开始 | | M5 | Shared-constraint candidate 在 shadow mode 中排序 | 有 precision/cost evidence;不自动触发 | 未开始 | | M6 | 可选 inferred trigger | 显式 maintainer decision 与量化 promotion threshold | 延后 | @@ -766,7 +862,11 @@ M1/M2 evidence 切片交付: result-layer 检查。覆盖反向配对 canonical identity、terminal coverage、attribution、 replay、input invalidation、stale experiment lineage 和真实 CLI/file-log 读回。 Projection 测试不证明 live remote effect 或模型自主研究行为。交付继续由 -[#5214](https://github.com/loopx-project/loopx/issues/5214) 跟踪,该 issue 保持打开。 +[#5214](https://github.com/loopx-project/loopx/issues/5214) 记录,该 issue 已关闭。 +关闭不补齐缺失的 live/M3 qualification。后续研究方向由 +[#4391](https://github.com/loopx-project/loopx/issues/4391) 与 +[#3243](https://github.com/loopx-project/loopx/issues/3243) 跟踪;本 RFC 保留明确 +milestone gap,不另建平行 task tree。 M3 是第一个 behavior-changing slice。它应单独成 PR,使 obligation 与 write gate 能够独立于 evidence schema 评审和回滚。 @@ -848,6 +948,7 @@ evidence-backed terminal result。 |---|---| | 2026-08-13 | 采用 Explore 作为 canonical research-topology owner;v0 选择 explicit-only composition candidate;joint work 表示为 experiment node;shared-constraint inference 延后到 shadow qualification。 | | 2026-08-13 | 将 eligibility 与 ranking 分离:控制面拥有合法、有界 candidate set,模型在多个 eligible candidate 中自主择优;selection receipt 只证明调度选择,不构成 research truth。该协议延后到 M4,不进入 #3173 的首期 runtime。 | +| 2026-10-04 | 提议 §11.5 作为有界 M3 prerequisite:精确 result -> 有范围 interpretation -> adopted task step/successor。复用 Explore 和 task/replan receipt;Turn settlement、原生评分与科学资格独立。本次 RFC 更新不交付自动策略。 | ## 20. RFC 验收标准