Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
82 changes: 82 additions & 0 deletions docs/architecture/rfcs/hierarchical-agent-stride-control-v0.md
Original file line number Diff line number Diff line change
Expand Up @@ -404,6 +404,84 @@ protected-operation policy remain authoritative for permission. Supervisors,
dashboards, and scheduler projections may recommend intervention but cannot
silently convert a proposal into authority.

### 7.5 Long-running work and decision checkpoints (proposed refinement)

This 2026-10-04 refinement is a proposal for M3/M5, not shipped adaptive
scheduling. It separates an **experiment boundary**, a **model decision
checkpoint**, and an **accountable Turn boundary**. One experiment can span
several Turns; every admitted accountable Turn still validates, writes back,
and settles its own identity. Waiting for the experiment does not defer an
owed debit, keep a Turn open indefinitely, or make earlier work free.

The nearest owning capability/provider must define, before launch:

- the question, candidate revision, comparison and bounded postcondition;
- the authorized job identity, resource/time limits and recovery owner;
- which observable result can change the next decision, and how to validate
that result; a training checkpoint, evaluation or build is meaningful only
under this declared contract;
- the last consumed result revision, next observation time/deadline and
stop/cancel behavior; elapsed time and log growth alone are not evidence.

These are capability/provider data. The kernel consumes identity, readiness,
freshness and existing effect/settlement facts; it does not interpret epochs,
benchmark scores, proof counts or process log prose. No generic job scheduler,
worker launcher or new built-in capability is needed for this proposal.

The proposed handoff is:

| Phase | Owner and required behavior |
|---|---|
| Start bounded work | A normal selected Todo/Turn admits the effect. The provider persists job/candidate identity and an operational receipt. Settle that Turn truthfully: accepted launch proves launch, not the final experiment outcome. |
| Observe while waiting | The existing host/runtime or capability observer checks the exact job without a reasoning-model call. Reuse one due monitor or existing observation channel; do not create a Todo per poll. Record observer CPU/IO cost even when no agent slot is spent. |
| Return a decision checkpoint | The provider validates a new result against the declared postcondition, binds its revision/artifact digest and emits a compact observation. Failure, cancellation and deadline expiry must also produce actionable observations. |
| Admit a decision | The existing scheduler/Turn path rechecks Goal lifetime, ownership, quota, capabilities and authority. A ready observation is input, not execution permission. Other runnable work can proceed while this experiment waits. |
| Consume and continue | In the accountable consuming Turn, commit result adoption and the next step through the normal writeback before acknowledging consumption. Replayed/duplicate delivery recovers that receipt without duplicating launch or adoption. Settlement replay is idempotent for the same Turn identity; a newly admitted recovery Turn still settles normally. |

Current Todo resume conditions include `monitor_changed` and `resume_at`; an
implementation should use the existing typed monitor generation and timeout
fallback where sufficient. **The missing integration is the provider-result
binding and material-generation update**, not a new arbitrary `job_done` token.
Only a validated decision checkpoint may advance that generation. The monitor
must name the exact job/result source and bound Todo; unrelated log changes
cannot release it. A safety observation deadline catches a lost notification
or crashed process; it is not a periodic request to reinterpret unchanged work.
Unsupported hosts retain current execution and show unavailable readiness,
rather than silently claiming model-wake suppression.

Identity must survive host/Agent restart: Goal lifetime, owning Agent/Todo,
provider/run id, attempt or job generation, candidate revision, result revision
and evidence/artifact refs. Keep raw logs and native job handles with the
provider; expose only bounded authorized references. A PID alone is not a
recoverable identity. Provider absence or an ambiguous external effect remains
unknown under existing recovery rules; it cannot justify an automatic relaunch.
The consumed cursor advances only with validated durable adoption, so a crash
after result delivery re-offers the result. Stopped/revoked Goals and expired
claims fence late results; replacement attempts cannot consume old results.

Persist the incumbent artifact reference and its qualifying result separately
from the running candidate. The provider updates it only after a comparable
valid result satisfies its declared constraints. Recovery verifies reference
availability and content before reuse; a missing candidate artifact does not
silently replace the incumbent or authorize another launch.

The first experiment is a **shadow observation** over one real capability
runtime: report unchanged polls, decision checkpoints, result-to-decision
latency, native outcome and total model/observer cost. Then qualify opt-in
model-admission suppression only for that supported provider/host combination.
Do not change the generic heartbeat prompt or all scheduler defaults. Turning
the policy off restores fixed admission while retaining job/evidence history;
it neither kills work implicitly nor changes acceptance or accounting.

Acceptance requires real controlled process execution and recovery, not just
serialized events: duplicate and out-of-order results, changed candidates,
restart, missed notification, deadline, cancellation, unknown effect, competing
claim and owner stop. A packaged frontend journey must show the bound work,
waiting reason, last/next observation, result and stale/unavailable recovery
through existing task/capability surfaces. CLI readback alone is a partial
slice. Scientific usefulness follows the independent qualification in
[research decisions](research-exploration-control-plane-v0.md#115-evidence-conditioned-experiment-decisions-proposed-refinement).

## 8. Measurement Model

The first implementation should measure before it controls.
Expand Down Expand Up @@ -652,6 +730,8 @@ policy framework.
transitions;
- validate recommendations against independent acceptance rules;
- retain Todo and replan as the only execution authorities.
- qualify the proposed §7.5 provider checkpoint/monitor binding in shadow;
distinguish model decisions from experiment completion and Turn settlement.

### M4: Authority-stride shadow recommendations

Expand All @@ -665,6 +745,8 @@ policy framework.
- retain hard ceilings and rollback;
- compare repeated results against the pinned fixed profile;
- publish limitations and failure modes with any claimed improvement.
- for long-running work, qualify §7.5 lost-result recovery and default-off
parity before suppressing model admission; observer cost remains measured.

## 13. Validation Criteria

Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -364,6 +364,75 @@ Goal 与 vision state 继续拥有 acceptance 权限。User gate 与 protected-o
policy 继续拥有 permission 权限。Supervisor、dashboard 与 scheduler projection
可以建议干预,但不能静默把 proposal 变成 authority。

### 7.5 长作业与决策 checkpoint(修订提案)

这项 2026-10-04 修订属于 M3/M5 提案,不代表 adaptive scheduling 已交付。
需要分开**实验边界**、**模型决策 checkpoint**和**可记账 Turn 边界**。
一个实验可以跨多个 Turn;每个已获准且需要记账的 Turn 仍以自己的 identity
完成 validation、writeback 和 settlement。等待实验不会延后欠付 debit、让
Turn 无限保持打开,或把此前工作变成免费。

最近的 capability/provider owner 应在启动前声明:

- 问题、candidate revision、比较对象和有界 postcondition;
- 已授权的 job identity、资源/时间上限和 recovery owner;
- 什么可观察结果会影响下一决策,以及如何验证;训练 checkpoint、评测或
构建只有在这份已声明的合同下才有决策意义;
- 已消费 result revision、下次观测时间/deadline、stop/cancel 行为;仅经过
一段时间或日志增长不是证据。

这些是 capability/provider 数据。Kernel 消费 identity、readiness、freshness
及现有 effect/settlement 事实,不解释 epoch、benchmark 分数、证明数量或进程
日志 prose。本提案不需要新增通用 job scheduler、worker launcher 或 built-in
capability。

拟议 handoff 如下:

| 阶段 | Owner 与必需行为 |
|---|---|
| 启动有界作业 | 普通已选 Todo/Turn 准入 effect。Provider 持久化 job/candidate identity 与 operational receipt。该 Turn 如实结算:启动成功只证明启动,不证明实验最终结果。 |
| 等待中的观测 | 现有 host/runtime 或 capability observer 无需 reasoning-model 调用即可检查精确 job。复用一个 due monitor 或已有 observation channel,不为每次 poll 创建 Todo。即使不消费 agent slot,也记录 observer CPU/IO 成本。 |
| 返回决策 checkpoint | Provider 按已声明的 postcondition 验证新结果,绑定 revision/artifact digest 并返回紧凑 observation。失败、取消和 deadline 到期同样必须形成可操作 observation。 |
| 准入决策 | 现有 scheduler/Turn 路径重新检查 Goal lifetime、ownership、quota、capability 与 authority。Ready observation 是输入,不是执行许可。等待此实验时仍可推进其他 runnable 工作。 |
| 消费并继续 | 在负责消费的可记账 Turn 中,通过普通 writeback 提交 result adoption 与下一步骤,再确认消费。重放/重复交付恢复原 receipt,不重复 launch 或 adoption。同一 Turn identity 的结算重放幂等;新获准 recovery Turn 仍正常结算。 |

当前 Todo resume condition 已有 `monitor_changed` 与 `resume_at`。若足够,
实现应复用 typed monitor generation 与 timeout fallback。
**缺失的是 provider-result binding 和 material-generation 更新集成**,不是
新增任意 `job_done` token。只有通过验证的决策 checkpoint 才能推进 generation。
Monitor 必须绑定精确 job/result source 与 Todo;无关日志变化不能释放等待。
Safety observation deadline 用来发现丢失通知或崩溃进程,不是周期性要求模型
重新解释未变工作。不支持的 host 保持现有执行并显示 readiness unavailable,
不能暗称已抑制模型唤醒。

Identity 必须跨 host/Agent restart 保留:Goal lifetime、owning Agent/Todo、
provider/run id、attempt 或 job generation、candidate revision、result revision
与 evidence/artifact ref。Raw log/native job handle 留在 provider,仅投影有界
且授权的引用。单个 PID 不是可恢复 identity。Provider 缺失或外部 effect 状态
不明确时,按现有 recovery 规则保留 unknown,不能据此自动 relaunch。只有
validated durable adoption 才推进 consumed cursor,因此交付后崩溃会重新提供
结果。Stopped/revoked Goal 和过期 claim 拦截 late result;替换 attempt 不能
消费旧结果。

Incumbent artifact reference 及其 qualifying result 与当前运行 candidate 分开
持久化。Provider 仅在可比有效结果满足已声明 constraint 后更新 incumbent。
Recovery 在复用前验证引用可用性与内容;candidate artifact 缺失不能静默
替换 incumbent,也不授权再次 launch。

先在一个真实 capability runtime 做 **shadow observation**:报告 unchanged poll、
decision checkpoint、result-to-decision latency、原生结果和模型/observer 总成本。
随后仅对该 provider/host 组合资格验证 opt-in model-admission suppression。
不改变通用 heartbeat prompt 或所有 scheduler 默认行为。关闭策略恢复 fixed
admission,保留 job/evidence history,不隐式杀作业,不改变 acceptance 或 accounting。

验收必须覆盖真实受控进程执行与 recovery,不只是 event 序列化:重复与乱序
结果、candidate 改变、restart、丢失通知、deadline、取消、unknown effect、竞争
claim 和 owner stop。Packaged frontend 旅程通过已有 task/capability surface 显示
绑定工作、等待原因、上次/下次观测、结果及 stale/unavailable 恢复;仅 CLI
readback 属于 partial slice。科学效用遵循
[研究决策](research-exploration-control-plane-v0.zh-CN.md#115-由证据驱动的实验决策修订提案)
的独立资格验证。

## 8. 测量模型

第一版实现应先测量,再控制。
Expand Down Expand Up @@ -597,6 +666,9 @@ executor、scheduler 或 generic policy framework。
- 用独立 acceptance rule 验证建议;
- Todo 与 replan 仍是唯一 execution authority。

- 在 shadow 中验证 §7.5 的 provider checkpoint/monitor binding,区分模型
决策、实验完成与 Turn settlement。

### M4:Authority-stride shadow recommendation

- 区分 report 与 authority-changing intervention;
Expand All @@ -610,6 +682,9 @@ executor、scheduler 或 generic policy framework。
- 对 pinned fixed profile 做 repeated comparison;
- 任何 improvement claim 都同时公开 limitation 与 failure mode。

- 长作业抑制 model admission 前,验证 §7.5 的 lost-result recovery 与
default-off parity;observer 成本继续纳入测量。

## 13. 验证标准

只有证明以下各项,研究计划才算合格:
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -655,6 +655,16 @@ layer at a time:
The study should estimate model- and work-class-specific response curves, not
one global tool-call or Todo-count threshold. Wider is not automatically better.

The proposed [long-running-work refinement](hierarchical-agent-stride-control-v0.md#75-long-running-work-and-decision-checkpoints-proposed-refinement)
tests **decision timing**, not deferred Turn accounting. First capture one
provider's validated result/checkpoint and bound monitor in shadow; only a
qualified opt-in arm may suppress unchanged model admission. Keep native
evaluation, job resources, deadlines and every owed Turn settlement identical.
Measure observer CPU/IO, model tokens/calls, result-to-decision latency, missed
results, duplicate execution/debits and native outcome per total budget. A
missing callback must recover through the declared observation deadline.
Neither a longer training run nor fewer bookkeeping calls alone qualifies it.

### 8.2 Evidence, stall detection, and semantic replan

The core hypothesis is that a durable coverage ledger plus semantic progress
Expand All @@ -668,6 +678,18 @@ whether direction changes create progress. DeepSWE validates whether the same
mechanism improves repository outcomes without reward-specific shortcuts. ALE
tests transfer to heterogeneous professional workflows.

The proposed [experiment decision path](research-exploration-control-plane-v0.md#115-evidence-conditioned-experiment-decisions-proposed-refinement)
qualifies result adoption separately from cadence. Hold session continuity,
planning, model and total budget fixed; compare delivered evidence alone with
evidence-linked continuation/successor adoption, then combine mechanisms only
after their independent comparisons. Include valid negative evidence,
inconclusive/conflicting measurements, failed prerequisites, justified repeats,
uncovered same-family probes and changed inputs reopening a retired scope.
Assess whether the executed next action follows the scoped result, alongside
native outcome and cost. Citation/schema compliance alone is insufficient;
forced pivots that suppress useful experiments count as failures. These are
proposed studies, not current C2/C4 evidence or permission to launch jobs.

### 8.3 Research exploration and composition

The [Research Exploration Control Plane RFC](./research-exploration-control-plane-v0.md)
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -552,6 +552,14 @@ non-material event mix;它不能证明 benchmark integrity,也不能复用
研究应估计 model/work-class-specific response curve,而不是一个全局 tool-call 或 Todo-count
threshold。更宽并不天然更好。

拟议[长作业修订](hierarchical-agent-stride-control-v0.zh-CN.md#75-长作业与决策-checkpoint修订提案)
测试**决策时机**,不延后 Turn 记账。先 shadow capture 一个 provider 的 validated
result/checkpoint 和绑定 monitor;只有通过资格验证的 opt-in arm 才能抑制
unchanged model admission。原生评测、job 资源、deadline 与所有欠付 Turn
settlement 保持一致。测量 observer CPU/IO、模型 token/call、result-to-decision
latency、漏结果、重复执行/debit 与总预算下原生结果。Callback 丢失必须按已声明
observation deadline 恢复。训练更久或 bookkeeping call 更少,单独都不够资格。

### 8.2 Evidence、空转检测与 semantic replan

核心假设是 durable coverage ledger 与 semantic progress observation 能阻止重复 maintenance
Expand All @@ -563,6 +571,15 @@ LHTB 是首要动态实验场,因为 partial reward 与 checkpoint 可以显
DeepSWE 验证相同机制能否改善 repository outcome,且不依赖 reward-specific shortcut。
ALE 验证它能否迁移到异构专业工作流。

拟议[实验决策路径](research-exploration-control-plane-v0.zh-CN.md#115-由证据驱动的实验决策修订提案)
把 result adoption 与 cadence 分别验证。固定 session continuity、planning、model
和总预算,对比仅交付 evidence 与 evidence-linked continuation/successor adoption,
独立比较成立后再组合机制。包含有效负证据、inconclusive/conflicting measurement、
失败前置条件、有依据复验、未覆盖的同 family probe 与 input 改变后重开 scope。
除原生结果和成本外,判断实际执行的 next action 是否遵循 scoped result。
Citation/schema compliance 单独不够;压制有用实验的强制 pivot 算失败。
这些是拟议 study,不是当前 C2/C4 evidence,也不授予 job launch 权限。

### 8.3 研究探索与组合

[研究型探索控制面 RFC](./research-exploration-control-plane-v0.zh-CN.md)定义 typed research
Expand Down
Loading
Loading