Conversation
Signed-off-by: hyk <4408344+hhyykk@users.noreply.github.com>
Signed-off-by: hyk <4408344+hhyykk@users.noreply.github.com>
|
Validation for 45ff752 (base Changed surfaces: goals checkpoint read/preflight composition, private runtime method dispatch/large-payload allowlists, generated registry IO locations, focused tests and bilingual RFC checkpoint. The complete basis, receipt protocol, source/provider lock boundaries, final commit/replay and defaults retain their existing owners.
Migration economics
On a 3 MiB prose fixture, total compact JSON payloads across checkpoint calls decrease from approximately 6.50 MB to 3.25 MB in each direction, for all three providers. These are serialized payload totals, not socket-byte or memory measurements; large payloads use the existing private-file transport. Full CLI evidence40 pairs per provider/size; alternating BASE/HEAD order, one excluded warmup per revision, p50/nearest-rank p95 in ms. No timing instrumentation or outlier removal. Each workload uses one fixed synthetic basis at the same path; repeated reads replace only its read receipt UUID. Snapshot restore was used for the separate complete-output parity check, not between timed invocations. Baseline is pinned
All six per-revision CLI p95 values improve. Large-workload medians improve by 6.3–7.3%. The original 10-pair pilot was mixed: legacy large-input p95 was 1136.3→1255.3 ms, and legacy small-input p50 was 895.6→1047.7 ms. That prompted this fixed 40-pair replication across all workloads; the pilot remains disclosed rather than discarded. Individual pairs can still be slower: p95 of paired HEAD-minus-BASE deltas is +38.6/+23.9/+31.8 ms for small legacy/File/SQLite and −11.7/+10.8/−3.4 ms for large inputs. These paired-difference percentiles are distinct from comparing each revision's CLI p95. Local results do not establish sustained or cross-host performance. Coverage limits: local macOS/Node 26.5.0 synthetic workloads; no sustained-host or PostgreSQL promotion claim. Cross-platform remote CI remains queued. Independent source review found no correctness blocker; maintainer review/merge remains pending. |
Goal And Delivered Outcome
Checkpoint context reading and commit preflight currently return the complete authoritative facts to Python, then send them back to TypeScript for reduction. Compose the existing source owner and reducer in one private effect: each read/check now uses one checkpoint request instead of two, preserving the complete basis, receipt identity and final provider-fenced commit.
main, merge basef35978e3a1289a146687346467b848bfca8fa2b0.Author Declaration
model_agent— OpenAI Codex (GPT-6 family), with human-directed scope.checkpoint_authority.ts::resolveCheckpointReadContext; real adapter read/check count testScope And Continuation
The existing goals checkpoint owner serves this internal boundary; no capability, provider or public API is added. Both private source/evaluate effects and their Python adapter calls retire. Provider selection remains legacy Markdown, File or SQLite. Reduction starts after the optimistic provider read fence releases, retaining the previous lock boundary; commit still rereads and validates inside its final fence.
Python retains local Markdown/history IO, source locks, settlement lookup and receipt persistence. Retiring those requires moving their real callers and durability contracts. The remaining RFC successor is versioned, bounded display paging; this PR does not claim all T3, durability promotion or long-running workload qualification. The adjacent refactor is the cohesive composition itself; no extra abstraction was needed.
Migration economics receipt
_source_factsorchestration function and both private source/evaluate RPC entries._evaluatebecomes_resolve: 2 net adapter lines to supplyruntime_root; no new bridge module.checkpoint-contextread: 2→1; commit preflight/read-after-stale: 2→1 per attempt. Final commit and exact replay remain 1 each. Real effect spy covers read/check; provider fence/recovery suites cover commit/replay.On a 3 MiB prose fixture, total compact JSON payloads across checkpoint calls decrease from approximately 6.50 MB to 3.25 MB in each direction, for all three providers. These are serialized payload totals, not socket-byte or memory measurements; large payloads use the existing private-file transport.
Validation
45ff752dadc1bc913808cff1e9cc6808c4838f48; earlier broader runs used the identical checkpoint implementation before rebasing ontof35978e3a.finished(local validation; remote CI pending). Inputs:synthetic,public_fixture.npm run typecheck:control-plane, changed-Python Ruff,git diff --check, semantic advisory and full semantic inventory smoke. Method allowlist replacement is private runtime vocabulary, not a new shared state classification.read_context_idnormalized.loopx canary premerge --from-git-diff --git-diff-base origin/main --format json: 18 executed, zero blocking failures, plus five direct checks. One inherited maintainability-ratchet advisory failure in unchangedgoal_topic_runtime.py, reproduced on baseline.1af7dbd43: 40 alternating pairs per workload, 480 invocations total. Initial 10-pair pilot retained separately. Separate effect probe measures call/payload counts.Known failing test:
test_checkpoint_only_recovery_bypasses_open_todo_completion_validation. This is a disclosed baseline failure, not an all-green test claim. Windows/Linux and sustained-production performance are untested locally; CI results are separate.Runtime measurements
Isolated synthetic legacy source, Node 26.5.0; p50/p95 in milliseconds. Cold samples include managed daemon startup; warm measurements use the production adapter within a request scope and exclude process startup. The checkpoint row includes the real local-source effect and reducer, so it is not a pure-kernel benchmark.
Daemon RSS after loading checkpoint modules / after 128 reads: 99.64 / 103.47 MiB baseline, 99.95 / 104.08 MiB candidate. This bounded burst is not a sustained-memory qualification, and is separate from the full-CLI comparison below.
Full CLI comparison
40 pairs per provider/size; alternating BASE/HEAD order, one excluded warmup per revision, p50/nearest-rank p95 in ms. No timing instrumentation or outlier removal. Each workload uses one fixed synthetic basis at the same path; repeated reads replace only its read receipt UUID. Snapshot restore was used for the separate complete-output parity check, not between timed invocations. Baseline is pinned
1af7dbd43; rebasing the candidate brought independent Todo completion changes while checkpoint implementation hashes stayed unchanged.All six per-revision CLI p95 values improve. Large-workload medians improve by 6.3–7.3%. The original 10-pair pilot was mixed: legacy large-input p95 was 1136.3→1255.3 ms, and legacy small-input p50 was 895.6→1047.7 ms. That prompted this fixed 40-pair replication across all workloads; the pilot remains disclosed rather than discarded.
Individual pairs can still be slower: p95 of paired HEAD-minus-BASE deltas is +38.6/+23.9/+31.8 ms for small legacy/File/SQLite and −11.7/+10.8/−3.4 ms for large inputs. These paired-difference percentiles are distinct from comparing each revision's CLI p95. Local results do not establish sustained or cross-host performance.
Frontend / Visual Evidence
UI impact: none. Existing CLI read and preflight use the composed boundary; output/configuration/authority contracts are unchanged. No frontend or Lark operation is introduced or extended; visual evidence and attention review are not applicable. Source data for UI: none.
Shared-authority RFC fixture impact
Existing checkpoint fixture schema, complete prose, canonical Todo and full acceptance are retained. No semantic dimension or provider default changes. Real legacy/File/SQLite arms ran with disposable synthetic state; PostgreSQL promotion rehearsal is not applicable to this composition-only change.
Type / Area
Boundary Checklist
Control-plane change: left for maintainer review and merge.