Skip to content

perf(goals): resolve checkpoint context in one TypeScript request - #5585

Open
hhyykk wants to merge 2 commits into
loopx-project:mainfrom
hhyykk:codex/checkpoint-source-reduction
Open

hhyykk wants to merge 2 commits into
loopx-project:mainfrom
hhyykk:codex/checkpoint-source-reduction

Conversation

@hhyykk

@hhyykk hhyykk commented Oct 4, 2026 •

Copy link
Copy Markdown
Contributor

Goal And Delivered Outcome

Checkpoint context reading and commit preflight currently return the complete authoritative facts to Python, then send them back to TypeScript for reduction. Compose the existing source owner and reducer in one private effect: each read/check now uses one checkpoint request instead of two, preserving the complete basis, receipt identity and final provider-fenced commit.

  • Source: the accepted TypeScript migration RFC, checkpoint transport boundary, T3/S2 with S7 performance evidence.
  • Intended base: main, merge base f35978e3a1289a146687346467b848bfca8fa2b0.
  • Expanded 40-pair comparisons pass the RFC §6 full-CLI p95 gate for all six measured workloads; the initial noisy observations and limitations are retained below. CI is still queued.

Author Declaration

  • Written by: model_agent — OpenAI Codex (GPT-6 family), with human-directed scope.
  • Implemented against: the RFC at the revision above, particularly the checkpoint transport boundary and §§5–6.
Criterion Disposition Implementation / evidence
Co-locate canonical source read and reduction implemented checkpoint_authority.ts::resolveCheckpointReadContext; real adapter read/check count test
Preserve complete basis and receipt/fence semantics implemented Complete small/3 MiB output parity on legacy/File/SQLite; stale head, concurrent provider write, uncertain commit and replay tests
Migration economics and packaged files implemented Receipt below; fresh wheel/sdist installs and deep semantic probes
No material full-CLI tail regression implemented 40 alternating pairs for each provider/size: every per-revision p95 improved; full results and pilot discrepancy below. Qualification is limited to these local workloads.
Bounded inspection manifest/pages deferred Existing RFC successor; full source remains receipt authority

Scope And Continuation

The existing goals checkpoint owner serves this internal boundary; no capability, provider or public API is added. Both private source/evaluate effects and their Python adapter calls retire. Provider selection remains legacy Markdown, File or SQLite. Reduction starts after the optimistic provider read fence releases, retaining the previous lock boundary; commit still rereads and validates inside its final fence.

Python retains local Markdown/history IO, source locks, settlement lookup and receipt persistence. Retiring those requires moving their real callers and durability contracts. The remaining RFC successor is versioned, bounded display paging; this PR does not claim all T3, durability promotion or long-running workload qualification. The adjacent refactor is the cohesive composition itself; no extra abstraction was needed.

Migration economics receipt

Field Before → after / evidence
Canonical owner Existing TypeScript authority reader + pure reducer, orchestrated by Python → the same owners composed inside TypeScript
Legacy code deleted No Python domain rule moved (already typed). Retires the 19-line private _source_facts orchestration function and both private source/evaluate RPC entries.
Bridge code added _evaluate becomes _resolve: 2 net adapter lines to supply runtime_root; no new bridge module.
Cross-runtime calls Public checkpoint-context read: 2→1; commit preflight/read-after-stale: 2→1 per attempt. Final commit and exact replay remain 1 each. Real effect spy covers read/check; provider fence/recovery suites cover commit/replay.
Product-code net change 20 added − 37 deleted = −17 LOC, excluding tests, generated manifest and docs.
Scaffolding No migration scaffolding added. One durable call-count regression and existing canonical-provider test strengthened through the public adapter. Temporary measurement harness is not shipped.
Facade exit Source-only facade removed now; remaining Python IO/locks/receipt facade described above.
Correctness/performance Full-output parity and packaging pass; expanded full-CLI distributions improve across all six workloads. Initial pilot variability and paired-sample tails are disclosed below.

On a 3 MiB prose fixture, total compact JSON payloads across checkpoint calls decrease from approximately 6.50 MB to 3.25 MB in each direction, for all three providers. These are serialized payload totals, not socket-byte or memory measurements; large payloads use the existing private-file transport.

Validation

  • Tested revision: 45ff752dadc1bc913808cff1e9cc6808c4838f48; earlier broader runs used the identical checkpoint implementation before rebasing onto f35978e3a.
  • Run state: finished (local validation; remote CI pending). Inputs: synthetic, public_fixture.
  • Host: macOS arm64, Python 3.13.7, Node 26.5.0, uv 0.8.17. No live Goal state or external model workload used.
Check kind Result Evidence / limitation
static passed npm run typecheck:control-plane, changed-Python Ruff, git diff --check, semantic advisory and full semantic inventory smoke. Method allowlist replacement is private runtime vocabulary, not a new shared state classification.
unit passed Four relevant TypeScript suites: 32 passed. Covers reducer, provider head, vision and lazy handler loading.
real_entrypoint / real_backend failed Rebased checkpoint context, provider fence and recovery suites: 56 passed, 1 existing failure. Broader pre-rebase checkpoint/isolation/effect-runtime/host recovery run: 145 passed, same 1 failure.
regression_parity passed The same recovery fixture fails on the original unmodified baseline (65 passed, 1 failed); #5573 already corrects its open-Todo setup. No unrelated fixture change included here.
regression_parity passed Complete basis equality on legacy/File/SQLite × small/3 MiB, restoring the same synthetic snapshot at the same path between implementations; only generated read_context_id normalized.
integration passed loopx canary premerge --from-git-diff --git-diff-base origin/main --format json: 18 executed, zero blocking failures, plus five direct checks. One inherited maintainability-ratchet advisory failure in unchanged goal_topic_runtime.py, reproduced on baseline.
real_entrypoint passed Fresh wheel and sdist installations each run packaged read/check, retain 3,040,000 prose bytes, reject changed-basis receipt and pass deep semantic readiness. Dashboard bundle rebuilt for packaging.
manual passed Independent code review found no blocker in identity binding, provider ownership, fence release or transport allowlists. Maintainer review is still required.
real_backend not_applicable No PostgreSQL adapter, provider routing or promotion change; checkpoint's existing local provider boundary remains legacy/File/SQLite.
regression_parity passed Uninstrumented full CLI against pinned baseline 1af7dbd43: 40 alternating pairs per workload, 480 invocations total. Initial 10-pair pilot retained separately. Separate effect probe measures call/payload counts.

Known failing test: test_checkpoint_only_recovery_bypasses_open_todo_completion_validation. This is a disclosed baseline failure, not an all-green test claim. Windows/Linux and sustained-production performance are untested locally; CI results are separate.

Runtime measurements

Isolated synthetic legacy source, Node 26.5.0; p50/p95 in milliseconds. Cold samples include managed daemon startup; warm measurements use the production adapter within a request scope and exclude process startup. The checkpoint row includes the real local-source effect and reducer, so it is not a pure-kernel benchmark.

Workload Samples per revision Baseline p50/p95 Candidate p50/p95
Cold managed ping 16 197.28 / 265.42 193.07 / 261.90
Warm typed ping 128 0.269 / 0.517 0.275 / 0.505
Warm small checkpoint read 128 2.408 / 7.083 1.169 / 2.158

Daemon RSS after loading checkpoint modules / after 128 reads: 99.64 / 103.47 MiB baseline, 99.95 / 104.08 MiB candidate. This bounded burst is not a sustained-memory qualification, and is separate from the full-CLI comparison below.

Full CLI comparison

40 pairs per provider/size; alternating BASE/HEAD order, one excluded warmup per revision, p50/nearest-rank p95 in ms. No timing instrumentation or outlier removal. Each workload uses one fixed synthetic basis at the same path; repeated reads replace only its read receipt UUID. Snapshot restore was used for the separate complete-output parity check, not between timed invocations. Baseline is pinned 1af7dbd43; rebasing the candidate brought independent Todo completion changes while checkpoint implementation hashes stayed unchanged.

Provider Workload Baseline p50/p95 Candidate p50/p95
Legacy Small 766.3 / 827.4 754.9 / 811.5
Legacy 3 MiB prose 1015.6 / 1075.7 944.6 / 1017.4
File Small 767.2 / 806.2 753.6 / 784.3
File 3 MiB prose 1024.1 / 1066.3 959.2 / 1031.0
SQLite Small 762.5 / 811.3 752.9 / 776.2
SQLite 3 MiB prose 1027.0 / 1094.3 952.4 / 1010.9

All six per-revision CLI p95 values improve. Large-workload medians improve by 6.3–7.3%. The original 10-pair pilot was mixed: legacy large-input p95 was 1136.3→1255.3 ms, and legacy small-input p50 was 895.6→1047.7 ms. That prompted this fixed 40-pair replication across all workloads; the pilot remains disclosed rather than discarded.

Individual pairs can still be slower: p95 of paired HEAD-minus-BASE deltas is +38.6/+23.9/+31.8 ms for small legacy/File/SQLite and −11.7/+10.8/−3.4 ms for large inputs. These paired-difference percentiles are distinct from comparing each revision's CLI p95. Local results do not establish sustained or cross-host performance.

Frontend / Visual Evidence

UI impact: none. Existing CLI read and preflight use the composed boundary; output/configuration/authority contracts are unchanged. No frontend or Lark operation is introduced or extended; visual evidence and attention review are not applicable. Source data for UI: none.

Shared-authority RFC fixture impact

Existing checkpoint fixture schema, complete prose, canonical Todo and full acceptance are retained. No semantic dimension or provider default changes. Real legacy/File/SQLite arms ran with disposable synthetic state; PostgreSQL promotion rehearsal is not applicable to this composition-only change.

Type / Area

  • Refactoring (no functional changes)
  • Documentation and test update
  • Control plane (goals / runtime)

Boundary Checklist

  • Diff and public evidence exclude private state, credentials, raw traces, internal links and local machine paths.
  • No maintainer-owned benchmark work duplicated; no benchmark/model jobs launched.
  • Scope matches the selected checkpoint request-consolidation slice.
  • UI impact is none.
  • Both commits carry DCO sign-off.

Control-plane change: left for maintainer review and merge.

hhyykk added 2 commits October 4, 2026 21:19
Signed-off-by: hyk <4408344+hhyykk@users.noreply.github.com>
Signed-off-by: hyk <4408344+hhyykk@users.noreply.github.com>
@hhyykk

hhyykk commented Oct 4, 2026

Copy link
Copy Markdown
Contributor Author

Validation for 45ff752 (base f35978e3a1289a146687346467b848bfca8fa2b0).

Changed surfaces: goals checkpoint read/preflight composition, private runtime method dispatch/large-payload allowlists, generated registry IO locations, focused tests and bilingual RFC checkpoint. The complete basis, receipt protocol, source/provider lock boundaries, final commit/replay and defaults retain their existing owners.

  • Rebased real checkpoint suites: 56 passed, 1 baseline failure. The same open-Todo recovery fixture fails before this change and is addressed by fix(goals): make checkpoint recovery hints match validation #5573. Broader earlier run of identical checkpoint code: 145 passed, same 1 failure. Four typed suites: 32 passed.
  • Typecheck, Ruff, full semantic inventory and whitespace checks passed. Canary: 18 executed, zero blocking failures, plus five direct checks; one inherited maintainability advisory in unchanged Lark runtime reproduced on baseline. Canary manual holds: none.
  • Real legacy/File/SQLite full-output parity passed at small and 3 MiB sizes. Fresh wheel/sdist installs passed complete read, receipt recheck, changed-basis rejection and deep semantic readiness. No production state or remote benchmark job was used.

Migration economics

Field Before → after / evidence
Canonical owner Existing TypeScript authority reader + pure reducer, orchestrated by Python → the same owners composed inside TypeScript
Legacy code deleted No Python domain rule moved (already typed). Retires the 19-line private _source_facts orchestration function and both private source/evaluate RPC entries.
Bridge code added _evaluate becomes _resolve: 2 net adapter lines to supply runtime_root; no new bridge module.
Cross-runtime calls Public checkpoint-context read: 2→1; commit preflight/read-after-stale: 2→1 per attempt. Final commit and exact replay remain 1 each. Real effect spy covers read/check; provider fence/recovery suites cover commit/replay.
Product-code net change 20 added − 37 deleted = −17 LOC, excluding tests, generated manifest and docs.
Scaffolding No migration scaffolding added. One durable call-count regression and existing canonical-provider test strengthened through the public adapter. Temporary measurement harness is not shipped.
Facade exit Source-only facade removed now; remaining Python IO/locks/receipt facade described above.
Correctness/performance Full-output parity and packaging pass; expanded full-CLI distributions improve across all six workloads. Initial pilot variability and paired-sample tails are disclosed below.

On a 3 MiB prose fixture, total compact JSON payloads across checkpoint calls decrease from approximately 6.50 MB to 3.25 MB in each direction, for all three providers. These are serialized payload totals, not socket-byte or memory measurements; large payloads use the existing private-file transport.

Full CLI evidence

40 pairs per provider/size; alternating BASE/HEAD order, one excluded warmup per revision, p50/nearest-rank p95 in ms. No timing instrumentation or outlier removal. Each workload uses one fixed synthetic basis at the same path; repeated reads replace only its read receipt UUID. Snapshot restore was used for the separate complete-output parity check, not between timed invocations. Baseline is pinned 1af7dbd43; rebasing the candidate brought independent Todo completion changes while checkpoint implementation hashes stayed unchanged.

Provider Workload Baseline p50/p95 Candidate p50/p95
Legacy Small 766.3 / 827.4 754.9 / 811.5
Legacy 3 MiB prose 1015.6 / 1075.7 944.6 / 1017.4
File Small 767.2 / 806.2 753.6 / 784.3
File 3 MiB prose 1024.1 / 1066.3 959.2 / 1031.0
SQLite Small 762.5 / 811.3 752.9 / 776.2
SQLite 3 MiB prose 1027.0 / 1094.3 952.4 / 1010.9

All six per-revision CLI p95 values improve. Large-workload medians improve by 6.3–7.3%. The original 10-pair pilot was mixed: legacy large-input p95 was 1136.3→1255.3 ms, and legacy small-input p50 was 895.6→1047.7 ms. That prompted this fixed 40-pair replication across all workloads; the pilot remains disclosed rather than discarded.

Individual pairs can still be slower: p95 of paired HEAD-minus-BASE deltas is +38.6/+23.9/+31.8 ms for small legacy/File/SQLite and −11.7/+10.8/−3.4 ms for large inputs. These paired-difference percentiles are distinct from comparing each revision's CLI p95. Local results do not establish sustained or cross-host performance.

Coverage limits: local macOS/Node 26.5.0 synthetic workloads; no sustained-host or PostgreSQL promotion claim. Cross-platform remote CI remains queued. Independent source review found no correctness blocker; maintainer review/merge remains pending.

@hhyykk
hhyykk marked this pull request as ready for review October 4, 2026 13:35
@hhyykk
hhyykk requested a review from huangruiteng as a code owner October 4, 2026 13:35

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant