Skip to content

perf: reduce VM, compilation, and garbage collection overhead - #81

Merged
owenthcarey merged 104 commits into
mainfrom
codex/performance-largest-gaps
Sep 27, 2026
Merged

owenthcarey merged 104 commits into
mainfrom
codex/performance-largest-gaps

Conversation

@owenthcarey

@owenthcarey owenthcarey commented Sep 23, 2026 •

Copy link
Copy Markdown
Contributor

WeavePy spends substantial time and memory on dispatch, allocation, cache bookkeeping, compilation, and garbage collection. This PR reduces those costs through incremental runtime changes with regression coverage and paired measurements. It also fixes several Python/C API compatibility defects found during validation. It does not establish universal superiority over CPython.

Changes

  • Reduce weakref allocations with shared call/repr methods, fewer captured helpers, and fixed getter/callback storage for exact 64-bit refs. Proxies and subclasses retain their dictionary representation. Callback ownership remains visible to GC. The Python C API shim now clears all watched wrappers without accidentally publishing callbacks or missing proxies.
  • Use guarded dispatch for saved native methods and zero-argument calls. Preserve native function identity, receiver ownership, observer fallback, errors, and weak cache ownership.
  • Use constant-time LRU recency updates for admitted scalar, tuple, and typed keys, with callback fallback and a repaired GIL-disabled tuple-cache race. Improve native set deletion and removal dispatch.
  • Reduce GC bookkeeping, tracked-object and instance storage, repeated attribute/field lookup, generator dispatch, parser/AST construction, and pickle bookkeeping. Reclaim retired JIT code and worker mappings while preserving embedding stack bounds.
  • Avoid repeated site/locale initialization and unnecessary compilation during later imports. Restore natural hot-code admission during startup after the original deferral regressed first-use Windows JIT work. Fix C attribute descriptor handling, native lazy-iterator C slots, and Union class subscription binding.

Fixed weakref storage

Commit 15d6ba2 compares with 02339c7 on Intel macOS and CPython 3.14.5. Twelve focused cases have seven initial cycles and eleven repeats, alternating engine order with discarded warmup and immutable warmed caches. Setup, assertions, and normal collection remain timed. Builds, tests, and profiles finish before timing; desktop activity is recorded.

Repeated candidate/reference ratios follow; lower is better. CPU and peak RSS cover the whole process.

Workload, 100,000 items JIT work Interpreter work JIT CPU JIT peak RSS
Weakrefs 0.886 0.916 0.892 0.786
Callback-bearing weakrefs 0.874 0.875 0.875 0.840
Watched cycles 0.872 0.894 0.877 0.789
Weak-set cleanup 0.868 0.854 0.871 0.809

The initial unchanged 24-application suite is roughly flat against the preceding runtime: JIT work/process elapsed/CPU/RSS geometric means are 1.002/0.999/1.004/0.995. Work excludes the startup-only fixture. Twelve selected applications repeated eleven times give 1.009/1.007/1.015/0.995; these are not repeated full-suite means.

Costs remain explicit. Long dictionary-work diagnostics are 1.063/1.037 in JIT/interpreter mode, subclass calls 1.031/1.034, and repeated no-site JIT startup elapsed/CPU 1.032/1.040. Large weakref work remains 16.04 times CPython and peak RSS 3.76 times. Full-suite JIT work/process elapsed/CPU/RSS against CPython are 1.023/1.384/1.239/1.552. Different increments' gains aren't additive.

The consolidated report (docs/PERFORMANCE-DISPATCH-MEMORY.md) summarizes every increment's changes, representative results, costs, rejected experiments, compatibility fixes, and validation. The per-increment reports are preserved in this PR's history at 7416faa.

CI follow-up: startup JIT admission

Commit 7416faa backs out blanket startup compilation deferral while preserving single site initialization, nested scopes, and later import budgets. It restores compilation of naturally hot startup code. No artificial warmup, fixture, baseline, threshold, or retry change is used.

Against 15d6ba2, fifteen paired local cycles give cold sumvm/nested-loops/jitloop workload ratios of 0.856/0.908/0.932 and warm ratios of 0.985/0.998/1.001. This has explicit costs: repeated ordinary startup elapsed/CPU/RSS is 1.123/1.175/1.201, and full-suite JIT work/elapsed/CPU/RSS is 0.970/1.034/1.045/1.055. Repeated PyAES interpreter work is 1.097. This rollback resolves an admission-policy tradeoff; it does not make the compiler faster. Those costs are relative to the removed deferral, not to main. A 31-cycle paired tools/bench_startup.py comparison of 7416faa against the merge base shows no net startup regression: JIT elapsed/CPU/RSS is 0.976/0.994/0.971 for -c pass, 0.996/0.987/0.981 without site, and 0.925/0.923/0.964 for the standard-library import probe.

Validation and readiness

The startup follow-up passes 414 VM tests, nine harness tests, 34 C API unit tests, direct weakref C API integration, explicit 1 MiB embedding, Clippy, no-default compilation, changed-file formatting, and the repository artifact check. Its frozen CLI passes 112 regression runs, ten additional GIL-disabled runs, 476 probe shapes, eight observer runs, and twelve native SQLAlchemy/Alembic checks.

The preceding fixed-storage increment's upstream validation is qualified: all 60 selected group commands exit successfully, but the strict log audit detects an ignored GIL-disabled SET_FUNCTION_ATTRIBUTE on a shared function exception. An identical-bytecode oracle reproduces the GC-snapshot/function-construction race three of three times on both weakref reference and candidate. A fresh check of the archived original PR baseline (82f7237) reproduces it two of three times, versus three of three on current head 7416faa; CPython passes. This establishes that the failure predates the PR, without claiming its frequency is unchanged. The report preserves this failure, along with existing weakref reflection, callback-attribute, and watched temporary-method differences.

All 28 checks in CI for exact head 7416faa passed, including all three benchmark gates, all nine ecosystem shards, platform Rust tests, blocking regressions, GIL-disabled lanes, distribution checks, formatting, Clippy, MSRV, the artifact policy, and conformance reporting.

Initial benchmark suite geometric means against the merge base are 0.934 on Linux, 0.916 on Windows, and 0.953 on ARM macOS. Windows sumvm is 0.981, resolving the previous retained 1.219 ratio. On macOS, DeltaBlue moves from 1.227 initially to 0.983 on the gate's standard retry; JSON moves from 1.199 to 1.008. Both batches remain in the CI log.

CI benchmark workloads, baselines, and thresholds are unchanged. Raw measurements, profiles, frozen binaries, source snapshots, and held experiments remain under target/performance/; only reusable tools, regressions, and one concise report are committed.

@owenthcarey owenthcarey changed the title Improve VM call throughput and measure remaining performance gaps Reduce VM call and native pickle overhead Sep 23, 2026
@owenthcarey owenthcarey changed the title Reduce VM call and native pickle overhead Reduce VM call and pickle overhead, including dynamic modules Sep 23, 2026
@owenthcarey owenthcarey changed the title Reduce VM call and pickle overhead, including dynamic modules Reduce VM call, attribute-store, and pickle overhead Sep 23, 2026
@owenthcarey owenthcarey changed the title Reduce VM call, attribute-store, and pickle overhead Reduce VM call, attribute, and pickle overhead Sep 23, 2026
@owenthcarey owenthcarey changed the title Reduce VM call, attribute, and pickle overhead Reduce startup, VM call, attribute, and pickle overhead Sep 23, 2026
@owenthcarey owenthcarey changed the title Reduce compilation, startup, VM call, attribute, and pickle overhead Reduce VM, compilation, startup, and garbage collection overhead Sep 26, 2026
@owenthcarey
owenthcarey marked this pull request as ready for review September 27, 2026 00:11
@owenthcarey owenthcarey changed the title Reduce VM, compilation, startup, and garbage collection overhead perf: reduce VM, compilation, and garbage collection overhead Sep 27, 2026
Replace the 66 per-increment performance reports and three compatibility
notes with docs/PERFORMANCE-DISPATCH-MEMORY.md, and record a net startup
comparison against the merge base.
@owenthcarey
owenthcarey merged commit 03d48b0 into main Sep 27, 2026
28 checks passed
@owenthcarey
owenthcarey deleted the codex/performance-largest-gaps branch September 27, 2026 18:02
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant