perf: reduce VM, compilation, and garbage collection overhead - #81
Merged
Merged
Conversation
The weakref liveness test bypassed constructor validation with a plain string. Concurrent collections cleared that invalid target because strings cannot be weakly referenced. Use the checked constructor with a user-defined class and force a weakref-only sweep before asserting liveness. The original fixture fails deterministically with the added sweep. All four weakref unit tests pass with the supported target. Runtime behavior is unchanged.
owenthcarey
marked this pull request as ready for review
September 27, 2026 00:11
Replace the 66 per-increment performance reports and three compatibility notes with docs/PERFORMANCE-DISPATCH-MEMORY.md, and record a net startup comparison against the merge base.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
WeavePy spends substantial time and memory on dispatch, allocation, cache bookkeeping, compilation, and garbage collection. This PR reduces those costs through incremental runtime changes with regression coverage and paired measurements. It also fixes several Python/C API compatibility defects found during validation. It does not establish universal superiority over CPython.
Changes
Fixed weakref storage
Commit
15d6ba2compares with02339c7on Intel macOS and CPython 3.14.5. Twelve focused cases have seven initial cycles and eleven repeats, alternating engine order with discarded warmup and immutable warmed caches. Setup, assertions, and normal collection remain timed. Builds, tests, and profiles finish before timing; desktop activity is recorded.Repeated candidate/reference ratios follow; lower is better. CPU and peak RSS cover the whole process.
The initial unchanged 24-application suite is roughly flat against the preceding runtime: JIT work/process elapsed/CPU/RSS geometric means are 1.002/0.999/1.004/0.995. Work excludes the startup-only fixture. Twelve selected applications repeated eleven times give 1.009/1.007/1.015/0.995; these are not repeated full-suite means.
Costs remain explicit. Long dictionary-work diagnostics are 1.063/1.037 in JIT/interpreter mode, subclass calls 1.031/1.034, and repeated no-site JIT startup elapsed/CPU 1.032/1.040. Large weakref work remains 16.04 times CPython and peak RSS 3.76 times. Full-suite JIT work/process elapsed/CPU/RSS against CPython are 1.023/1.384/1.239/1.552. Different increments' gains aren't additive.
The consolidated report (
docs/PERFORMANCE-DISPATCH-MEMORY.md) summarizes every increment's changes, representative results, costs, rejected experiments, compatibility fixes, and validation. The per-increment reports are preserved in this PR's history at7416faa.CI follow-up: startup JIT admission
Commit
7416faabacks out blanket startup compilation deferral while preserving single site initialization, nested scopes, and later import budgets. It restores compilation of naturally hot startup code. No artificial warmup, fixture, baseline, threshold, or retry change is used.Against
15d6ba2, fifteen paired local cycles give cold sumvm/nested-loops/jitloop workload ratios of 0.856/0.908/0.932 and warm ratios of 0.985/0.998/1.001. This has explicit costs: repeated ordinary startup elapsed/CPU/RSS is 1.123/1.175/1.201, and full-suite JIT work/elapsed/CPU/RSS is 0.970/1.034/1.045/1.055. Repeated PyAES interpreter work is 1.097. This rollback resolves an admission-policy tradeoff; it does not make the compiler faster. Those costs are relative to the removed deferral, not tomain. A 31-cycle pairedtools/bench_startup.pycomparison of7416faaagainst the merge base shows no net startup regression: JIT elapsed/CPU/RSS is 0.976/0.994/0.971 for-c pass, 0.996/0.987/0.981 withoutsite, and 0.925/0.923/0.964 for the standard-library import probe.Validation and readiness
The startup follow-up passes 414 VM tests, nine harness tests, 34 C API unit tests, direct weakref C API integration, explicit 1 MiB embedding, Clippy, no-default compilation, changed-file formatting, and the repository artifact check. Its frozen CLI passes 112 regression runs, ten additional GIL-disabled runs, 476 probe shapes, eight observer runs, and twelve native SQLAlchemy/Alembic checks.
The preceding fixed-storage increment's upstream validation is qualified: all 60 selected group commands exit successfully, but the strict log audit detects an ignored GIL-disabled
SET_FUNCTION_ATTRIBUTE on a shared functionexception. An identical-bytecode oracle reproduces the GC-snapshot/function-construction race three of three times on both weakref reference and candidate. A fresh check of the archived original PR baseline (82f7237) reproduces it two of three times, versus three of three on current head7416faa; CPython passes. This establishes that the failure predates the PR, without claiming its frequency is unchanged. The report preserves this failure, along with existing weakref reflection, callback-attribute, and watched temporary-method differences.All 28 checks in CI for exact head
7416faapassed, including all three benchmark gates, all nine ecosystem shards, platform Rust tests, blocking regressions, GIL-disabled lanes, distribution checks, formatting, Clippy, MSRV, the artifact policy, and conformance reporting.Initial benchmark suite geometric means against the merge base are 0.934 on Linux, 0.916 on Windows, and 0.953 on ARM macOS. Windows sumvm is 0.981, resolving the previous retained 1.219 ratio. On macOS, DeltaBlue moves from 1.227 initially to 0.983 on the gate's standard retry; JSON moves from 1.199 to 1.008. Both batches remain in the CI log.
CI benchmark workloads, baselines, and thresholds are unchanged. Raw measurements, profiles, frozen binaries, source snapshots, and held experiments remain under
target/performance/; only reusable tools, regressions, and one concise report are committed.