perf: close the gap with CPython across the benchmark suite - #82
Draft
owenthcarey wants to merge 65 commits into
Draft
owenthcarey wants to merge 65 commits into
owenthcarey wants to merge 65 commits into
Conversation
The VM's Rc was an alias for std::sync::Arc, so every clone and drop of a heap object was a locked read-modify-write. Wrap Arc and Weak in newtypes whose strong-count updates are plain loads and stores until a second thread registers with the GIL, free-threading starts, or a Python thread is spawned. Deallocation still goes through Arc's own drop. Shared strings, bytes, and tuples use the same biased updates. Debug builds keep atomics because the unit-test binary runs interpreters on concurrent threads.
- Port CPython's `_abc` accelerator and use CPython's own `abc.py`. Registration and ABC `isinstance` checks no longer run through Python-level `WeakSet`s, which also takes `import site` from 25 ms to 15 ms. The `io` types drop a private `register` that shadowed `ABCMeta.register`. - Port CPython's list sort (timsort with the powersort merge policy) with its type-specialized comparisons for `str`, `int`, `float`, and tuple keys. Sorting a list that mixes NaNs with other numbers no longer panics, and inconsistent comparisons behave as in CPython. - Use mimalloc as the process allocator and remove the thread-cache front end over the system allocator. - Run container subscripts, subscript stores, and comprehension appends in the core loop, out of line so the rest of the loop keeps its register allocation. Inserts into dicts of any size stay native when the probe proves no key of another kind could compare equal, and report the mutation to dict watchers and the builtins rare-event counter. - Memoize whether a type overrides `__eq__` or `__hash__` per type version instead of looking the names up for every dict operation. - Store suspended generator frames as `Box<Frame>` rather than type-erased boxes that were downcast on every resume. - Deliver work queued by a dropped value, such as an unclosed file's `ResourceWarning`, at the instruction that dropped it.
…e loop - Step dict, key, value, and item iterators in the core loop, including their first step and exhaustion. An item unpacked by the following `UNPACK_SEQUENCE 2` goes straight to the stack, so no tuple is built. Iterating a 1,000-key dict's items is 8 times as fast. - Unpack exact tuples and lists of the expected length, and get iterators for dicts and dict views, without leaving the core loop. - Compare pairs of ints, floats, or strings natively before consulting any comparison protocol.
- Make `builtin_types()` return a leaked `&'static` registry instead of cloning an `Rc` through a thread-local `RefCell` at every call site. - Run `isinstance` in the core loop's in-place builtin arm and grade dropped type operands inline. - Serve the JIT's dict reads with the native `LeafProbe` lookup. - Retire compiled callees that keep calling back into the interpreter from native-to-native entries (deltablue's method pattern ran 4x slower compiled than interpreted). - Run `random.Random` on its state bytearray in place instead of copying and reallocating 2.5 KB per call (`random()` was 14x CPython). - Derive the core loop's cold per-activation state lazily, so calls, returns and helper handoffs reload less. - Park yielding generators and deliver clean returns directly, without the generic exit plumbing; pass the recursion depth cell through. - Construct built-in exceptions before the conversion-constructor chain, share code objects in traceback entries, and intern the slot keys every raise sets. - Allocate function `__dict__`s on first use and stop mirroring every disk-cached frozen module's code in memory. - Add an opt-in allocation-site profiler (`--features alloc-profile`). Also includes the getattr/hasattr miss precheck, the slot wrapper miss cache, single-pass exception class matching, and deferring JIT warm compiles during startup.
Compiled code called pure-leaf functions and methods (getters, predicates, small arithmetic) through a full native activation and then counted each one as a reason to retire the calling loop to the interpreter. Call sites now evaluate them with the interpreter's frameless pure-leaf evaluator right after the site's guards, so such loops stay native; the verdict is mirrored on the code's `JitHint` so other callees pay one relaxed load. A method with a callback-free scalar field update plan (`self.n += k; return self.n`) likewise runs the update directly after the method guard instead of through the native-call preflight. Instructions retired, JIT on: a `c.m()` getter loop -32%, a two-argument function call loop -25%, richards -20%, attr_access -4%.
- Keep up to three earlier resolutions in method slots and call slots, so call sites over several receiver classes stop re-resolving on every class change (a four-class `o.f()` site: -48% instructions), and widen the instance-attribute polymorphic cache to four classes. - Switch to `CALL_KW` callees inside the core loop, and cache each keyword site's validated binding instead of re-checking the names on every call; unpack bound-method callees in place like CPython's `CALL_BOUND_METHOD_EXACT_ARGS`. - Serve `dict.get`, `in`, and the generic dict lookup with the native `LeafProbe` for `str` and `int` keys, and format `str % scalars` in the leaf loop (dict_ops: -25% instructions, now at parity). - Raises: memoize whether an exception class has its own `__init__`, recognize an inert handled exception without the reap walks, and re-probe whether the frame object is observed right after `POP_EXCEPT` so the frame returns to the fast loop (raise loop: -22%). - Charge JIT-driven generator resumes like interpreter calls so such loops retire to the interpreter's cheaper inline resume (-21%), mirror the pure-leaf verdict on `JitHint`, and fill fresh locals in line.
A compiled self-recursive scalar frame (fib's shape) now calls itself with a native call: the callee's JitFrame and buffers live on the caller's stack frame, the activation is charged through small enter and exit helpers (GIL countdown, observers, recursion depth), and a deopt or raise finishes through a slow helper that rebuilds the callee's interpreter frame. The running compilation's layout rides on the call context so that rebuild never depends on the tier cache, which may retire the code mid-recursion. A call site whose callee is an already compiled, guard-free scalar leaf (spectral_norm's `_eval_a`) enters that leaf's code directly, falling back to the ordinary call helper when the enter helper declines or the leaf deopts (the leaf only computes, so it restarts from the top). Keyword calls no longer count toward retiring a native driver loop: the interpreter binds them through the same permutation, so tier 1 would not run them any cheaper (call_overhead's loop now stays compiled). fib: 445M -> 96M instructions; spectral_norm: 657M -> 228M.
- Inline-cache reads on x86_64 without AVX use the epoch snapshot's plain loads instead of portable-atomic's run-time dispatched 128-bit load (an out-of-line indirect call on every cache read). - Dicts keep a one-word summary of their `str` keys' hash classes, built once per key layout, so the per-call check that an instance dict doesn't shadow a method usually skips the probe. - A frame's small, untracked, uniquely held list, tuple or dict of atomic values (a `**kwargs` dict) is released by a plain drop instead of the prompt-reap cascade. - Keyword calls from native code bind into pooled vectors and run the callee as a lean activation (a pure leaf evaluates frameless), as the interpreter's own `CALL_KW` does, instead of a full framed run. - `datetime` natives read fields straight from the shared slot layout, shift with 64-bit arithmetic, and build exact `date`/`datetime` instances from int fields without running the Python `__new__`. call_overhead: 2280M -> 1805M instructions; spectral_norm, dict_ops, deltablue and datetime_ops improve 1-10%.
The `mimalloc` crate's `GlobalAlloc` sends every Rust allocation through `mi_malloc_aligned`, whose fast path needs the size class's free list to offer a suitably aligned block; otherwise it takes a generic path that can over-allocate (mimalloc v3 treats only power-of-two block sizes as naturally aligned). Every mimalloc block is word aligned, so a small allocator shim over `libmimalloc-sys` uses `mi_malloc`/`mi_zalloc`/ `mi_realloc` for alignments up to a word and the aligned entry points only above that. 1-4% fewer instructions on allocation-heavy fixtures, and float_math's peak RSS drops 1.5 MB.
The core loop now loads a `dict`, `set` or `str` receiver's site-cached native method itself (it already did for lists), and calls every leaf builtin kind whose body releases no reference beyond its operands (the `str` methods, `len`, `dict.get` and views, `list.append`/`pop`/ `insert`/`reverse`/`copy`, `set.pop`) through `leaf_builtin_call` right there. Each such call used to leave the loop twice, once for the method load and once for the call, and rerun the loop's prologue both times. About 800 fewer instructions per call: str_methods 955M -> 828M, call_overhead and datetime_ops improve as well.
The frameless leaf evaluator, which ran pure getters and predicates in place, now also covers: - Effect leaves: bodies whose only side effects are attribute stores. Stores are buffered as owned values (later loads of the same attribute read them back) and committed at the return, after every target is re-validated, so a bail-out anywhere leaves nothing behind and the ordinary call reruns the body. Setters get a dedicated shape. - Local variables (STORE_FAST and its fused forms). - Nested calls to pure leaves (methods and functions), evaluated in place up to a small depth while no store is buffered. - A code that keeps bailing (64 consecutive misses) loses its leaf verdict, so calls stop paying for the attempt. Pure method calls also take the interned names from the loaded constant table instead of re-deriving it per attribute check. `len()`, truth tests, and iteration over an instance whose class resolves the dunder to a native builtin call that builtin directly, and the core loop runs a native iterator's `__next__` in place. Deque iteration reads its slots by position. A setter call: 2001 -> 924 instructions; deltablue's `execute` kernel: 3976 -> 3313; deque iteration: 3278 -> 1562 per item.
- The JIT's attribute helpers read `__slots__` storage with an unguarded peek, as they already did for instance dicts. - An indexed store replacing an existing key's value no longer bumps the dict's stamp (the interpreter's store already used the unstamped and value-store accessors), so stamp-keyed caches survive it. - The callback-free field update (`self.n += k; return self.n`) also serves `__slots__` classes. - Slot keys are interned, so guards settle them by identity and no instance allocates its own copy of each slot name. attr_access: 922M -> 730M instructions.
Cross-crate inlining between the VM, compiler, and JIT crates makes the benchmark fixtures 1-10% faster in wall time (pickle_bench 10%, float_math 6%, deltablue 3%) for about a third more build time.
A dying list of plain objects no longer walks each element's fields: scalars and deferred-tracking instances can't anchor anything the cascade must visit. The scan also reuses the handles it finds instead of looking each child up twice, and hashes ids with the VM's id hasher.
An instance that assigns its attributes in its class's usual order keeps them as a plain vector of values whose names live once per class, like CPython's shared keys. Constructing such an instance allocates no hash table, and an attribute keeps the position a __dict__ entry would have, so the indexed attribute caches serve both layouts. Anything that needs the real dictionary (vars(), del, an out-of-order attribute, the C API) materializes it, and the instance keeps it from then on. A three-attribute instance drops from 490 to 249 bytes, and constructors run 15-20% fewer instructions (float_math: -11% instructions, -20% wall). The rarely set native-value slot becomes a pointer-sized write-once cell, so PyInstance grows by only one word.
A keyword call that skips a defaulted parameter (f(x, c=1) over def f(a, b=0, c=0)) used to fall back to the generic dynamic-call path, which re-verified the keyword names and bound a fresh frame every call. The JIT now marshals it like any keyword call and tags the skipped slots; the call helper fills them from the callee's current scalar defaults (or, for anything else, binds the call generically by name), so the compiled callee runs natively. Compiled scalar leaves are also preferred over the frameless evaluator when both apply. call_overhead's f(i, c=5) shape: 256 ns -> 52 ns per call (CPython: 99).
- Retire an exhausted range or list iterator in the core loop instead of handing the loop exit to the full handler, as the leaf arm does. An interpreted loop over a short list costs 36% fewer instructions. - Build list literals in the core loop, and release the last reference to an untracked list or dict of scalars there as a plain drop instead of taking the prompt-reap path. - Reclaim dead deferred containers at the tail of the deferral list as new ones arrive, so container churn neither grows the list nor parks freed allocations until a sweep. - Allocate 16-byte-aligned layouts through mimalloc's plain entry points: every block of 16 bytes or more already has that alignment, and the aligned entry points fell to their generic path for iterators.
- Recognize the process-wide NotImplemented and Ellipsis singletons by identity: they carry the type registry of the thread that built them, so a class check against the current thread's registry failed for every later thread (the native _abc raised AssertionError). - Restore compiling sustained start-up and import work: warm compiles during those phases count lean intervals again, rather than deferring every one, and frameless leaf calls count toward the warm-up, handing the call at which a compile falls due to the framed path. - Count direct self and leaf calls and frameless call-site evaluations in the native-call statistics, and scalar field updates' boxed returns on the method path, which the tests assert on. All 419 VM unit tests pass with the CI stack size.
- Fuse x.m(simple args) for a class's registered native method (a deque's append, popleft, pop, ...): the builtin runs on bitwise views of the borrowed receiver and arguments, with no method object, no receiver or argument references taken, and no operand stack traffic. - Serve a subscript whose class has a cached native fast __getitem__ in the core loop instead of handing it to the full arm. - Cache how each class answers a boolean context (native __bool__ or __len__, always true, or the full path) by attribute version, instead of looking both names up for every test. - Give deque iteration, integer indexing and appendleft guard-free fast paths, and keep an emptied deque's free prefix for the next appendleft. deque_ops: 2.58x -> 2.03x CPython in instructions; append/popleft pairs -29%, indexing -44%, truth tests -39%.
- Don't compile a loop-free function that makes a dynamic Python call: from native code the call takes the interpreter's generic path (a nested run rather than the inline activation an interpreted caller uses), and compiling was the largest cost a short-lived method paid. deltablue (default size) drops from 97 ms to 81 ms, with JIT compile time cut from 17 ms to about 10 ms. - Switch generator state in and out of its cell without guard atomics on inline resumes and yields (about 5% per resume/yield pair).
A LOAD_ATTR or STORE_ATTR site that proved a split-layout field records the class version and shared-key position, so later hits skip the inline cache decode and the key-name check. The leaf evaluator's operand values also get a primitive layout, so copying one is two words.
The frameless leaf evaluator now translates each pure or effect leaf once into register-addressed ops (locals and stack positions get fixed registers, jumps name their target op, and local loads, POP_TOP, COPY and SWAP mostly disappear) and runs them with one unchecked dispatch per operation. A path that reaches an instruction the plan can't express still declines only when it runs.
A constructor whose __init__ is a pure or effect leaf returning None now runs it through the leaf evaluator on the fresh instance, storing straight into it (nothing else can see it, so a decline just discards it). New-attribute stores take the split-layout shortcut by appending the class's next shared key, and an effect leaf's multi-store commit skips the latest-store scan when every store goes to self under a distinct name.
Interpreter call sites keep evaluating a pure or effect leaf frameless after the JIT compiles it (native callers still enter the compiled code). The collector's tracked-handle references use the biased Rc, its Bloom-filter inserts skip already-set bits, and its population counters update without locked read-modify-writes while the GIL serializes them. The generation-0 threshold now matches CPython 3.14's (2000).
The math module's functions register as leaf builtins over plain int, float and bool arguments, so math.sqrt(x) no longer drops to the generic call path; module attribute reads hit the site's cached index in the core loop's LOAD_ATTR arm; and a polymorphic method site now remembers each receiver class the class cache resolves for it, so its fused frameless leaf call keeps serving every class.
…ispatch len() of an instance whose class serves __len__ with a registered native leaf builtin calls it through a version-keyed class cache; the core loop's TO_BOOL and the JIT's truth helper answer instances through the native truth cache; and generic iteration (sum, list, ...) calls a registered native __next__ directly, with class misses cached too. The deque's slot probes use interned names.
…ut a lock The core loop's BINARY_OP serves str + str, small str * int, and str % over scalar and string arguments through an out-of-line helper instead of handing the instruction to the full leaf arms. Dict mutation stamps come from a plain load and store while the GIL serializes mutation (the locked increment stays for free-threaded mode and the debug test binary).
…ore loop A datetime-style instance's public field read, its binary operators and its comparisons now run from the core loop's LOAD_ATTR, BINARY_OP and COMPARE_OP arms (through out-of-line calls into the native bodies) instead of leaving the loop twice for the helper and the full leaf arms.
The tier-2 attribute get and set helpers try the common shape first (a scalar lane over an indexed instance field whose class is unchanged) before the general lane classification, and the class guard is always inlined.
… checks The core loop's fused LOAD_GLOBAL len; LOAD_FAST x; CALL 1 now also serves an instance whose class's __len__ is a registered native body (a deque's), skipping the separate push, call, and pop: len(q) on a deque drops from about 1110 to 620 instructions. The pickle decoder's first pass verified every instance's __slots__ names against the class MRO, once per instance. It now remembers the members it verified for the stream, so a list of slotted instances checks each name once (records loads: 4% fewer instructions).
tools/pgo_build.py builds an instrumented CLI, trains it on the benchmark fixtures at the harness's work sizes (JIT on and off), the bundled regression suite, and a stdlib import sweep, then rebuilds the release binary with the merged profile, as CPython's release builds do. On the measured macOS host this cuts wall time on most fixtures by 15% to 30% (attr_access 1.24x to 0.84x of CPython, call_overhead 1.10x to 0.87x, float_math 1.27x to 1.01x). Training at toy work sizes left large-integer multiplication cold and slowed pidigits by 38%, hence the harness sizes.
The frameless evaluator copied each plan op and decoded all of its operands before dispatching on it, keeping the copy's fields and the instruction index in stack slots: about 21 instructions per op. Matching the op in place lets each arm load only its own operands, and dispatch drops to 5 instructions. The evaluator's scratch buffers also stop calling their out-of-line drops when they are empty (the usual case). The collector's temporary id sets, its suspect map, and the descriptor registry hashed object addresses with SipHash. They now use the address mixer the collector's main index already uses (these tables never hold untrusted keys), and so does the prompt reaper's cascade marker set. Measured with callgrind: DeltaBlue retires 5.6% fewer instructions, and the pickle benchmark 1.9% fewer.
A forwarding wrapper's f(*args, **kwargs) compiles to BUILD_MAP 0, LOAD_FAST kwargs, DICT_MERGE and CALL_FUNCTION_EX: a fresh dict, tracked by the collector, filled by copying kwargs, only to be read by the callee's binder. When the callee is a Python function the core loop now binds straight from the local mapping and skips the copy; anything else runs the instructions as before, untouched. The keyword binder also writes the parameters into the activation's locals directly (it staged them in a temporary vector and copied them over), sizes a **kwargs dictionary up front, and releases the call's operands in place instead of collecting them first. A callable instance forwarding w(5, protocol=5) through __call__(self, *args, **kwargs) runs 30% fewer instructions. Also make the shared-string thread handoff test revoke the reference count bias before it spawns threads, as the VM does: under the biased counts its raw threads raced with the owner's plain increments, and the test failed depending on test order.
… once The deque natives resolved each of _data, _head, _maxlen and _state by a separate hinted name lookup, several per operation. SlotStorage now offers leading/leading_mut, which match the first N slots against interned names by thin pointer in one pass, and the deque's fast paths and iterator step work on that view. The core loop's container subscript no longer grades an instance operand it then declines (the native subscript grades it), shared instances skip the full drop grade, and native method calls stage operands straight into their slice. deque_ops: about 13% fewer instructions.
…preter A compiled function entered directly from the interpreter (try_call_native_direct) never had its interpreter round-trips counted: only framed entries and native-to-native callees did. deltablue's incremental_add ran every constraint.satisfy call through an activation shell and a framed interpreter call, several times the cost of the interpreter's inline call. Direct entries now feed the same callee round-trip backoff. deltablue (JIT on): 3% fewer instructions.
- The reload prologue no longer caches the globals and builtins handles (only LOAD_GLOBAL and LOAD_NAME read them, off the frame), so every call, return, resume and yield spills four fewer values. - VmExt caches the extension table's thin address, so code_vm_ext is one load instead of a fat dyn pointer and its alignment arithmetic. - An inline generator resume writes its slot's fields without drop glue, the yield parks the frame over a Running state without it, and the yield's consumer protocol runs without a QuietEntry round trip. Frame::push inlines. A for loop over a generator: 7% fewer instructions per item.
…l drop grade A pinned receiver or argument that other references still hold, with no weakref watching it, can't be at its dead line (past two owners only a weakref clone could put it there). gc_trace::drop_survives_plainly says so in two loads; the JIT's pin drain and the core loop's operand check take it before the prompt-reap and drop grades, which cost about 120 instructions per pin. The fused-pair bisection flag moves into the per-code table, off the frame switch. float_math (JIT on): 6% fewer instructions.
- The core loop's clean-return test counts each object's own slots instead of assuming every heap local aliases every other, so frames holding a few distinct objects skip the exit reap. - The exit reap's fast path untracks a dying collector-tracked list or dict whose contents can't finalize (deltablue's drained todo lists) instead of running the full prompt-reap cascade, and skips its alias count for objects held past every slot. - A leaf plan's nested pure-leaf call returns its result borrowed when the callee's scratch doesn't hold it (see LeafRet), so the caller neither clones nor later releases a getter's result. deltablue: about 6% fewer instructions.
Each dumps call grew a fresh memo table from empty (rehashing as it doubled) and a fresh output buffer (copying as it doubled). The encoder now takes the thread's spare table, emptied with its capacity kept, and reserves the previous call's output size. pickle_bench: about 3% faster.
- A local list, dict, set or str receiver of a site-cached leaf method (x.append(v), d.get(k), s.startswith(p)) is called on the borrowed local, as instance receivers already were, so no receiver clone is pushed and graded on release. - The unpickler builds plain instances with deferred collector tracking, as ordinary construction does; slot state tracks them only when a value isn't atomic (dict state already did). A decoded Point of scalars now dies by plain drop instead of the prompt-reap cascade. - The prompt-reap cascade allocates its per-node scratch once per cascade. pickle_bench: about 5% fewer instructions; list.append/pop loop: 10%.
The run fixture for weakrefs and gc still expected the old default collection threshold of 700; the branch adopted CPython 3.14's 2000 (the fixture's output now matches CPython's). The remaining changes are rustfmt and clippy fixes, with alignment and mutable-borrow allowances where the pointer casts are sound by construction.
The branch stopped compiling standard library functions during startup, so a program's first compile now pays the code generator's cold start (about 0.7 ms: its code paged in, its tables built) inside its hot loop; the Windows bench gate flagged sumvm at +18% against the merge base. The CLI now compiles and discards a small loop on a background thread when it starts a program file or module, and any other run does so when its first code object is halfway to the compile threshold. -c pass is unaffected (no extra memory). sumvm, jitloop, nested_loops: back within 1-3% of the merge base.
Resuming a generator switched activations (the frame made the running one, the pending-caller and recursion bookkeeping, the core loop's reload) and back at each yield: most of the cost of each item for a body that only moves scalars through locals. A new stepper (gen_fast) runs such a body directly on its frame to its next yield: locals, constants, int and float arithmetic and comparisons, jumps, and for loops over ranges, lists, tuples and other such generators. It never raises or runs Python code; anything else ends the step at an instruction boundary and the ordinary resume continues from there, so a fast step is always a prefix of the ordinary run. The core loop's FOR_ITER over a generator and the lean resume path (next(), and sum()/list() folding) take fast steps first. A nested generator whose step stops partway is parked having consumed its sent value (Frame::sent_consumed), which every resume path honors. generators: 28% faster; a for loop over a generator: 34% fewer instructions per item.
The decoder's push and the encoder's extend reserved capacity (a fallible call) before every append. They now reserve only when the buffer is full, and the encoder's frame writer inlines. pickle_bench: about 8% fewer instructions.
The eligibility scan now also proves every jump target, local index and constant index in range and that the last instruction can't fall through, so the step reads instructions, locals and constants without per-use checks. The scalar arithmetic and comparison helpers inline, which keeps the step's operand stack in registers. A for loop over a generator: 7% fewer instructions per item; sum() and list() over generator expressions: 5-9%.
The core loop's fused `x.m(<args>)` step now folds an argument like `i + 1` (checked add, subtract, multiply, and bitwise ops on two ints) in place, and a local instance whose method is a native slot calls it directly before the leaf-site lookup. The deque fixture drops 8.4% in instructions, and `q.append(i + 1)` drops 31%.
- A class whose __init__ takes trailing default parameters now takes the frameless constructor paths (and the lean framed one): the call fills the missing arguments from the function's defaults. - Leaf plans allocate empty list and dict literals, so an __init__ like `self.items = []` stays a leaf. - A leaf plan runs a pure-leaf callee as a frame of its own evaluation, saving only the caller's live registers, instead of re-entering the plan runner. - The GC's finalizer check on tracking uses the class's cached __del__ verdict rather than an MRO lookup per instance. Constructing a DeltaBlue Variable drops from 12.5K to 7.4K instructions, and building its constraint chain drops 7%.
A leaf plan's in-place callee frames now apply only to general plans. A certified tiny shape (a constant or field return, a field comparison) keeps its dedicated evaluator, as before the frames: it's at least as fast, and the unit test that counts cached field predicate hits on DeltaBlue passes again.
Slicing a tuple, bytes, bytearray, or a string with surrogates used to copy the entire sequence into a vector of objects and then select from it, so every slice cost the length of its source. `slice_seq` is now generic over the element type and copies a unit-step slice in one piece, and each type slices its own storage. `re.compile` with IGNORECASE slices a 64K charset map 256 times, which made `import logging` cost 1.19B instructions; it now costs 525M.
The private frozen-stdlib cache stored CPython marshal data, so every import unmarshalled it and transcoded the CPython bytecode back into WeavePy instructions. The new `native_code` module encodes a code object's fields directly (varints, delta-coded line tables, nested code inline) and decodes them with no transcoding; filenames are stamped at load. The source check hashes eight bytes at a time. `import json` drops from 269M to 205M instructions, `-c pass` from 106M to 91M, and `import dataclasses` from 454M to 349M.
Suspects live in two maps, active and dormant, so the sweep that runs at drop safe points walks only the entries with probe budget left; the dormant ones are visited on the stride, as before. During imports the old single map made every sweep walk up to 256 entries. `import logging` drops 3.5% and `import json` 7%.
- A class's scalar attribute read through an instance (`self.LIMIT`) is served from the site's stamp slot while the class version holds and the instance doesn't shadow the name, instead of resolving the name on every read. - The fused `x.attr` arm reads the site's field shortcut in line. - `if x is None`, `if x is not None`, and `if x:` on a local decide the branch from the local in place (no load, clone, or release), and a conditional jump steps over the `NOT_TAKEN` that follows it. - `POP_TOP` of a scalar no longer calls into the drop glue. A class scalar read through an instance drops from 730 to 280 instructions, and an `is None` branch by a fifth.
`STORE_GLOBAL` and module-level `STORE_NAME` of an existing name used to stamp the module dict like any mutation, so every cached global load in the module missed after each store. Rebinding now replaces the value in place: the key layout, and so the stamp the global-load caches check, stays put. The JIT's burned-in globals, which remember values rather than positions, also watch a process-wide epoch that each in-place rebinding advances. The core loop runs `STORE_GLOBAL` itself, as it already did module-level `STORE_NAME`. A global store drops from 935 to 320 instructions, on par with CPython, and a call to a function that updates a global counter from 3.7K to 1.8K.
Record the paired wall-time checkpoint for the benchmark fixtures, the startup and import costs, the per-operation gaps that remain, and the method behind each number.
A fast step stops at a back edge when the GIL countdown runs out. For an inner generator that left its frame partway, and `gen_fast_next` then declined it for good, so an outer generator fell back to the general loop for the rest of its run. Fast steps now continue a partial frame where it stopped, and a draining consumer's send services the eval breaker itself (the countdown reset, the dormant suspect probe, the GIL checkpoint) and keeps stepping unless the loop generation moved. Summing a generator over a generator drops 6% to 8%.
The core loop's fused `xs.append(x)` and `xs.pop()` on a local list now push and pop directly instead of admitting the call and then entering the generic builtin, about 8% off each call. `list.pop` also accepts a bool index, as CPython's does (`[1, 2, 3].pop(True)` returns 2; it raised TypeError).
CI's stable toolchain moved to 1.98.1, whose rustfmt orders the `initconfig` imports differently and whose clippy asks for `as_chunks` in the frozen-cache hash and `usize::midpoint` in the binary insertion sort. Both are stable well before the 1.93 MSRV.
owenthcarey
force-pushed
the
perf/beat-cpython
branch
from
September 30, 2026 18:09
f31f2ef to
9e7ed7f
Compare
…e GIL
Under `-X gil=0`, two threads putting into the same
`queue.SimpleQueue` could both see a parked getter's `_locked` flag,
both clear it, and both release the gate; the second release raised
`RuntimeError('release unlocked lock')`. `multiprocessing.dummy`'s
thread pool returns results that way, so the free-threaded lane's
`test_multiprocessing_dummy` failed intermittently (about 2% of runs
on main, 5% here).
The flag's test, clear, and release now happen under one lock, as
CPython's critical section makes them. The append stays outside it,
since it may collect and run Python code. The test passes 400 of 400
stressed runs.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
This branch makes WeavePy faster than CPython 3.14 on 19 of the 23 benchmark fixtures (geometric mean 0.59x CPython's time) and cuts import costs by 20% to 60%. Four fixtures still trail CPython; closing them needs an interpreter redesign (frames and calls, inline caches, object layout), which is left for a follow-up PR.
The largest changes:
**kwargshelpers, bodies that call read-only builtins such aslen,isinstance,mathfunctions,dict.get, andstrmethods) run from a pre-translated register plan without an activation. Method call sites remember their verified callee per class version.*args,**kwargs, keyword-only,f(*args),f(*args, **kwargs), and callable-instance calls run as inline activations in the core loop instead of leaving it for the generic binder.sum(),list(), and other draining consumers fold a generator's yields in place.re.compile(..., re.I)and soimport loggingpathological), and the collector's per-drop suspect sweep skips dormant entries.tools/pgo_build.pybuilds a PGO release, as CPython's release builds are. It trains on the benchmark fixtures at harness work sizes, the bundled regression suite, and a stdlib import sweep.Two correctness fixes came out of the work: a
sum()fold that could claim a nested generator's yields, and delayed finalizers after spread calls. The interned empty tuple now also holds for an empty*args.Current standing
Wall time relative to CPython 3.14 (lower is better) for a profile-guided release build (
tools/pgo_build.py) on the macOS x86-64 development host, JIT on: the median of five paired, interleaved cycles per fixture at the harness work sizes (geometric mean 0.592).deltabluedeque_opspickle_benchgeneratorsdatetime_opsstr_methodsfloat_mathdict_opscall_overheadfannkuchrichardsnbodyattr_accesslist_opsjson_benchpidigitspyaesjitkernelsfibspectral_normnested_loopsjitloopsumvmFour fixtures still trail CPython (see Follow-up).
Imports
Retired instructions for a whole process, WeavePy (this branch, warm cache) against CPython 3.14:
-c passimport jsonimport loggingimport asyncioimport unittestStartup is faster than CPython's; importing large stdlib modules still costs about 2x, and peak memory after those imports is about 2x CPython's (mostly compiled code objects).
Validation
RUST_MIN_STACK=8388608 cargo test -p weavepy-vm --lib --all-features).test_extcall,test_call,test_keywordonlyarg,test_positional_only_arg,test_pickle,test_functools,test_inspect,test_generators,test_gc,test_descr,test_dataclasses,test_typing) pass.Follow-up
See CPython parity performance for the method and the per-operation gaps. The remaining work, for a separate PR:
deltablue: the plan interpreter and the core loop's per-instruction overhead (spilled loop state), spread across method calls, attribute access and frame switches.generators: each resume and yield pair still costs several hundred instructions of activation bookkeeping plus two frame-switch prologues.deque_ops: the deque is a Python class with native methods; its method dispatch and the core loop's per-instruction overhead.pickle_bench: the decoder's two passes and the collector bookkeeping for decoded graphs.