Skip to content

perf: close the gap with CPython across the benchmark suite - #82

Draft
owenthcarey wants to merge 65 commits into
mainfrom
perf/beat-cpython
Draft

owenthcarey wants to merge 65 commits into
mainfrom
perf/beat-cpython

Conversation

@owenthcarey

@owenthcarey owenthcarey commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor

Summary

This branch makes WeavePy faster than CPython 3.14 on 19 of the 23 benchmark fixtures (geometric mean 0.59x CPython's time) and cuts import costs by 20% to 60%. Four fixtures still trail CPython; closing them needs an interpreter redesign (frames and calls, inline caches, object layout), which is left for a follow-up PR.

The largest changes:

  • Frameless leaf evaluation. Small functions and methods whose bodies can't run Python code (getters, predicates, setters, **kwargs helpers, bodies that call read-only builtins such as len, isinstance, math functions, dict.get, and str methods) run from a pre-translated register plan without an activation. Method call sites remember their verified callee per class version.
  • Inline calls for every common shape. Positional, keyword, *args, **kwargs, keyword-only, f(*args), f(*args, **kwargs), and callable-instance calls run as inline activations in the core loop instead of leaving it for the generic binder.
  • Core-loop coverage. Container methods, native iterators, string formatting, slicing, datetime natives, and generator resumes run in the core loop. sum(), list(), and other draining consumers fold a generator's yields in place.
  • Object model and memory. Class-shared split instance layouts with per-site field shortcuts, non-atomic reference counts while one thread owns the heap, mimalloc, and cheaper collector bookkeeping.
  • Import time. Compiled stdlib modules are cached in a WeavePy-native code format that loads without transcoding CPython bytecode, slicing no longer copies whole sequences (which made re.compile(..., re.I) and so import logging pathological), and the collector's per-drop suspect sweep skips dormant entries.
  • Profile-guided builds. tools/pgo_build.py builds a PGO release, as CPython's release builds are. It trains on the benchmark fixtures at harness work sizes, the bundled regression suite, and a stdlib import sweep.

Two correctness fixes came out of the work: a sum() fold that could claim a nested generator's yields, and delayed finalizers after spread calls. The interned empty tuple now also holds for an empty *args.

Current standing

Wall time relative to CPython 3.14 (lower is better) for a profile-guided release build (tools/pgo_build.py) on the macOS x86-64 development host, JIT on: the median of five paired, interleaved cycles per fixture at the harness work sizes (geometric mean 0.592).

Fixture Ratio
deltablue 3.15x
deque_ops 1.53x
pickle_bench 1.34x
generators 1.18x
datetime_ops 0.97x
str_methods 0.95x
float_math 0.94x
dict_ops 0.90x
call_overhead 0.84x
fannkuch 0.83x
richards 0.80x
nbody 0.77x
attr_access 0.77x
list_ops 0.75x
json_bench 0.71x
pidigits 0.67x
pyaes 0.54x
jitkernels 0.48x
fib 0.37x
spectral_norm 0.21x
nested_loops 0.09x
jitloop 0.08x
sumvm 0.05x

Four fixtures still trail CPython (see Follow-up).

Imports

Retired instructions for a whole process, WeavePy (this branch, warm cache) against CPython 3.14:

Command WeavePy CPython
-c pass 91M 127M
import json 205M 153M
import logging 513M 258M
import asyncio 689M 365M
import unittest 496M 270M

Startup is faster than CPython's; importing large stdlib modules still costs about 2x, and peak memory after those imports is about 2x CPython's (mostly compiled code objects).

Validation

  • VM unit tests: 420 passed (RUST_MIN_STACK=8388608 cargo test -p weavepy-vm --lib --all-features).
  • Bundled regression suite: 249 of 249 pass.
  • Formatting and Clippy pass on Rust 1.98.1, CI's stable toolchain.
  • Full CPython regression sweep: no new unexpected results relative to the branch's baseline (the same seven pre-existing ones).
  • Targeted CPython suites for the call and pickle changes (test_extcall, test_call, test_keywordonlyarg, test_positional_only_arg, test_pickle, test_functools, test_inspect, test_generators, test_gc, test_descr, test_dataclasses, test_typing) pass.

Follow-up

See CPython parity performance for the method and the per-operation gaps. The remaining work, for a separate PR:

  • deltablue: the plan interpreter and the core loop's per-instruction overhead (spilled loop state), spread across method calls, attribute access and frame switches.
  • generators: each resume and yield pair still costs several hundred instructions of activation bookkeeping plus two frame-switch prologues.
  • deque_ops: the deque is a Python class with native methods; its method dispatch and the core loop's per-instruction overhead.
  • pickle_bench: the decoder's two passes and the collector bookkeeping for decoded graphs.
  • Imports: module-body execution and class creation, and the memory held by compiled code objects.

The VM's Rc was an alias for std::sync::Arc, so every clone and drop of a
heap object was a locked read-modify-write. Wrap Arc and Weak in newtypes
whose strong-count updates are plain loads and stores until a second
thread registers with the GIL, free-threading starts, or a Python thread
is spawned. Deallocation still goes through Arc's own drop.

Shared strings, bytes, and tuples use the same biased updates. Debug
builds keep atomics because the unit-test binary runs interpreters on
concurrent threads.
- Port CPython's `_abc` accelerator and use CPython's own `abc.py`.
  Registration and ABC `isinstance` checks no longer run through
  Python-level `WeakSet`s, which also takes `import site` from 25 ms to
  15 ms. The `io` types drop a private `register` that shadowed
  `ABCMeta.register`.
- Port CPython's list sort (timsort with the powersort merge policy)
  with its type-specialized comparisons for `str`, `int`, `float`, and
  tuple keys. Sorting a list that mixes NaNs with other numbers no
  longer panics, and inconsistent comparisons behave as in CPython.
- Use mimalloc as the process allocator and remove the thread-cache
  front end over the system allocator.
- Run container subscripts, subscript stores, and comprehension appends
  in the core loop, out of line so the rest of the loop keeps its
  register allocation. Inserts into dicts of any size stay native when
  the probe proves no key of another kind could compare equal, and
  report the mutation to dict watchers and the builtins rare-event
  counter.
- Memoize whether a type overrides `__eq__` or `__hash__` per type
  version instead of looking the names up for every dict operation.
- Store suspended generator frames as `Box<Frame>` rather than
  type-erased boxes that were downcast on every resume.
- Deliver work queued by a dropped value, such as an unclosed file's
  `ResourceWarning`, at the instruction that dropped it.
…e loop

- Step dict, key, value, and item iterators in the core loop, including
  their first step and exhaustion. An item unpacked by the following
  `UNPACK_SEQUENCE 2` goes straight to the stack, so no tuple is built.
  Iterating a 1,000-key dict's items is 8 times as fast.
- Unpack exact tuples and lists of the expected length, and get
  iterators for dicts and dict views, without leaving the core loop.
- Compare pairs of ints, floats, or strings natively before consulting
  any comparison protocol.
- Make `builtin_types()` return a leaked `&'static` registry instead of
  cloning an `Rc` through a thread-local `RefCell` at every call site.
- Run `isinstance` in the core loop's in-place builtin arm and grade
  dropped type operands inline.
- Serve the JIT's dict reads with the native `LeafProbe` lookup.
- Retire compiled callees that keep calling back into the interpreter
  from native-to-native entries (deltablue's method pattern ran 4x
  slower compiled than interpreted).
- Run `random.Random` on its state bytearray in place instead of
  copying and reallocating 2.5 KB per call (`random()` was 14x CPython).
- Derive the core loop's cold per-activation state lazily, so calls,
  returns and helper handoffs reload less.
- Park yielding generators and deliver clean returns directly, without
  the generic exit plumbing; pass the recursion depth cell through.
- Construct built-in exceptions before the conversion-constructor chain,
  share code objects in traceback entries, and intern the slot keys
  every raise sets.
- Allocate function `__dict__`s on first use and stop mirroring every
  disk-cached frozen module's code in memory.
- Add an opt-in allocation-site profiler (`--features alloc-profile`).

Also includes the getattr/hasattr miss precheck, the slot wrapper miss
cache, single-pass exception class matching, and deferring JIT warm
compiles during startup.
Compiled code called pure-leaf functions and methods (getters,
predicates, small arithmetic) through a full native activation and then
counted each one as a reason to retire the calling loop to the
interpreter. Call sites now evaluate them with the interpreter's
frameless pure-leaf evaluator right after the site's guards, so such
loops stay native; the verdict is mirrored on the code's `JitHint` so
other callees pay one relaxed load. A method with a callback-free
scalar field update plan (`self.n += k; return self.n`) likewise runs
the update directly after the method guard instead of through the
native-call preflight.

Instructions retired, JIT on: a `c.m()` getter loop -32%, a two-argument
function call loop -25%, richards -20%, attr_access -4%.
- Keep up to three earlier resolutions in method slots and call slots, so
  call sites over several receiver classes stop re-resolving on every
  class change (a four-class `o.f()` site: -48% instructions), and widen
  the instance-attribute polymorphic cache to four classes.
- Switch to `CALL_KW` callees inside the core loop, and cache each
  keyword site's validated binding instead of re-checking the names on
  every call; unpack bound-method callees in place like CPython's
  `CALL_BOUND_METHOD_EXACT_ARGS`.
- Serve `dict.get`, `in`, and the generic dict lookup with the native
  `LeafProbe` for `str` and `int` keys, and format `str % scalars` in
  the leaf loop (dict_ops: -25% instructions, now at parity).
- Raises: memoize whether an exception class has its own `__init__`,
  recognize an inert handled exception without the reap walks, and
  re-probe whether the frame object is observed right after `POP_EXCEPT`
  so the frame returns to the fast loop (raise loop: -22%).
- Charge JIT-driven generator resumes like interpreter calls so such
  loops retire to the interpreter's cheaper inline resume (-21%), mirror
  the pure-leaf verdict on `JitHint`, and fill fresh locals in line.
A compiled self-recursive scalar frame (fib's shape) now calls itself
with a native call: the callee's JitFrame and buffers live on the
caller's stack frame, the activation is charged through small enter and
exit helpers (GIL countdown, observers, recursion depth), and a deopt or
raise finishes through a slow helper that rebuilds the callee's
interpreter frame. The running compilation's layout rides on the call
context so that rebuild never depends on the tier cache, which may
retire the code mid-recursion.

A call site whose callee is an already compiled, guard-free scalar leaf
(spectral_norm's `_eval_a`) enters that leaf's code directly, falling
back to the ordinary call helper when the enter helper declines or the
leaf deopts (the leaf only computes, so it restarts from the top).

Keyword calls no longer count toward retiring a native driver loop: the
interpreter binds them through the same permutation, so tier 1 would
not run them any cheaper (call_overhead's loop now stays compiled).

fib: 445M -> 96M instructions; spectral_norm: 657M -> 228M.
- Inline-cache reads on x86_64 without AVX use the epoch snapshot's
  plain loads instead of portable-atomic's run-time dispatched 128-bit
  load (an out-of-line indirect call on every cache read).
- Dicts keep a one-word summary of their `str` keys' hash classes,
  built once per key layout, so the per-call check that an instance
  dict doesn't shadow a method usually skips the probe.
- A frame's small, untracked, uniquely held list, tuple or dict of
  atomic values (a `**kwargs` dict) is released by a plain drop instead
  of the prompt-reap cascade.
- Keyword calls from native code bind into pooled vectors and run the
  callee as a lean activation (a pure leaf evaluates frameless), as the
  interpreter's own `CALL_KW` does, instead of a full framed run.
- `datetime` natives read fields straight from the shared slot layout,
  shift with 64-bit arithmetic, and build exact `date`/`datetime`
  instances from int fields without running the Python `__new__`.

call_overhead: 2280M -> 1805M instructions; spectral_norm, dict_ops,
deltablue and datetime_ops improve 1-10%.
The `mimalloc` crate's `GlobalAlloc` sends every Rust allocation through
`mi_malloc_aligned`, whose fast path needs the size class's free list to
offer a suitably aligned block; otherwise it takes a generic path that
can over-allocate (mimalloc v3 treats only power-of-two block sizes as
naturally aligned). Every mimalloc block is word aligned, so a small
allocator shim over `libmimalloc-sys` uses `mi_malloc`/`mi_zalloc`/
`mi_realloc` for alignments up to a word and the aligned entry points
only above that.

1-4% fewer instructions on allocation-heavy fixtures, and float_math's
peak RSS drops 1.5 MB.
The core loop now loads a `dict`, `set` or `str` receiver's site-cached
native method itself (it already did for lists), and calls every leaf
builtin kind whose body releases no reference beyond its operands (the
`str` methods, `len`, `dict.get` and views, `list.append`/`pop`/
`insert`/`reverse`/`copy`, `set.pop`) through `leaf_builtin_call` right
there. Each such call used to leave the loop twice, once for the method
load and once for the call, and rerun the loop's prologue both times.

About 800 fewer instructions per call: str_methods 955M -> 828M,
call_overhead and datetime_ops improve as well.
The frameless leaf evaluator, which ran pure getters and predicates in
place, now also covers:

- Effect leaves: bodies whose only side effects are attribute stores.
  Stores are buffered as owned values (later loads of the same
  attribute read them back) and committed at the return, after every
  target is re-validated, so a bail-out anywhere leaves nothing
  behind and the ordinary call reruns the body. Setters get a
  dedicated shape.
- Local variables (STORE_FAST and its fused forms).
- Nested calls to pure leaves (methods and functions), evaluated in
  place up to a small depth while no store is buffered.
- A code that keeps bailing (64 consecutive misses) loses its leaf
  verdict, so calls stop paying for the attempt.

Pure method calls also take the interned names from the loaded
constant table instead of re-deriving it per attribute check.

`len()`, truth tests, and iteration over an instance whose class
resolves the dunder to a native builtin call that builtin directly, and
the core loop runs a native iterator's `__next__` in place. Deque
iteration reads its slots by position.

A setter call: 2001 -> 924 instructions; deltablue's `execute`
kernel: 3976 -> 3313; deque iteration: 3278 -> 1562 per item.
- The JIT's attribute helpers read `__slots__` storage with an
  unguarded peek, as they already did for instance dicts.
- An indexed store replacing an existing key's value no longer bumps
  the dict's stamp (the interpreter's store already used the unstamped
  and value-store accessors), so stamp-keyed caches survive it.
- The callback-free field update (`self.n += k; return self.n`) also
  serves `__slots__` classes.
- Slot keys are interned, so guards settle them by identity and no
  instance allocates its own copy of each slot name.

attr_access: 922M -> 730M instructions.
Cross-crate inlining between the VM, compiler, and JIT crates makes the
benchmark fixtures 1-10% faster in wall time (pickle_bench 10%,
float_math 6%, deltablue 3%) for about a third more build time.
A dying list of plain objects no longer walks each element's fields:
scalars and deferred-tracking instances can't anchor anything the
cascade must visit. The scan also reuses the handles it finds instead
of looking each child up twice, and hashes ids with the VM's id hasher.
An instance that assigns its attributes in its class's usual order keeps
them as a plain vector of values whose names live once per class, like
CPython's shared keys. Constructing such an instance allocates no hash
table, and an attribute keeps the position a __dict__ entry would have,
so the indexed attribute caches serve both layouts. Anything that needs
the real dictionary (vars(), del, an out-of-order attribute, the C API)
materializes it, and the instance keeps it from then on.

A three-attribute instance drops from 490 to 249 bytes, and constructors
run 15-20% fewer instructions (float_math: -11% instructions, -20% wall).
The rarely set native-value slot becomes a pointer-sized write-once cell,
so PyInstance grows by only one word.
A keyword call that skips a defaulted parameter (f(x, c=1) over
def f(a, b=0, c=0)) used to fall back to the generic dynamic-call path,
which re-verified the keyword names and bound a fresh frame every call.
The JIT now marshals it like any keyword call and tags the skipped
slots; the call helper fills them from the callee's current scalar
defaults (or, for anything else, binds the call generically by name),
so the compiled callee runs natively. Compiled scalar leaves are also
preferred over the frameless evaluator when both apply.

call_overhead's f(i, c=5) shape: 256 ns -> 52 ns per call (CPython: 99).
- Retire an exhausted range or list iterator in the core loop instead
  of handing the loop exit to the full handler, as the leaf arm does.
  An interpreted loop over a short list costs 36% fewer instructions.
- Build list literals in the core loop, and release the last reference
  to an untracked list or dict of scalars there as a plain drop instead
  of taking the prompt-reap path.
- Reclaim dead deferred containers at the tail of the deferral list as
  new ones arrive, so container churn neither grows the list nor parks
  freed allocations until a sweep.
- Allocate 16-byte-aligned layouts through mimalloc's plain entry points:
  every block of 16 bytes or more already has that alignment, and the
  aligned entry points fell to their generic path for iterators.
- Recognize the process-wide NotImplemented and Ellipsis singletons by
  identity: they carry the type registry of the thread that built them,
  so a class check against the current thread's registry failed for
  every later thread (the native _abc raised AssertionError).
- Restore compiling sustained start-up and import work: warm compiles
  during those phases count lean intervals again, rather than deferring
  every one, and frameless leaf calls count toward the warm-up, handing
  the call at which a compile falls due to the framed path.
- Count direct self and leaf calls and frameless call-site evaluations
  in the native-call statistics, and scalar field updates' boxed returns
  on the method path, which the tests assert on.

All 419 VM unit tests pass with the CI stack size.
- Fuse x.m(simple args) for a class's registered native method (a
  deque's append, popleft, pop, ...): the builtin runs on bitwise views
  of the borrowed receiver and arguments, with no method object, no
  receiver or argument references taken, and no operand stack traffic.
- Serve a subscript whose class has a cached native fast __getitem__ in
  the core loop instead of handing it to the full arm.
- Cache how each class answers a boolean context (native __bool__ or
  __len__, always true, or the full path) by attribute version, instead
  of looking both names up for every test.
- Give deque iteration, integer indexing and appendleft guard-free fast
  paths, and keep an emptied deque's free prefix for the next appendleft.

deque_ops: 2.58x -> 2.03x CPython in instructions; append/popleft pairs
-29%, indexing -44%, truth tests -39%.
- Don't compile a loop-free function that makes a dynamic Python call:
  from native code the call takes the interpreter's generic path (a
  nested run rather than the inline activation an interpreted caller
  uses), and compiling was the largest cost a short-lived method paid.
  deltablue (default size) drops from 97 ms to 81 ms, with JIT compile
  time cut from 17 ms to about 10 ms.
- Switch generator state in and out of its cell without guard atomics
  on inline resumes and yields (about 5% per resume/yield pair).
A LOAD_ATTR or STORE_ATTR site that proved a split-layout field records
the class version and shared-key position, so later hits skip the inline
cache decode and the key-name check. The leaf evaluator's operand values
also get a primitive layout, so copying one is two words.
The frameless leaf evaluator now translates each pure or effect leaf
once into register-addressed ops (locals and stack positions get fixed
registers, jumps name their target op, and local loads, POP_TOP, COPY
and SWAP mostly disappear) and runs them with one unchecked dispatch per
operation. A path that reaches an instruction the plan can't express
still declines only when it runs.
A constructor whose __init__ is a pure or effect leaf returning None now
runs it through the leaf evaluator on the fresh instance, storing
straight into it (nothing else can see it, so a decline just discards
it). New-attribute stores take the split-layout shortcut by appending
the class's next shared key, and an effect leaf's multi-store commit
skips the latest-store scan when every store goes to self under a
distinct name.
Interpreter call sites keep evaluating a pure or effect leaf frameless
after the JIT compiles it (native callers still enter the compiled
code). The collector's tracked-handle references use the biased Rc, its
Bloom-filter inserts skip already-set bits, and its population counters
update without locked read-modify-writes while the GIL serializes them.
The generation-0 threshold now matches CPython 3.14's (2000).
The math module's functions register as leaf builtins over plain int,
float and bool arguments, so math.sqrt(x) no longer drops to the
generic call path; module attribute reads hit the site's cached index
in the core loop's LOAD_ATTR arm; and a polymorphic method site now
remembers each receiver class the class cache resolves for it, so its
fused frameless leaf call keeps serving every class.
…ispatch

len() of an instance whose class serves __len__ with a registered native
leaf builtin calls it through a version-keyed class cache; the core
loop's TO_BOOL and the JIT's truth helper answer instances through the
native truth cache; and generic iteration (sum, list, ...) calls a
registered native __next__ directly, with class misses cached too. The
deque's slot probes use interned names.
…ut a lock

The core loop's BINARY_OP serves str + str, small str * int, and str %
over scalar and string arguments through an out-of-line helper instead
of handing the instruction to the full leaf arms. Dict mutation stamps
come from a plain load and store while the GIL serializes mutation (the
locked increment stays for free-threaded mode and the debug test
binary).
…ore loop

A datetime-style instance's public field read, its binary operators and
its comparisons now run from the core loop's LOAD_ATTR, BINARY_OP and
COMPARE_OP arms (through out-of-line calls into the native bodies)
instead of leaving the loop twice for the helper and the full leaf arms.
The tier-2 attribute get and set helpers try the common shape first (a
scalar lane over an indexed instance field whose class is unchanged)
before the general lane classification, and the class guard is always
inlined.
… checks

The core loop's fused LOAD_GLOBAL len; LOAD_FAST x; CALL 1 now also
serves an instance whose class's __len__ is a registered native body
(a deque's), skipping the separate push, call, and pop: len(q) on a
deque drops from about 1110 to 620 instructions.

The pickle decoder's first pass verified every instance's __slots__
names against the class MRO, once per instance. It now remembers the
members it verified for the stream, so a list of slotted instances
checks each name once (records loads: 4% fewer instructions).
tools/pgo_build.py builds an instrumented CLI, trains it on the
benchmark fixtures at the harness's work sizes (JIT on and off), the
bundled regression suite, and a stdlib import sweep, then rebuilds the
release binary with the merged profile, as CPython's release builds
do. On the measured macOS host this cuts wall time on most fixtures by
15% to 30% (attr_access 1.24x to 0.84x of CPython, call_overhead 1.10x
to 0.87x, float_math 1.27x to 1.01x). Training at toy work sizes left
large-integer multiplication cold and slowed pidigits by 38%, hence the
harness sizes.
The frameless evaluator copied each plan op and decoded all of its
operands before dispatching on it, keeping the copy's fields and the
instruction index in stack slots: about 21 instructions per op.
Matching the op in place lets each arm load only its own operands, and
dispatch drops to 5 instructions. The evaluator's scratch buffers also
stop calling their out-of-line drops when they are empty (the usual
case).

The collector's temporary id sets, its suspect map, and the descriptor
registry hashed object addresses with SipHash. They now use the
address mixer the collector's main index already uses (these tables
never hold untrusted keys), and so does the prompt reaper's cascade
marker set.

Measured with callgrind: DeltaBlue retires 5.6% fewer instructions, and
the pickle benchmark 1.9% fewer.
A forwarding wrapper's f(*args, **kwargs) compiles to BUILD_MAP 0,
LOAD_FAST kwargs, DICT_MERGE and CALL_FUNCTION_EX: a fresh dict,
tracked by the collector, filled by copying kwargs, only to be read by
the callee's binder. When the callee is a Python function the core loop
now binds straight from the local mapping and skips the copy; anything
else runs the instructions as before, untouched.

The keyword binder also writes the parameters into the activation's
locals directly (it staged them in a temporary vector and copied them
over), sizes a **kwargs dictionary up front, and releases the call's
operands in place instead of collecting them first. A callable
instance forwarding w(5, protocol=5) through
__call__(self, *args, **kwargs) runs 30% fewer instructions.

Also make the shared-string thread handoff test revoke the reference
count bias before it spawns threads, as the VM does: under the biased
counts its raw threads raced with the owner's plain increments, and the
test failed depending on test order.
… once

The deque natives resolved each of _data, _head, _maxlen and _state by
a separate hinted name lookup, several per operation. SlotStorage now
offers leading/leading_mut, which match the first N slots against
interned names by thin pointer in one pass, and the deque's fast paths
and iterator step work on that view. The core loop's container
subscript no longer grades an instance operand it then declines (the
native subscript grades it), shared instances skip the full drop grade,
and native method calls stage operands straight into their slice.

deque_ops: about 13% fewer instructions.
…preter

A compiled function entered directly from the interpreter
(try_call_native_direct) never had its interpreter round-trips counted:
only framed entries and native-to-native callees did. deltablue's
incremental_add ran every constraint.satisfy call through an activation
shell and a framed interpreter call, several times the cost of the
interpreter's inline call. Direct entries now feed the same callee
round-trip backoff.

deltablue (JIT on): 3% fewer instructions.
- The reload prologue no longer caches the globals and builtins handles
  (only LOAD_GLOBAL and LOAD_NAME read them, off the frame), so every
  call, return, resume and yield spills four fewer values.
- VmExt caches the extension table's thin address, so code_vm_ext is one
  load instead of a fat dyn pointer and its alignment arithmetic.
- An inline generator resume writes its slot's fields without drop glue,
  the yield parks the frame over a Running state without it, and the
  yield's consumer protocol runs without a QuietEntry round trip.
  Frame::push inlines.

A for loop over a generator: 7% fewer instructions per item.
…l drop grade

A pinned receiver or argument that other references still hold, with no
weakref watching it, can't be at its dead line (past two owners only a
weakref clone could put it there). gc_trace::drop_survives_plainly says
so in two loads; the JIT's pin drain and the core loop's operand check
take it before the prompt-reap and drop grades, which cost about 120
instructions per pin. The fused-pair bisection flag moves into the
per-code table, off the frame switch.

float_math (JIT on): 6% fewer instructions.
- The core loop's clean-return test counts each object's own slots
  instead of assuming every heap local aliases every other, so frames
  holding a few distinct objects skip the exit reap.
- The exit reap's fast path untracks a dying collector-tracked list or
  dict whose contents can't finalize (deltablue's drained todo lists)
  instead of running the full prompt-reap cascade, and skips its alias
  count for objects held past every slot.
- A leaf plan's nested pure-leaf call returns its result borrowed when
  the callee's scratch doesn't hold it (see LeafRet), so the caller
  neither clones nor later releases a getter's result.

deltablue: about 6% fewer instructions.
Each dumps call grew a fresh memo table from empty (rehashing as it
doubled) and a fresh output buffer (copying as it doubled). The encoder
now takes the thread's spare table, emptied with its capacity kept, and
reserves the previous call's output size.

pickle_bench: about 3% faster.
- A local list, dict, set or str receiver of a site-cached leaf method
  (x.append(v), d.get(k), s.startswith(p)) is called on the borrowed
  local, as instance receivers already were, so no receiver clone is
  pushed and graded on release.
- The unpickler builds plain instances with deferred collector tracking,
  as ordinary construction does; slot state tracks them only when a value
  isn't atomic (dict state already did). A decoded Point of scalars now
  dies by plain drop instead of the prompt-reap cascade.
- The prompt-reap cascade allocates its per-node scratch once per
  cascade.

pickle_bench: about 5% fewer instructions; list.append/pop loop: 10%.
The run fixture for weakrefs and gc still expected the old default
collection threshold of 700; the branch adopted CPython 3.14's 2000 (the
fixture's output now matches CPython's). The remaining changes are
rustfmt and clippy fixes, with alignment and mutable-borrow allowances
where the pointer casts are sound by construction.
The branch stopped compiling standard library functions during startup,
so a program's first compile now pays the code generator's cold start
(about 0.7 ms: its code paged in, its tables built) inside its hot loop;
the Windows bench gate flagged sumvm at +18% against the merge base. The
CLI now compiles and discards a small loop on a background thread when
it starts a program file or module, and any other run does so when its
first code object is halfway to the compile threshold. -c pass is
unaffected (no extra memory).

sumvm, jitloop, nested_loops: back within 1-3% of the merge base.
Resuming a generator switched activations (the frame made the running
one, the pending-caller and recursion bookkeeping, the core loop's
reload) and back at each yield: most of the cost of each item for a
body that only moves scalars through locals. A new stepper (gen_fast)
runs such a body directly on its frame to its next yield: locals,
constants, int and float arithmetic and comparisons, jumps, and for
loops over ranges, lists, tuples and other such generators. It never
raises or runs Python code; anything else ends the step at an
instruction boundary and the ordinary resume continues from there, so
a fast step is always a prefix of the ordinary run.

The core loop's FOR_ITER over a generator and the lean resume path
(next(), and sum()/list() folding) take fast steps first. A nested
generator whose step stops partway is parked having consumed its sent
value (Frame::sent_consumed), which every resume path honors.

generators: 28% faster; a for loop over a generator: 34% fewer
instructions per item.
The decoder's push and the encoder's extend reserved capacity (a
fallible call) before every append. They now reserve only when the
buffer is full, and the encoder's frame writer inlines.

pickle_bench: about 8% fewer instructions.
The eligibility scan now also proves every jump target, local index and
constant index in range and that the last instruction can't fall
through, so the step reads instructions, locals and constants without
per-use checks. The scalar arithmetic and comparison helpers inline,
which keeps the step's operand stack in registers.

A for loop over a generator: 7% fewer instructions per item; sum() and
list() over generator expressions: 5-9%.
The core loop's fused `x.m(<args>)` step now folds an argument like
`i + 1` (checked add, subtract, multiply, and bitwise ops on two ints)
in place, and a local instance whose method is a native slot calls it
directly before the leaf-site lookup. The deque fixture drops 8.4% in
instructions, and `q.append(i + 1)` drops 31%.
- A class whose __init__ takes trailing default parameters now takes
  the frameless constructor paths (and the lean framed one): the call
  fills the missing arguments from the function's defaults.
- Leaf plans allocate empty list and dict literals, so an __init__
  like `self.items = []` stays a leaf.
- A leaf plan runs a pure-leaf callee as a frame of its own
  evaluation, saving only the caller's live registers, instead of
  re-entering the plan runner.
- The GC's finalizer check on tracking uses the class's cached
  __del__ verdict rather than an MRO lookup per instance.

Constructing a DeltaBlue Variable drops from 12.5K to 7.4K
instructions, and building its constraint chain drops 7%.
A leaf plan's in-place callee frames now apply only to general plans.
A certified tiny shape (a constant or field return, a field
comparison) keeps its dedicated evaluator, as before the frames: it's
at least as fast, and the unit test that counts cached field
predicate hits on DeltaBlue passes again.
Slicing a tuple, bytes, bytearray, or a string with surrogates used to
copy the entire sequence into a vector of objects and then select from
it, so every slice cost the length of its source. `slice_seq` is now
generic over the element type and copies a unit-step slice in one
piece, and each type slices its own storage.

`re.compile` with IGNORECASE slices a 64K charset map 256 times, which
made `import logging` cost 1.19B instructions; it now costs 525M.
The private frozen-stdlib cache stored CPython marshal data, so every
import unmarshalled it and transcoded the CPython bytecode back into
WeavePy instructions. The new `native_code` module encodes a code
object's fields directly (varints, delta-coded line tables, nested
code inline) and decodes them with no transcoding; filenames are
stamped at load. The source check hashes eight bytes at a time.

`import json` drops from 269M to 205M instructions, `-c pass` from
106M to 91M, and `import dataclasses` from 454M to 349M.
Suspects live in two maps, active and dormant, so the sweep that runs
at drop safe points walks only the entries with probe budget left; the
dormant ones are visited on the stride, as before. During imports the
old single map made every sweep walk up to 256 entries.

`import logging` drops 3.5% and `import json` 7%.
- A class's scalar attribute read through an instance (`self.LIMIT`)
  is served from the site's stamp slot while the class version holds
  and the instance doesn't shadow the name, instead of resolving the
  name on every read.
- The fused `x.attr` arm reads the site's field shortcut in line.
- `if x is None`, `if x is not None`, and `if x:` on a local decide
  the branch from the local in place (no load, clone, or release),
  and a conditional jump steps over the `NOT_TAKEN` that follows it.
- `POP_TOP` of a scalar no longer calls into the drop glue.

A class scalar read through an instance drops from 730 to 280
instructions, and an `is None` branch by a fifth.
`STORE_GLOBAL` and module-level `STORE_NAME` of an existing name used
to stamp the module dict like any mutation, so every cached global
load in the module missed after each store. Rebinding now replaces
the value in place: the key layout, and so the stamp the global-load
caches check, stays put. The JIT's burned-in globals, which remember
values rather than positions, also watch a process-wide epoch that
each in-place rebinding advances. The core loop runs `STORE_GLOBAL`
itself, as it already did module-level `STORE_NAME`.

A global store drops from 935 to 320 instructions, on par with
CPython, and a call to a function that updates a global counter from
3.7K to 1.8K.
Record the paired wall-time checkpoint for the benchmark fixtures, the
startup and import costs, the per-operation gaps that remain, and the
method behind each number.
A fast step stops at a back edge when the GIL countdown runs out. For
an inner generator that left its frame partway, and `gen_fast_next`
then declined it for good, so an outer generator fell back to the
general loop for the rest of its run. Fast steps now continue a
partial frame where it stopped, and a draining consumer's send
services the eval breaker itself (the countdown reset, the dormant
suspect probe, the GIL checkpoint) and keeps stepping unless the loop
generation moved.

Summing a generator over a generator drops 6% to 8%.
The core loop's fused `xs.append(x)` and `xs.pop()` on a local list
now push and pop directly instead of admitting the call and then
entering the generic builtin, about 8% off each call.

`list.pop` also accepts a bool index, as CPython's does
(`[1, 2, 3].pop(True)` returns 2; it raised TypeError).
CI's stable toolchain moved to 1.98.1, whose rustfmt orders the
`initconfig` imports differently and whose clippy asks for
`as_chunks` in the frozen-cache hash and `usize::midpoint` in the
binary insertion sort. Both are stable well before the 1.93 MSRV.
…e GIL

Under `-X gil=0`, two threads putting into the same
`queue.SimpleQueue` could both see a parked getter's `_locked` flag,
both clear it, and both release the gate; the second release raised
`RuntimeError('release unlocked lock')`. `multiprocessing.dummy`'s
thread pool returns results that way, so the free-threaded lane's
`test_multiprocessing_dummy` failed intermittently (about 2% of runs
on main, 5% here).

The flag's test, clear, and release now happen under one lock, as
CPython's critical section makes them. The append stays outside it,
since it may collect and run Python code. The test passes 400 of 400
stressed runs.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant