Skip to content

Compile whole programs to machine code - #2

Draft
treeform wants to merge 40 commits into
masterfrom
jit
Draft

treeform wants to merge 40 commits into
masterfrom
jit

Conversation

@treeform

@treeform treeform commented Sep 22, 2026 •

Copy link
Copy Markdown
Owner

Compiles a whole BASIC program to AArch64 or x86-64 machine code, emitted
directly from Nim. No external compiler and no JIT library, so the package
keeps its only dependency.

What runs natively

Everything. Every bytecode offset becomes a native block, so jumps, calls,
GOSUB, returns, and budget meters go straight between blocks and the
interpreter loop never runs. Values stay exactly where the interpreter keeps
them: globals, register file, arguments, array cells.

Inline: whole-number and fixed-point arithmetic (mixed kinds promoted exactly
as the interpreter does), comparisons, fused branches, arrays, calls,
returns, EXIT SUB, meters.

Strings, printing, host calls, and any instruction about to fail call the
interpreter's own code for that one instruction. The run loop's instruction
body is now a template shared by both paths, so errors, budgets, output, and
even exception types cannot diverge. Interpreter speed is unchanged.

A loop that touches few globals (8 on arm64, 3 on x86-64) also gets a
specialised copy that keeps them in registers, proves they are whole numbers
on entry, and checks the budget once per pass at every backward-branch
target. Anything unexpected writes the registers back and continues in the
general code at that very instruction.

This replaces the earlier loop-region compiler, which also reported a host
exception as a plain BasicError instead of the original type.

Speed (M4, local)

arm64 x86-64 (Rosetta)
tight loops 26-33x 43-58x
raytracer 4.2x 4.8x

The raytracer compiled nothing before. Only host calls leave native code now.

Why it stays honest

  • -d:bassyNative compiles every runtime, so the whole suite runs native
    (release, danger, fixedChecks). A refused program fails loudly.
  • tests/test_native.nim fuzzes ~1100 generated programs mixing kinds,
    strings, arrays, subs, raising host calls, and budget exhaustion, and
    compares every global, array cell, print event, budget, stop offset, and
    exception, including resuming after a failure and host code writing
    globals mid-run.
  • tests/test_jit_safety.nim hands the compiler malformed bytecode and
    requires refusal.
  • Encoders are checked against the system assembler.

Windows x64 only cross-compiles locally; this CI run is its first real test.

🤖 Generated with Claude Code

treeform and others added 30 commits September 22, 2026 13:11
Add a native backend that turns backward-branching integer loops in the
register bytecode into machine code emitted directly from Nim, with no
external compiler or JIT library.

- machine.nim reserves write-then-execute pages, handling Apple silicon's
  MAP_JIT and per-thread write protection, mprotect on Linux, and
  VirtualProtect on Windows.
- arm64.nim encodes the AArch64 subset the compiler needs, with labels and
  forward-reference patching. Every encoding is checked against clang's
  assembler in test_arm64.nim.
- bytecode.nim holds Op and Instruction so the interpreter and the native
  compiler can share them.
- jit.nim proves each participating global still holds an integer, hoists
  it into a callee-saved register, and runs the loop without tags, memory
  traffic, or dispatch. Unmodelled operations and non-integer values fall
  back to the interpreter, and the meter is charged exactly as before, so
  a script cannot tell which path ran.

Measured on an M4 Pro: arithmetic 16.2 ms to 0.69 ms, branches 26.0 ms to
1.34 ms, nested loops 15.8 ms to 0.67 ms, with identical results and
identical instruction and work accounting.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Gate every platform declaration behind a single NativeCode constant, so a
target without a code generator emits no mmap, VirtualAlloc or cache
intrinsic at all. Add -d:bassyNoJit to force the interpreter on a target
that would otherwise qualify.

Verified three ways on this machine: native AArch64, -d:bassyNoJit, and a
real emscripten wasm32 build run under node. All three agree with the
interpreter on results and on budgets.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Restructure the region compiler around a small set of emitters that each
architecture supplies, so the walk over the bytecode is written once and
AArch64 and x86-64 cannot drift apart.

- amd64.nim encodes the x86-64 subset the compiler needs. Its test checks
  that each encoding decodes to the intended instruction rather than
  matching clang byte for byte, because clang picks shorter immediate and
  displacement forms that mean the same thing.
- jit.nim gains the System V register assignment. Windows x64 passes its
  first argument elsewhere and preserves a different register set, so it
  keeps using the interpreter.
- Dividing by one or minus one leaves no remainder, and minus one traps on
  x86, so neither ever reaches a divide instruction.
- bench_jit.nim now times itself instead of pulling in benchy, so CI can
  report speedups on every runner.

The matrix now covers arm64 and x86-64 on both Linux and macOS plus
x86-64 Windows, which exercises the interpreter-only path.

Verified on linux/amd64 in Docker: results and both budgets match the
interpreter on every case, including budget exhaustion and the guard
fallback.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The scene, lights and algorithm follow tests/bench_raytracer.nim in vmath.
Three things change because of the language rather than the algorithm:
numbers are Q16.16 fixed point, the scene lives in parallel arrays because
there are no records, and anything that must survive a recursive call is
passed as an argument because every scalar that is not a parameter is
global.

It compiles no loops at all. Every hot path is fixed-point arithmetic,
array indexing, or a subroutine call, and the native compiler models none
of those yet, so interpreted and native run at the same speed. That is the
point of adding it: it marks where the speedup does not reach.

It doubles as a cross-architecture determinism check. Fixed point is the
reason this VM has no floats, so the pinned checksum has to match
everywhere. Verified identical on arm64 macOS and amd64 Linux.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Comparing the two paths against each other only shows that they agree on
the machine running them. The raytracer made this obvious: it compiles no
loops, so its native run was the interpreter and the comparison proved
nothing.

Fold every global into one value and pin it. Now the benchmark fails
unless generated machine code produces the same answers everywhere, and
the check keeps its meaning on targets that compile nothing.

Verified three ways: arm64 macOS with the native backend, the same machine
with -d:bassyNoJit, and amd64 Linux with the native backend.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The port divided the screen offset by half the image size where the
reference divides by twice it, so the camera had four times the intended
field of view. That is what made the floor fan out. With it corrected the
composition matches the reference: two spheres above a checkerboard.

Three other fidelity gaps close with it. The reflection depth is five
again, the reference roughness values of 250 and 150 are affordable now
that a host function raises to them by squaring instead of a multiply loop
in BASIC, and shadow rays start slightly off the surface so fixed point
cannot make a point shadow itself.

Square roots now use the digit-by-digit method on the raw Q16.16 bits,
which needs no division and no iteration count.

Then four optimizations, each leaving the image byte for byte identical:

- The floor is the only plane and is tested first, so its intersection is
  written out directly and the object loop starts at the spheres. Nothing
  in the loop is a plane any more, so the kind test goes too.
- applyLight took fifteen parameters it could read as globals, since it is
  never on the recursive path.
- A light facing away from the surface that also makes no highlight adds
  nothing whether or not it is blocked, so its shadow ray is skipped.
- Six values were copied before uses that nothing could disturb.

Together: 4,986,746 instructions down to 3,833,996, and 24.6 ms down to
20.0 ms at 48 by 48 on an M4 Pro.

Profiling says where the rest goes. Loading and storing globals is 36% of
everything executed, because every BASIC scalar is a global and so every
mention of one is a memory round trip. Two script-level attempts to beat
that made it worse: replacing three divides with a reciprocal and three
multiplies, and hoisting array reads out of a branch, both traded cheaper
arithmetic for more instructions and lost. In this VM the instruction
count is the cost.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The generated code spent most of itself on metering. Twenty-six
instructions ran per pass of the counting benchmark to do two
instructions of work: two budget checks that each rebuilt their limits
from scratch, a comparison constant rebuilt every pass, and an
unconditional branch after every block to hop over an inline stub.

- Exit stubs move out of line, so a conditional branch goes straight to
  one and the common path falls through with nothing to jump over.
- A loop whose body has no internal branch costs the same every pass, so
  the budget is settled on entry by dividing: how many passes can both
  budgets certainly afford. The loop then only counts. One pass is held
  back, because the pass that finally leaves charges on the way out.
- A loop whose passes differ adds up what it spends and asks once a pass
  whether another whole pass is still affordable. Blocks inside it charge
  with two adds and no branch.
- Compare constants too wide for an immediate are hoisted into registers
  before the loop rather than rebuilt inside it.
- A remainder against a power of two is a bit test, not a divide and a
  multiply, and that holds for negative values under truncating division.
- Compiled loops are found through a sequence indexed by offset rather
  than a table.

The counting loop is now eight instructions a pass:

    cmp x13, x14      ; passes left
    b.ge handback
    cmp w22, w1       ; i against a constant already in a register
    b.ge exit
    add w23, w23, w22
    add w22, w22, #1
    add x13, x13, #1
    b loop

On an M4 Pro: counting 0.67 ms to 0.25 ms, 24x to 65x. Branching 1.22 ms
to 1.02 ms. Nested 0.53 ms to 0.46 ms. x86-64, which keeps the per-block
check, still gains the out-of-line stubs.

Budgets are charged to the unit as before, and refusals still land on the
exact offset, because a loop that cannot afford a whole pass is handed
back to the interpreter rather than approximated.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
macos-13 is being retired and never left the queue, which kept the whole
run from publishing its logs. macos-15-intel is the label GitHub offers
for x86-64 macOS now.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Windows x64 reaches the same code generator as System V. The two differ
only in which register carries the argument and which ones a callee has
to preserve: Windows takes it in rcx and must return rsi and rdi intact,
System V takes it in rdi and may use both freely. Everything else, down
to the encodings, is shared.

The kernel32 declarations now follow the Windows header types exactly.
DWORD is an unsigned long there, a different type from an unsigned int
even where the two are the same width, and a cross build with mingw
rejected the old signatures.

Also drop macos-15-intel from the matrix. It never failed on this code:
setup-nim-action installs an arm64 Nim on the Intel runners, so nim
itself will not start, reporting a bad CPU type. The comment in the
workflow says to put it back once the action picks its download by
architecture.

Windows is verified here only as far as a mingw cross build of every
test; the runner is what will actually execute it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The note about restoring an Intel runner implied a plan to cover x86-64
macOS again. There is none, so it goes.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Compiled code indexes global storage without checking and writes through
offsets worked out at compile time. The interpreter checks every access as
it runs and compiled code cannot, so what makes that safe is the proving
done before any of it is emitted, and that proving was not thorough
enough.

The escape: a loop whose passes differ in cost looked at its spending only
at the region's first offset. A loop nested inside one of those reaches
that offset once for however many times it goes round, so it could spend
as long as it liked in between. A script could exceed its instruction
budget by whatever factor its inner trip count gave it, which is the
denial of service the meter exists to prevent. The spending is now looked
at wherever a backward branch lands, so what runs between two looks
contains no backward branch and can charge no more than the one pass the
limit is set against. tests/test_jit_safety.nim found this; it showed
twice the work done for the same budget.

Also proved rather than assumed:

- Every global index is checked against the storage that exists. Out of
  range, or negative, and the loop is refused. This was an encoding range
  check on AArch64 and nothing at all on x86-64.
- An operation naming a global the gathering pass did not see now refuses
  instead of reaching for whichever register came next, which would have
  been a budget register.
- The value and context layouts the generator writes by hand are confirmed
  at run time before any loop is compiled. These offsets were read off
  this Nim version; nothing holds them there, and a quiet change would put
  every compiled store at the wrong address.
- A pass that charges nothing cannot turn forever against an unmoving
  total.

The new tests come in two halves: bytecode the language's own compiler
could not produce, which must be refused rather than emitted, and scripts
run down both paths whose globals, budgets and failures must match, since
a script that could tell the two apart could be written to exploit the
difference. Four hundred generated scripts are included. CI runs them
under both -d:release and -d:danger, the second being where the
interpreter's own bounds checks are gone.

Speed is unchanged: 0.24 ms, 0.94 ms and 0.45 ms on an M4 Pro.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Compiled code proves its globals are integers and loads them into
registers on the way in, so arriving anywhere but the first offset would
skip the proof and read registers that were never filled. The check for
that only knew the branches the generator compiles, and three operations
it refuses to compile can still name an offset inside a loop it did:
a subroutine call, a label return, and a register test.

Nothing could reach those entries today, because a loop is only ever
called at its first offset, but the check is what the argument rests on,
so it should be the one that is complete.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The check builds its values in a sequence and reads them through copyMem
instead of casting the address of a local. That is not a style choice:
Nim 2.2.6 and 2.2.10 both fail to compile a procedure that takes the
address of a converter-initialised variant local and also returns early,
reporting an index error with no location. Anyone tidying this back into
casts would hit it, so the reason is now written down beside it.

Also drop an import the tests stopped using.

Verified on Nim 2.2.10, which is the minimum the package asks for:
the suite, both encoder checks, the equivalence and safety tests under
-d:release and -d:danger, the interpreter-only build, an emscripten wasm32
build under node, and the same set on x86-64 Linux in a 2.2.10 container.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Three findings from review, none of which the tests would have caught.

Pages were mapped readable, writable and executable everywhere POSIX.
Linux dropped write when sealing, but macOS did not, so an Intel Mac left
every region writable and executable for as long as it was mapped. Only
Apple silicon needs that combination, because there the writing is gated
per thread instead; everywhere else the pages are now writable until
sealed and executable afterwards, never both.

Nothing ever released a region. compileNative drops the previous ones on
each call, so a program compiling many scripts accumulated executable
mappings. The buffer now owns its pages: it cannot be copied, and its
last owner unmaps them. Twenty thousand compiled programs hold steady at
1.4 MB where they would have leaked about eighty.

A block charging more than an add-immediate can hold raised out of
compileNative instead of leaving that one loop interpreted. Such a block
now keeps the per-block check, and compiling many loops survives one it
cannot finish either way, because the interpreter runs everything the
generator declines. A global whose byte offset would not fit the
displacement it is reached through is refused as well; the limits do not
allow one, but nothing was checking.

The fourth finding, that a cycle inside a loop could run without meeting
the budget, was fixed in "Close a budget escape" two commits earlier and
is verified again here. Four shapes are exercised, each under a budget far
smaller than it wants: nested while, a backward goto inside a while, three
levels of nesting, and a program that is one goto cycle with no while at
all. All four charge exactly what the interpreter charges and refuse at
exactly the same point.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Coverage was one loop in five. The generator took loops whose every
operation was one of a dozen fused global forms the bytecode compiler
happens to emit, so `n = n - 3` was already a miss: subtraction has no
fused global form, and the loop fell out over a load, a subtract and a
store. The benchmarks flattered it, being written by accident in its own
dialect.

The frame's register slots now reach compiled code. They stay where they
are rather than being hoisted, which costs a memory round trip per operand
but keeps two properties worth more than the speed: a slot holding
something other than an integer hands that offset straight back to the
interpreter, and nothing has to be written back when it does, because
every result was already written where the interpreter would look for it.
The charge for the pass so far is the one every other exit already uses.

Added: load immediate, move, load and store global, add, subtract,
multiply, negate, the six comparisons, and jump if zero. Every slot index
is proved inside the frame before anything is emitted, as global indices
already were.

    subtract      21.4 -> 2.6 ms    8.4x   was not compiled at all
    multiply      27.6 -> 2.0 ms   14.0x   was not compiled at all
    comparison    19.2 -> 0.8 ms   23.3x   was not compiled at all
    fused          15.1 -> 0.2 ms   63.7x   unchanged

Three loops in five now, the remaining two blocked by arrays and by calls.
Slower than the fused forms, as it should be: those keep their values in
machine registers across the whole loop, these fetch and store each one.

Verified on arm64 macOS and x86-64 Linux, under -d:release and -d:danger,
with the interpreter forced, as a wasm32 build under node, and as a
Windows cross build. The four budget-escape shapes are still bounded.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Coverage was three loops in five, arrays being one of the two that were
left. A loop as ordinary as summing cells fell out over a single fused
operation the generator did not model.

Reading and writing a cell copies the value entire, whatever kind it
holds, because that is what the interpreter does: it assigns the value
across and never asks what is in it. So these copy sixteen bytes and ask
nothing either, which means an array of fixed-point numbers compiles as
readily as one of integers. Only the two fused forms that add through a
cell need an integer, and those guard for it.

The bounds check is the interpreter's: one unsigned comparison covering
both ends at once. Out of range hands the offset back rather than raising,
so the message still names the array and its real extent, and still comes
from the one place that knows them.

    array read    17.8 -> 0.76 ms   23.4x   was not compiled at all
    array write   20.3 -> 1.13 ms   17.9x   was not compiled at all

Five loops in six now. Calls are the only thing left, and the only thing
left that needs a different machine.

Every array named is proved to exist and to sit where a displacement can
reach it. Scripts that walk off either end of an array, index one
negatively, or fill one with fixed-point values are in the safety tests
and agree on both paths.

Verified on arm64 macOS and x86-64 Linux, under -d:release and -d:danger,
with the interpreter forced, as a wasm32 build under node, and as a
Windows cross build. The four budget-escape shapes remain bounded.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Remainder and integer division were missing, which an array benchmark
caught by measuring nothing: it indexed with `i mod 1024` and so compiled
no loop at all. Both now compile, with the three divisors that are not
plain division going back to the interpreter: zero, which it refuses, and
minus one, which traps on one of the two architectures.

    remainder         28.5 -> 2.1 ms   13.4x   was not compiled at all
    integer divide    29.2 -> 1.8 ms   16.2x   was not compiled at all

Fixed-point numbers turn out to need very little. Adding, subtracting,
negating and comparing them is the same instruction on the stored bits as
for whole numbers, so one path serves both: a pair of operands whose tags
agree needs no further telling apart, and the answer keeps the tag they
agreed on. Only multiplication differs, widening to sixty-four bits and
rounding as the fixed-point library does. A mixed pair would have to be
promoted, which can fail, so it goes back to the interpreter.

    fixed in cells    46.2 -> 3.7 ms   12.5x   was not compiled at all

That benchmark says "in cells" because of a limit worth stating plainly.
Globals are hoisted into machine registers on the way in and guarded to be
whole numbers, so a loop holding fixed-point values in globals still fails
that guard and hands itself back. Which is to say the common shape,

    x = x + 0.5

does not compile yet, while the same arithmetic between array cells does,
because only the loop counter beside it has to be hoisted. Letting globals
carry fixed-point values means either keeping them in memory for such
loops or compiling once their kinds are known. That is the next decision,
and it is a real one rather than more of the same.

Under -d:fixedChecks the interpreter asserts on fixed-point overflow where
this wraps, so fixed point is not modelled in that build at all, and CI
runs the safety tests there as well to say so.

Verified on arm64 macOS and x86-64 Linux, under -d:release, -d:danger and
-d:fixedChecks, as a wasm32 build under node, and as a Windows cross
build. 47 AArch64 and 46 x86-64 encodings check against the assembler.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Five operations that were interpreted for no reason beyond not having
been written. A fixed-point literal is a constant the compiler already
knows, so it is stored straight into a slot rather than read from a table
at run time. Host data is copied entire, like an array cell, because the
interpreter copies it entire and neither needs to know its kind.

    host data in a loop    15.1 -> 0.74 ms   20.2x   was not compiled
    fixed literal in cells 20.7 -> 1.15 ms   18.0x   was not compiled

Division of one fixed-point number by another is left alone on purpose.
Its rounding normalises the signs and then corrects a floor, and getting
that subtly wrong would show up as two machines disagreeing rather than as
a failure, which is the one kind of bug this VM cannot afford. It is one
and a half percent of the raytracer and can wait for its own change.

Thirty-nine of sixty-one opcodes now compile. Of what the raytracer
executes, calls are all that is left in any quantity, and they are the
reason none of its loops compile at all: one operation a region does not
model refuses the whole region.

Verified on arm64 macOS and x86-64 Linux, under -d:release, -d:danger and
-d:fixedChecks, as a wasm32 build under node, and as a Windows cross
build. The four budget-escape shapes remain bounded.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Compiling a call means taking in the body of whatever is called, and that
body sits elsewhere in the code. A region can therefore no longer be a
pair of bounds. Membership and the block map are now stated as a set of
offsets, and whether a branch stays inside is a question asked of that
set rather than of a range.

Nothing changes yet: the set is still exactly the loop, every test passes
unaltered, and the counting, branching and nested benchmarks hold at
65.6x, 26.3x and 32.8x. This is the shape the next change needs.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
A call moves the register base, pushes a frame, and changes which routine
is running. Compiled code will have to do all three where the interpreter
can read the result, so the context now carries the frame array, the
argument array, the whole register file, where the current frame starts,
how deep the calls are, and which routine is running. Those three
counters are read back on every return from compiled code, alongside the
budgets and the offset.

The frame layout is checked before anything is compiled, the same way the
value layout already was. Compiled code will push frames the interpreter
then reads, so a layout it does not recognise has to mean nothing is
compiled rather than a corrupted call stack.

Nothing is compiled differently yet. Every test passes unchanged and the
counting and nested benchmarks hold at 60.6x and 30.5x.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
A region now takes in the body of whatever it calls, and the bodies of
whatever those call in turn, so a loop that calls a subroutine compiles
instead of being refused whole.

Nothing is called in the machine's sense. The frame goes into the
interpreter's own array, the base and depth and current routine into its
own fields, and control simply jumps to the callee's compiled code. Every
piece of state a call moves therefore stays exactly where the interpreter
looks for it, which is what keeps leaving part way through a call free:
there is nothing held anywhere else to put back.

Returning is the one thing that cannot be decided when compiling, since a
subroutine called from two places has two places to go back to. A table
from offset to address answers it, and every offset the region did not
compile points at the one stub that hands control back, so returning into
interpreted code needs no test of its own.

Both ceilings are the interpreter's: the call depth, and the register
stack running past its limit. Either sends the offset back so the
interpreter refuses it with its own message.

    loop with a call     14.4 -> 1.65 ms   8.7x   was not compiled at all
    two levels of call   12.9 -> 1.89 ms   6.8x   was not compiled at all
    recursion            17.1 -> 2.48 ms   6.9x   was not compiled at all

Six shapes in six now compile, where it was one in five three changes
ago. Modest speedups next to a tight loop's sixty, because a call still
writes a frame and clears the callee's slots exactly as the interpreter
does; what changed is that these run at all.

Finding calls also turned up a real bug. touchedSlots read jump-if-zero's
operands the wrong way round, taking the branch target for the slot. The
target was being bounds checked as if it were a slot, which refused
perfectly good loops, and the slot itself was never checked at all. That
is the kind of gap the rest of this pass exists to close, so it is now
refused by name in the safety tests.

A loop whose region takes in a callee cannot settle its budget by how far
a pass has got, since that stops meaning anything once the region spans
more than the loop. Such regions charge block by block.

Calls are AArch64 only so far. Elsewhere a region containing one is not
compiled, exactly as before they were written anywhere, and x86-64 and
the Windows cross build are unaffected.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Host code is the one thing compiled code leaves for code it did not
write, so it is the one place that has to keep the calling convention.
The context and the frame base go on the stack because a callee may use
those registers; the budgets, the globals base and the hoisted globals
are ones a callee must leave alone.

Host code can also refuse, and it cannot refuse by raising back through a
frame that nothing described. A trampoline catches whatever it raised and
reports it as an answer, and the offset goes back with the refusal
already made, so the function is not called a second time on the way out.
A host function that raises now stops the script with its own message
from either path.

    a host call in a loop   3.3x   was not compiled at all

Division to a fixed-point answer follows the library exactly: the signs
are put right first, half the divisor is added, and the truncating divide
is corrected back to a floor. Either kind may appear on either side, so
both are widened first, and a whole number too large to widen goes back
rather than being approximated.

    divide between cells    6.2x   was not compiled at all
    divide by negatives     5.5x   was not compiled at all

Fifty-two AArch64 encodings now, the new ones being the sign extension
and the two shifts the widening needs, all checked against the assembler.

The raytracer still compiles nothing, and now for one reason only: its
loops touch more globals than there are registers to hold them. Every
operation in them is modelled. Letting a region keep its globals in
memory when there are too many to hoist is the next change, and the last
one the raytracer is waiting on.

Both are AArch64 only. Elsewhere a region containing either is not
compiled, and x86-64, the wasm build and the Windows cross build are
unaffected.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@treeform treeform changed the title Compile hot integer loops to machine code Compile whole programs to machine code Sep 23, 2026
@treeform
treeform marked this pull request as draft September 25, 2026 15:02
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant