Skip to content

GotA neural policies (stacked on #78) - #76

Open
daveey wants to merge 29 commits into
Metta-AI:polyworld-neural-tierfrom
daveey:daveey/neural-policies-upstream
Open

daveey wants to merge 29 commits into
Metta-AI:polyworld-neural-tierfrom
daveey:daveey/neural-policies-upstream

Conversation

@daveey

@daveey daveey commented Sep 25, 2026 •

Copy link
Copy Markdown

Stacked on #78 (the shared neural tier), which is stacked on #72. The base is #78's branch. Review the commits after Merge upstream/main into the neural tier; the merge brings main (mailboxes, towers) into the stack, because #72 predates them.

This PR adds neural policies to Gods of the Arena: a package seat runs a small FP32 network next to its BASIC program. It also adds a native training environment with a C ABI. GotA is the first client of the tier in #78: the package loader, model format, op budget, telemetry line, recurrent-state lifecycle and neural BASIC functions all live there. Plain .bas submissions are unchanged, and matches without neural seats replay byte-identically to main (a new CI step checks this).

Specs: examples/gods_of_the_arena/neural_basic.md (GotA) and docs/neural-policies.md (tier). C ABI: examples/gods_of_the_arena/native_env.h.

What's here

  • Contract (neural_contract.nim): observation v1 (1407 floats, ego-centric), action v1 (heads verb 8, target 25, point 49, ability 4, item 6), the demonstration encoder, and the GotA package keys: goal, decoder.defer_script, decoder.mask_empty_targets + mask_mode (a MaskMode enum).
  • Seats (neural_host_hooks.nim): GotaContract is built with the tier's initNeuralContract, and NeuralSeat derives from the tier's NeuralBrain. Seat kinds: hosted package, learner, capture, override, shadow, defer and mask. Every neural seat gets the same per-tick BASIC budget as a plain .bas seat (20,000 instructions, 50,000 work units); the network runs on its own 4,000,000-op budget.
  • Native training env (native_env.nim / .h): a thread-safe C ABI for PufferLib-style trainers. Any learner/script mix, BC labels, DAgger shadows, defer and override seats, action masks, package seats, and gota_net_* for bit-exact actor checks.
    • Opaque GotaEnv / GotaNet handles with magic checks, gota_mask_size, and a contract-checked gota_net_load.
    • gota_last_error is thread-local and set on every failure path, including gota_net_infer and gota_set_policy_script (which keeps the compile message).
    • Package rejection reasons and compile messages survive gota_reset. After gota_step returns -3 the caller must gota_reset before stepping again (documented).
    • Per-world vision blockers (no camp sight leaks between worlds in one process) and a lock around gota_create.
    • Package seats keep their manifest goal, and read the match length when they decide, not at install.
    • The observation frame is frozen before neural seats observe, in the same order as the hosted server, so a native package seat plays exactly like a hosted one.
  • Python: coworld/gota/runtime/neural_package.py is a thin wrapper over the shared validator. tools/native_env.py raises NativeEnvError on every negative return code and checks 1407/92/187/24 against the library.
  • sim.nim / bots.nim: behaviour-preserving refactors. deferConsult runs inside runHeroVm's error handling. HeroVm.neural stays RootRef, so sim.nim doesn't import float code.
  • Tools: parity_tier.py (byte-identity battery between two libraries), parity_upstream.sh (no-neural parity with main), test_defer.py, test_mask.py, test_package_goal.py, test_thread_isolation.py, test_native_env.py, test_native_concurrency.py, mapping_ceiling.py, canary.py, replay_diff.nim.
  • CI (build.yml): tests/test_gota_neural.nim, a golden-hash no-neural parity check against main (tests/gota_golden.sh, base/puller/rusher, hashes recorded from main), and a compile-only build of the native library.

Intentional behaviour changes (each with its own proof)

Everything else is byte-identical to the previous implementation of this PR (see Evidence).

change before after proof
An invalid override on a defer seat (a non-defer verb that decodes to nothing, e.g. an empty target slot) absorbed the script for the whole window; the hero idled counts as invalid and defers for that window The L2 battery (always-defer + mixed defer/override packages) differs on 19/20 seeds; with only this change reverted it is identical on 20/20. test_defer.py d: a defer seat fed only invalid overrides (62,110 decisions over 10 full matches) plays byte-identical to plain base.bas on 10/10, with 0 overrides; the previous library matches on 0/10
Telemetry line of package seats ... ticks=<t> inferences=<n> adds decisions=<n> invalid=<n>, plus defer=<n> override=<n> on defer seats canary.py on the -d:coworld server, full 28,800-tick matches, 5 package seats each: argmax w128 and defer w64. Every line matches the format, every seat exits 0, and the replays re-simulate with 0 mismatches
A neural package that fails validation on the hosted server compiled a fake broken source; failure.json said "BASIC compilation failed for player slot N" failure.json says "neural package rejected for player slot N"; the reason is in the seat log New case in coworld/tools/test_runtime.nim: passes here. The same test against the previous implementation fails with "BASIC compilation failed for player slot 0"
An exception raised while a defer seat consults its network escaped runHeroVm disables that VM like any other BASIC error No match in any battery raises there; all batteries are unchanged
ABI error reporting gota_net_infer left gota_last_error empty; an invalid policy script lost its message; gota_reset could drop a package rejection reason all three reported New error-path checks in test_native_env.py: pass here; the previous library fails exactly these three

Evidence

Linux, Nim 2.2.10. "Previous implementation" is this PR's code before it moved onto the tier.

check result
build.yml test list, test_gota_neural.nim, golden hashes, native library build, coworld/tools/test_runtime.nim all pass
GitHub Actions build.yml at this PR's head (ubuntu, macOS, Windows; dispatched, since the workflow only runs on PRs into main) pass: run 36199384968. The golden check now strips a CRLF line end, which a Windows checkout adds
Battery vs the previous implementation (parity_tier.py), 20 seeds per lineup, full 28,800 ticks. Compares the state hash after every step, the replay bytes, seat stats, and defer/override counts L0 no neural seats 20/20. L1 hosted packages (w64 and w128 argmax, w64 sampling) 20/20. L2 defer 1/20 (the invalid-override change; 20/20 with it reverted). L3 action mask, conditional and static, 20/20. L4 trainer seats (two learners, a shadow, an override seat, capture, a package seat) 20/20. L5 always-defer only 20/20
No neural seats vs main (parity_upstream.sh), 20 seeds, full length replays byte-identical 20/20; the native library's replay is verified by main's binary 20/20
Native ABI seat == hosted package seat test_native_env.py: a package seat equals a learner driven by gota_net_infer at w64/w128/w256. test_package_goal.py: native package seat == hosted server, 20/20 (10 seeds x manifest goal / default goal): state hash and action streams, final hash, and the native replay verified by the binary. test_defer.py b: ABI defer seat == package defer seat on a mixed net, 10/10
Defer test_defer.py a: an always-defer net equals plain base.bas, ABI and package, 20/20 (hashes + replay). c: seats without the option match the previous implementation 10/10
Action mask ABI mask == host mask, 20/20 seeds (134,257 decisions: mask bytes, heads, per-step hashes). Conditional mask: 0 invalid choices in 10/10 games
Package validation GotA wrapper and Nim loader agree on 16 GotA corruption cases (goal, defer and mask keys, float integers, temperature: true, a non-finite goal, bad JSON, truncation), plus the tier's 50-case suite
Error paths (test_native_env.py) gota_net_infer on a NaN observation returns -2 with "nonfinite neural input"; an invalid policy script returns 1 with its compile message; a package rejection reason and a script compile message are still reported after gota_reset; the binding raises NativeEnvError on the -2 reset. All pass here; the previous library fails the first three
Threads 16 handles on 4 threads: no vision leak between worlds (16/16 distinct tower histories). 400 concurrent creates on 16 threads: no mismatch. 12 handles on 6 threads with handle migration: serial hashes

Not in this PR

  • Coworld release. Accepting neural packages in the hosted league needs a release built from this stack.

Caveats

  • Thread scaling is about 5x at 16 threads, against about 8x with separate processes (atomicArc refcount traffic).
  • The action contract is approximate for some commands. mapping_ceiling.py (100 seeds, in neural_basic.md's Evidence) shows base.bas loses nothing when routed through it.

🤖 Generated with Claude Code

Review fixes

(After first rebasing/merging onto the updated polyworld-neural-tier with #78's review fixes.)

fix evidence
gota_reset returned -2 on a seat compile failure but never set lastError native_env.nim's resetEnv: when the result is -2, builds "seat compilation failed: seat N: <reason>; seat M: <reason>..." from every seat whose status.code == 2 and sets the thread-local lastError. tools/native_env.py's NativeEnvError now stores the message as .message. Verified in test_native_env.py's test_error_paths: "reset with a failed seat raises NativeEnvError" passes.
A package with a valid manifest but a broken policy.bas was reported as "BASIC compilation failed", not "neural package rejected" bots.nim's installPackageSeat now compiles policy.bas itself (not through compilePlayer, whose own failPlayer call ran unconditionally with the wrong message) and reports "neural package rejected for player slot N: policy.bas failed to compile: <reason>" through the same failPlayer path. New case in coworld/tools/test_runtime.nim, next to the existing package-rejection case: builds a valid GotA package (real manifest + model.bin) with a syntactically broken policy.bas and checks the new message plus failPlayer's status.json output (reason, exit code, other slots untouched). Run against a real -d:coworld gota binary built on the cluster: passes.
tools/native_env.py's set_package/set_script returned a status code silently on failure Both now take check=True and raise NativeEnvError (carrying the status message) on any nonzero return by default; check=False opts back into the old return-code behavior for the three call sites that intentionally probe a rejection (test_native_env.py's validation-suite loop and the two explicit "-> 1"/"-> 2" error-path checks). All other callers already asserted == 0, so their behavior on success is unchanged, and now fails loudly instead of via a bare assert if a package they expect to succeed doesn't.
Width support carried into GotA coworld/gota/runtime/neural_package.py's WIDTHS already inherits from the shared tier validator (fixed in #78), so no edit was needed there. examples/gods_of_the_arena/neural_basic.md: width list (64/128/256/384/512) and op counts (218,496 / 486,144 / 1,168,896 / 2,048,256 / 3,124,224), plus the w512 model size note. tools/test_native_env.py: ported test_widths (ACCEPT_WIDTHS/REJECT_WIDTHS, 15 neighboring widths rejected by both the Python validator and the raw Nim loader, and each accepted width also checked through a hosted seat), and widened test_package_parity's width list to include w384/w512.

Invariants (with evidence)

  • No behaviour change for existing w64-w256 packages: same evidence as polyworld: shared neural policy tier #78 (unchanged buffer contents/bounds for h <= 256); test_native_env.py's "package seat == ABI-driven seat" passes at w64/w128/w256 with the same op counts as before (218,496 / 486,144 / 1,168,896).
  • Plain .bas seats stay byte-identical: bash tests/gota_golden.sh (no-neural parity against main) passes 3/3 seeds on the built library at this PR's head. test_native_env.py's existing suites (labels, capture identity, shadow, goals, rewards) are unaffected and pass.
  • Full CI test list passes: tests/test_gota_neural.nim, the native training library build (--app:lib -d:headless -d:gotaTrainingStats --mm:atomicArc --threads:on -d:useMalloc -u:nimTypeNames), tests/gota_golden.sh (3/3), and examples/gods_of_the_arena/tools/test_native_env.py (quick mode, ALL PASS including the new width and error-path checks) all pass on the metta cluster against a fully nimby.lock-pinned dependency graph. coworld/tools/test_runtime.nim's gota case (including the new package-policy-compile-failure case) passes against a real -d:coworld gota binary built on the cluster. GitHub Actions dispatched at this PR's head (7863481): pass on ubuntu-latest, macos-latest and windows-latest — https://github.com/daveey/polyworld/actions/runs/36342890904

Note on test robustness

test_widths's width-640 case (deliberately far outside ActorWidths) is rejected by both validators for different, equally correct reasons: the package-level Python validator hits model.bin's own size cap first (a ZIP entry's declared size is checked before it's parsed), while the raw gota_net_load call used in this test (no package wrapper) only ever reaches the dimensions check. The assertion accepts either failure mode; the 13 other neighboring widths (1, 32, 63, 65, 96, 192, 255, 257, 320, 383, 385, 448, 511, 513) are all rejected by the dimensions check on both sides, as expected.

daveey and others added 15 commits September 25, 2026 15:20
Gods of the Arena's neural seats, rebuilt as the first client of the
polyworld neural tier (src/polyworld/neural_*.nim). Behaviour is
byte-identical to the previous GotA-only implementation.

- neural_contract.nim: observation v1, action v1, the demonstration
  encoder, and the GotA package options (goal, defer_script, a
  MaskMode enum for mask_empty_targets / mask_mode).
- neural_host_hooks.nim: GotaContract built with initNeuralContract;
  NeuralSeat derives from NeuralBrain; the observe/decode/mask
  callbacks; learner, capture, override, shadow, defer and mask seats.
  The enum value NeuralPackage is now NeuralHosted.
- bots.nim / sim.nim: package seats through parseGotaPackage; a
  rejected package ends a hosted episode with its real reason
  (failPlayer); deferConsult runs inside runHeroVm's error handling.
- native_env.nim / .h: the training ABI with opaque handle types and
  magic checks, gota_mask_size, contract-checked gota_net_load,
  lastError on every failure path, rejection reasons kept across
  reset.
- coworld/gota/runtime/neural_package.py: a thin GotA wrapper over
  the shared validator; tools/native_env.py raises on every negative
  return code and asserts its sizes against the library.
- tests: test_gota_neural.nim (contract identity, GotA corruption
  cases), gota_golden.sh (no-neural final hashes recorded from main),
  and a compile check of the native library in CI.
- tools/parity_tier.py: the byte-identity battery.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A defer seat whose network picks an override that decodes to nothing
(e.g. an empty target slot) used to absorb the script's commands for
the whole decision window, so the hero idled. It now counts as
invalid and defers for that window. Always-defer seats and seats
without decoder.defer_script are unchanged.

The GotA telemetry line gains decisions= and invalid= counts, plus
defer= and override= for defer seats.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ead-safe gota_create, isolation tests

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…l seats get the plain per-tick BASIC budget; metrics use each seat's own limit; lastError thread-local doc; goal parity test

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…on RNG by kind)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…serve (hosted order; fixes native != hosted package play); replay_diff tool; goal/hosted parity test covers default goals

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ats saw maxTicks 0: time features and telemetry wrong)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…e with the package reason)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…in base.bas)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…er/override in defer mode)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…, policy script message, reasons kept across reset); binding types gota_set_policy_script

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@daveey daveey changed the title GotA: neural policy packages + native training env (builds on #72) GotA neural policies (stacked on #78) Sep 25, 2026
@daveey
daveey changed the base branch from main to polyworld-neural-tier September 25, 2026 22:36
@daveey
daveey force-pushed the daveey/neural-policies-upstream branch from 1898578 to feeca58 Compare September 25, 2026 22:36
daveey and others added 2 commits September 25, 2026 16:02
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
# Conflicts:
#	.github/workflows/build.yml
#	coworld/tools/test_runtime.nim
daveey and others added 3 commits September 27, 2026 11:44
…idth carry, native_env.py check=True

- examples/gods_of_the_arena/native_env.nim: resetEnv now sets the
  thread-local lastError to a message naming every seat that failed to
  compile and its reason before returning -2 from gota_reset (it only
  updated per-seat status before). tools/native_env.py's NativeEnvError
  now carries that message as `.message`.
- examples/gods_of_the_arena/bots.nim: installPackageSeat compiles a
  package's policy.bas outside compilePlayer's own failPlayer call, so a
  valid manifest with a broken policy.bas is reported as "neural package
  rejected for player slot N: policy.bas failed to compile: <reason>"
  through failPlayer, instead of the generic "BASIC compilation failed for
  player slot N" (compilePlayer's own message, which ran unconditionally
  regardless of why compilation was being attempted).
- coworld/tools/test_runtime.nim: a new case next to the existing
  package-rejection one builds a valid GotA package (manifest + model.bin)
  with a syntactically broken policy.bas and checks the new message and
  failPlayer's status.json output; failureMessage is now a startsWith
  check so it can assert a prefix when the compiler's own error text
  varies. Verified against a real -d:coworld gota binary on the cluster.
- examples/gods_of_the_arena/tools/native_env.py: set_package/set_script
  take `check=True` and raise NativeEnvError (carrying the status message)
  on any nonzero return by default; `check=False` keeps the old
  return-code behavior for the tests that intentionally probe rejection.
  Updated the three callers that rely on inspecting a rejection.
- Width support carried into GotA: examples/gods_of_the_arena/neural_basic.md
  (widths, op counts, w512 model size) and
  examples/gods_of_the_arena/tools/test_native_env.py (ported test_widths:
  w64-w512 accepted, 15 neighboring widths rejected by both the Python
  validator and the raw Nim loader, plus the hosted seat). GotA's own
  Python validator (coworld/gota/runtime/neural_package.py) already
  inherits WIDTHS from the shared tier validator fixed in Metta-AI#78, so no
  separate edit was needed there.

Rebases/merges cleanly onto the updated polyworld-neural-tier (Metta-AI#78 fixes).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants