Skip to content

state_exec/ariadne: render-pass park safety — per-step cap + read-lane guard (§13.8) - #477

Merged
transfix merged 1 commit into
masterfrom
feat/ari-render-park-budget
Sep 29, 2026
Merged

transfix merged 1 commit into
masterfrom
feat/ari-render-park-budget

Conversation

@transfix

Copy link
Copy Markdown
Owner

Keeps a render-pass state_exec program that awaits/parks from breaking the render thread — the last piece of the §13.8 async/await arc (transparent http-get #474, generic fetch #475 both merged).

This is a pivot. The adversarial review of the first attempt (a per-activation time budget that would kill a runaway resident) found it genuinely unsafe, so this ships the safe subset and defers the hard part:

  • HIGH — a buffered input-event burst drains through the msg-recv fast-path sharing one activation budget → a legitimate input resident is killed, and ensure_input_residents never resubmits it → input silently dies.
  • HIGH — a sleep(0)/self-wake spin resets its budget every wake → escapes the kill.

Distinguishing a genuine external wait from a self-spin needs more than a wake count, so killing a persistent spinner is left as future work behind a safe backstop.

What ships (two layers)

  1. Read-lane guard. The reactive lane (visible_when/computed) is structurally park-proof — a private idle scheduler that's never pumped, and an allowlist that excludes await/msg-recv/http-get/fetch. So a park verb in a predicate is denied (unbound → fail-safe hidden, render returns), not parked into a lane that would silent-nil. Asserted over both msg-recv and await.
  2. Per-STEP cap. Residents/actions run in drain() off the draw walk, bounded by the pump's per-drain cap (a per-node DSL loop degrades but never hangs). The gap the pump can't cover is a single evaluator step that loops in place (a runaway native builtin never returns → the pump's between-step check never fires). Fix: process::max_step_time — a per-step deadline armed fresh each step via the same eval_deadline_guard an action's max_time uses, so it aborts an in-step loop + kills the process, yet never accumulates and so never kills a long-lived resident. ensure_resident wires it at 50 ms.

Also fixes a latent double-count: a parking intrinsic already books its running slice into accumulated_time + sets waiting, but execute_process_step re-added the same slice unconditionally — now guarded on status == running (mirroring elapsed_time()).

Deferred (documented)

Killing a persistently spinning resident (vs the pump bounding it each drain) — needs real-wait-vs-self-spin accounting the review showed the naive per-activation budget botches.

Tests

PerStepCapLeavesHealthyWorkUntouched (no false-kill + plumbs through), ParkVerbInPredicateIsDeniedNotHung (msg-recv and await denied), RunawayTickResidentIsBoundedAndSchedulerRecovers (a runaway on:tick is pump-bounded, drain() returns, a replacement resident fires). Full regression across async/scheduler/intrinsics/integration green; 0/25 flaky on the timing cases.

…e guard (§13.8)

Keep a render-pass state_exec program that awaits/parks from breaking the render
thread, in two layers.

(1) The reactive READ lane (visible_when/computed) is structurally park-proof: it
evaluates on a private idle scheduler that is never pumped, and its allowlist
excludes await/msg-recv/http-get/fetch, so a park verb in a predicate is DENIED
(unbound -> fail-safe hidden, render returns) rather than parking a lane that would
silent-nil. Asserted by a unit test over both msg-recv and await.

(2) Residents + actions run in drain(), off the draw walk, bounded by the pump's
per-drain cap, so a per-node DSL loop degrades but never hangs. The one gap the pump
can't cover is a single evaluator STEP that loops in place (a runaway native builtin
never returns, so the pump's between-step time check never fires). Fix: a per-STEP
wall-clock cap (process::max_step_time), armed fresh each step via the SAME
eval_deadline_guard an action's max_time uses — so it aborts an in-step loop + kills
the process, but being fresh each step it never accumulates and so never kills a
long-lived resident. ensure_resident wires it at 50ms; a healthy on:tick/on_key body
is far under it.

Also fixes a latent double-count: a parking intrinsic (sleep/yield-frame/
receive_message) already books its running slice into accumulated_time and sets
status=waiting, but execute_process_step then re-added the same slice
unconditionally. Guard the accrual on status==running (mirroring
process::elapsed_time()'s own guard).

NOT killing a PERSISTENTLY spinning resident (vs the pump bounding it each drain) is
deferred. The naive per-activation budget that would do so was found unsafe by
adversarial review: a buffered input-event burst drains through the msg-recv
fast-path sharing one budget -> a legitimate input resident is killed (and
ensure_input_residents never resubmits it -> input silently dies), and a sleep(0)/
self-wake spin resets its budget every wake -> escapes. Distinguishing a genuine
external wait from a self-spin needs more than a wake count, so it is left as future
work behind the safe per-step backstop.

Tests: PerStepCapLeavesHealthyWorkUntouched (the cap doesn't false-kill normal
work + plumbs through); ParkVerbInPredicateIsDeniedNotHung (msg-recv AND await
denied); RunawayTickResidentIsBoundedAndSchedulerRecovers (a runaway on:tick is
pump-bounded, drain returns, a replacement resident fires). Full regression across
the async/scheduler/intrinsics/integration suites green; 0/25 flaky on the timing
cases. Completes the §13.8 async/await arc.
@transfix
transfix merged commit 55c671b into master Sep 29, 2026
13 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant