Skip to content

fix(machine): separated browsers wait for a real extension exchange and never block native reconnects (0.60.1) - #1766

Merged
chronoai-kai merged 3 commits into
mainfrom
fix/machine-context-browser-startup
Oct 5, 2026
Merged

chronoai-kai merged 3 commits into
mainfrom
fix/machine-context-browser-startup

Conversation

@chronoai-kai

Copy link
Copy Markdown
Contributor

Summary

Hotfix for the v0.60.0 Publish Images failures (main run 37231036156 and tag run 37231044043). The arm64 machine image smoke and the X64 Machine Container E2E both failed at context secure navigation: 12413 secure browser response timed out.

Root cause

  1. A freshly provisioned separated browser profile could report the native-host hello before its MV3 extension could service a request. The first navigation then raced extension startup.
  2. Browser::exchange held the connection mutex across socket I/O. During a native-host reconnect, the accept loop could not install the replacement stream, so the first request sat on a wedged pipe until its 20 s timeout.

The Authentication failed: Unexpected message type lines in the same logs were transient authority-reconnect retries that later authenticated. They were not the browser failure.

Fix

  • Separated browsers prove one bounded startup tabs exchange (12 s grace, separate from each operation's 20 s deadline) before the first action, and cache the verification per connection.
  • Transport I/O no longer holds the connection mutex. A connection generation prevents a late exchange from overwriting a replacement stream.
  • Secure-browser operations remain serialized by the runtime's browser lock, so concurrent exchanges cannot interleave.
  • Shared legacy browsers keep their previous readiness check and timing, and liveness is still checked first so a dead browser is repaired immediately.

Validation

  • Implementer:
    • Rust 1.98.1 fmt and workspace clippy.
    • 1,419 CLI unit, 65 CLI integration and 27 machine tests.
    • Reduced-stack browser tests (17).
    • Native arm64 full machine container e2e: 3 consecutive passes with the shipped seccomp profile and 3 with the mount-denying profile, including separated browsers, cross-display/profile/socket denial, login binding, quarantine, authority reconnect and updater checks.
    • New regression tests cover cached startup verification and a replacement connection during an in-flight exchange.
  • Reviewer: kept the liveness short-circuit for legacy repair. This PR's Machine Container E2E and the post-merge Publish Images runs are the acceptance gate.

chrono-kw added 3 commits October 5, 2026 07:12
…nd never block native reconnects

Fresh separated profiles could report the native hello before the MV3
extension could service a request, and Browser::exchange held the connection
mutex across socket I/O so the accept loop could not install a replacement
stream during a reconnect, leaving the first navigation on a wedged pipe until
timeout (v0.60.0 Publish Images context secure navigation 12413).

Separated browsers now prove one bounded startup exchange before the first
action; transport I/O no longer holds the mutex, and a connection generation
prevents a late exchange from overwriting a replacement stream. Shared legacy
browsers keep their readiness check and timing, and secure-browser operations
remain serialized by the runtime's browser lock.
@github-actions

github-actions Bot commented Oct 5, 2026

Copy link
Copy Markdown

📊 Code coverage

Component Lines Threshold Status Δ vs base
Backend (nyxid) 87.07% 73% ✅ n/a
CLI (nyxid-cli) 69.26% 64% ✅ 🔺 +0.03
Frontend (vitest) 72.21% 15% ✅ — 0.00

Gate: line coverage must stay at or above the threshold. Ratchet plan (W21): Backend → 55%, CLI → 50%, Frontend → 30% by quarter end.

@chronoai-kai
chronoai-kai merged commit 5331899 into main Oct 5, 2026
35 of 36 checks passed
chronoai-kai added a commit that referenced this pull request Oct 5, 2026
…nostics on failure (#1769)

Coverage (Backend) and Coverage (Backend Base) lost mongod at the same
cumulative point (~840-960 s, auth_device_service tests) on #1765/#1766 with
~7 GB host memory still free; mongod held ~8 GB when it vanished. Raise the
descriptor limit far above the nextest peak (thousands of per-test databases'
WiredTiger files stay open in one llvm-cov process) and cap the WiredTiger
cache at 2 GB. On failure or cancellation, print the container state
(exit code, OOMKilled), mongod descriptor count/limits, kernel OOM lines and
the server log tail so any recurrence is diagnosable.

Co-authored-by: chrono-kw <chrono-kw@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant