Skip to content

fix(cli): failed updates and the scope fixture name their cause - #2589

Merged
DeusData merged 2 commits into
mainfrom
fix/cli-flake-diagnostics
Oct 11, 2026
Merged

DeusData merged 2 commits into
mainfrom
fix/cli-flake-diagnostics

Conversation

@DeusData

@DeusData DeusData commented Oct 11, 2026 •

Copy link
Copy Markdown
Owner

What

Two failures that said nothing about their cause now name it.

  1. update on macOS names the step that failed, and the OS error. The signed-candidate preparation folded every step into one boolean. A full /tmp therefore printed the same line as a broken codesign: signed update candidate preparation failed. Now, for example:

    error: signed update candidate preparation failed while writing the unsigned candidate: No space left on device
    

    The staging functions in activation_transaction.c keep the errno of the failing write (or sync, close, read) through their cleanup, so a caller can name it. On the non-macOS path the existing failed to stage verified update line gains the OS error (POSIX only; Windows reports through GetLastError). A failed fingerprint of the staged file now gets its own message, instead of ...: activation transaction completed.

  2. The cli_scope test fixture names its failing setup step, and no timeout decides readiness. cli_scope_fixture_start returned false silently from any of its setup steps, and a 90 s deadline decided the wait for the host child's readiness. It now records and prints the step, with the OS error where one applies, and the host child prints which of its own steps failed before it answers E. The readiness wait has no deadline: the child answers R or E, or the pipe reaches EOF when it dies. A wedged child is left to the harness's per-suite ceiling.

Why

Two cli tests failed once each on 2026-10-09 in local multi-suite runs. The logs couldn't say why:

  • cli_update_agent_configs_finish_before_guard_release failed at ASSERT(replacement_kept), after signed update candidate preparation failed;
  • cli_install_skip_binary_unchanged_in_host_namespace_quiesces_nothing failed at ASSERT(ready), with nothing else printed.

The attribution:

  • First failure: a disk that filled up mid-test. It was reproduced line for line by putting /tmp on a small volume and filling it between the test's two updates. A multi-GB index was running on a nearly full disk at the time.
  • Second failure: not reproduced, in 30 runs of the same suite order (with and without full CPU load, down to 0 MB free, under ulimit -n 256).

So the changes here make the next occurrence say what happened. The product already behaved safely: it refused the update and kept the old executable.

Tests

  • cli_update_names_the_failing_staging_step_and_os_error: a test seam fails the update's first staging write with ENOSPC. The test asserts the step and No space left on device in the output, and that the old executable stays. It fails on main's cli.c (strstr(output, strerror(ENOSPC)) is NULL) and passes with this change. It is skipped on Windows (GetLastError, not errno).
  • cli_scope_fixture_names_its_failing_setup_step: a tag naming a missing directory makes the fixture's mkdtemp fail. The test asserts the recorded step.
  • Real input, with /tmp on a 4 GB sparse volume (a local-only redirect) and the disk filled between the two updates:
    • 900 MB filled: ...failed while writing the unsigned candidate: No space left on device;
    • 600 MB filled: ...failed while signing the candidate ad hoc.

Checks

  • Affected suites (daemon_version daemon_application activation_transaction cli, macOS ASan): 461 passed, 0 failed.
  • make lint-ci clean; cppcheck 2.20 (CI's version, from the lint image) clean on the two changed .c files.
  • Local 3-OS ladder on this exact tree:
    • macOS (scripts/test.sh): 9,135 passed, 0 failed, 11 skipped (177 suites); cli 356/0 including both new tests
    • Linux arm64 container (run.sh test): 8,973 passed, 0 failed, 10 skipped (173 suites); cli 357/0 including both new tests
    • Windows arm64 VM (win.sh test, 173 suites; correction after the merge: the list was taken from tests/test_main.c, which leaves out the opt-in cs_lsp_bench, incremental, py_lsp_bench and py_lsp_scale suites, none of which touch this change. The follow-up's Windows leg runs all 177): 8,794 passed, 0 tests failed, 91 skipped (the new update test is among the skips, by design). The runner's count of 4 failures is four suite names I passed that exist only in coverage builds (coverage_*, compiled out on Windows); each unknown name counts as one failure. Windows guards (CI's TEST_SEAMS=1 build): 8/8 green.

On macOS the signed-candidate preparation folded every step into one
boolean, so each failure printed the same line: "signed update candidate
preparation failed". A full /tmp read exactly like a broken codesign. That
is what a cli test hit on 2026-10-09, while a multi-GB index was filling a
nearly full disk.

Each step now records why it failed, and the update prints, for example:

  error: signed update candidate preparation failed while writing the
  unsigned candidate: No space left on device

The staging functions in activation_transaction.c keep the errno of the
failing write (or sync, close, read) through their cleanup, so a caller can
name it. The non-macOS path adds the OS error to its existing "failed to
stage verified update" line (POSIX only; Windows reports through
GetLastError). A failed fingerprint of the staged file now gets its own
message instead of "...: activation transaction completed".

A test seam makes the next staging write fail with a chosen errno. The new
test cli_update_names_the_failing_staging_step_and_os_error fails without
the cli.c change and passes with it, and checks that the old executable
stays in place. Also checked on real input: with /tmp on a small volume
that fills between two updates, the update prints the line above.

Signed-off-by: Martin Vogel <martin.vogel.tech@gmail.com>
…es readiness

cli_install_skip_binary_unchanged_in_host_namespace_quiesces_nothing
failed once on 2026-10-09 at ASSERT(ready), with nothing else in the log.
cli_scope_fixture_start returned false silently from any of its setup
steps, and a 90 s deadline decided the wait for the host child's
readiness.

The fixture now records and prints the step that failed: the fixture
directory, the host directories, the runner's own image fingerprint, the
pipes, the fork, the host's start, or the client connect, with the OS error
where one applies. The host child prints which of its own steps failed
before it answers 'E'.

The readiness wait has no deadline any more, because no timeout may decide
a test. The child answers 'R' or 'E', or the pipe reaches EOF when it dies.
A child that wedges is left to the harness's per-suite ceiling.

New test: cli_scope_fixture_names_its_failing_setup_step.

Signed-off-by: Martin Vogel <martin.vogel.tech@gmail.com>
@DeusData
DeusData merged commit f216a10 into main Oct 11, 2026
36 checks passed
neil1123-vip pushed a commit to neil1123-vip/codebase-memory-mcp that referenced this pull request Oct 11, 2026
…ronment

Follow-up to DeusData#2589.

No timeout decides a cli_scope test any more:
- The status probe and the parent's client connect repeat an attempt that
  reaches a live but mute endpoint (a loaded runner answering slowly). A
  reply or a definite refusal decides. Each attempt keeps the API's finite
  5 s budget. This replaces the 15 s status budget and the single 5 s
  connect.
- The host child is reaped without a deadline (it used to get SIGKILL after
  60 s).
- The host child joins its version cohort with UINT64_MAX, which waits for
  locks but answers a conflicting holder at once (it used to have a 45 s
  deadline).
A wedged child is left to the harness's per-suite ceiling.

cli_scope_fixture_finish restored HOME, CBM_CACHE_DIR and SHELL even when
start had failed before saving them, which unset them for every later test
in the suite. It now restores only what start saved.

New tests:
- cli_scope_fixture_finish_keeps_env_after_an_early_failure fails on main
  at ASSERT(home_kept);
- cli_scope_retries_a_mute_endpoint_until_it_answers checks the retry rule
  on scripted attempt outcomes.

Signed-off-by: Martin Vogel <martin.vogel.tech@gmail.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant