Repository navigation
fix(cli): failed updates and the scope fixture name their cause - #2589
Merged
Merged
Conversation
On macOS the signed-candidate preparation folded every step into one boolean, so each failure printed the same line: "signed update candidate preparation failed". A full /tmp read exactly like a broken codesign. That is what a cli test hit on 2026-10-09, while a multi-GB index was filling a nearly full disk. Each step now records why it failed, and the update prints, for example: error: signed update candidate preparation failed while writing the unsigned candidate: No space left on device The staging functions in activation_transaction.c keep the errno of the failing write (or sync, close, read) through their cleanup, so a caller can name it. The non-macOS path adds the OS error to its existing "failed to stage verified update" line (POSIX only; Windows reports through GetLastError). A failed fingerprint of the staged file now gets its own message instead of "...: activation transaction completed". A test seam makes the next staging write fail with a chosen errno. The new test cli_update_names_the_failing_staging_step_and_os_error fails without the cli.c change and passes with it, and checks that the old executable stays in place. Also checked on real input: with /tmp on a small volume that fills between two updates, the update prints the line above. Signed-off-by: Martin Vogel <martin.vogel.tech@gmail.com>
…es readiness cli_install_skip_binary_unchanged_in_host_namespace_quiesces_nothing failed once on 2026-10-09 at ASSERT(ready), with nothing else in the log. cli_scope_fixture_start returned false silently from any of its setup steps, and a 90 s deadline decided the wait for the host child's readiness. The fixture now records and prints the step that failed: the fixture directory, the host directories, the runner's own image fingerprint, the pipes, the fork, the host's start, or the client connect, with the OS error where one applies. The host child prints which of its own steps failed before it answers 'E'. The readiness wait has no deadline any more, because no timeout may decide a test. The child answers 'R' or 'E', or the pipe reaches EOF when it dies. A child that wedges is left to the harness's per-suite ceiling. New test: cli_scope_fixture_names_its_failing_setup_step. Signed-off-by: Martin Vogel <martin.vogel.tech@gmail.com>
neil1123-vip
pushed a commit
to neil1123-vip/codebase-memory-mcp
that referenced
this pull request
Oct 11, 2026
…ronment Follow-up to DeusData#2589. No timeout decides a cli_scope test any more: - The status probe and the parent's client connect repeat an attempt that reaches a live but mute endpoint (a loaded runner answering slowly). A reply or a definite refusal decides. Each attempt keeps the API's finite 5 s budget. This replaces the 15 s status budget and the single 5 s connect. - The host child is reaped without a deadline (it used to get SIGKILL after 60 s). - The host child joins its version cohort with UINT64_MAX, which waits for locks but answers a conflicting holder at once (it used to have a 45 s deadline). A wedged child is left to the harness's per-suite ceiling. cli_scope_fixture_finish restored HOME, CBM_CACHE_DIR and SHELL even when start had failed before saving them, which unset them for every later test in the suite. It now restores only what start saved. New tests: - cli_scope_fixture_finish_keeps_env_after_an_early_failure fails on main at ASSERT(home_kept); - cli_scope_retries_a_mute_endpoint_until_it_answers checks the retry rule on scripted attempt outcomes. Signed-off-by: Martin Vogel <martin.vogel.tech@gmail.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What
Two failures that said nothing about their cause now name it.
updateon macOS names the step that failed, and the OS error. The signed-candidate preparation folded every step into one boolean. A full/tmptherefore printed the same line as a broken codesign:signed update candidate preparation failed. Now, for example:The staging functions in
activation_transaction.ckeep the errno of the failing write (or sync, close, read) through their cleanup, so a caller can name it. On the non-macOS path the existingfailed to stage verified updateline gains the OS error (POSIX only; Windows reports through GetLastError). A failed fingerprint of the staged file now gets its own message, instead of...: activation transaction completed.The
cli_scopetest fixture names its failing setup step, and no timeout decides readiness.cli_scope_fixture_startreturned false silently from any of its setup steps, and a 90 s deadline decided the wait for the host child's readiness. It now records and prints the step, with the OS error where one applies, and the host child prints which of its own steps failed before it answersE. The readiness wait has no deadline: the child answersRorE, or the pipe reaches EOF when it dies. A wedged child is left to the harness's per-suite ceiling.Why
Two
clitests failed once each on 2026-10-09 in local multi-suite runs. The logs couldn't say why:cli_update_agent_configs_finish_before_guard_releasefailed atASSERT(replacement_kept), aftersigned update candidate preparation failed;cli_install_skip_binary_unchanged_in_host_namespace_quiesces_nothingfailed atASSERT(ready), with nothing else printed.The attribution:
/tmpon a small volume and filling it between the test's two updates. A multi-GB index was running on a nearly full disk at the time.ulimit -n 256).So the changes here make the next occurrence say what happened. The product already behaved safely: it refused the update and kept the old executable.
Tests
cli_update_names_the_failing_staging_step_and_os_error: a test seam fails the update's first staging write withENOSPC. The test asserts the step andNo space left on devicein the output, and that the old executable stays. It fails onmain'scli.c(strstr(output, strerror(ENOSPC)) is NULL) and passes with this change. It is skipped on Windows (GetLastError, not errno).cli_scope_fixture_names_its_failing_setup_step: a tag naming a missing directory makes the fixture's mkdtemp fail. The test asserts the recorded step./tmpon a 4 GB sparse volume (a local-only redirect) and the disk filled between the two updates:...failed while writing the unsigned candidate: No space left on device;...failed while signing the candidate ad hoc.Checks
daemon_version daemon_application activation_transaction cli, macOS ASan): 461 passed, 0 failed.make lint-ciclean; cppcheck 2.20 (CI's version, from the lint image) clean on the two changed.cfiles.scripts/test.sh): 9,135 passed, 0 failed, 11 skipped (177 suites);cli356/0 including both new testsrun.sh test): 8,973 passed, 0 failed, 10 skipped (173 suites);cli357/0 including both new testswin.sh test, 173 suites; correction after the merge: the list was taken fromtests/test_main.c, which leaves out the opt-incs_lsp_bench,incremental,py_lsp_benchandpy_lsp_scalesuites, none of which touch this change. The follow-up's Windows leg runs all 177): 8,794 passed, 0 tests failed, 91 skipped (the new update test is among the skips, by design). The runner's count of 4 failures is four suite names I passed that exist only in coverage builds (coverage_*, compiled out on Windows); each unknown name counts as one failure. Windows guards (CI'sTEST_SEAMS=1build): 8/8 green.