Skip to content

Release 3.4.3: external ownership, truthful evidence, and lifecycle fixes - #43

Merged
ajhcs merged 41 commits into
mainfrom
codex/coengineer-autonomous-ownership-20260910
Sep 11, 2026
Merged

ajhcs merged 41 commits into
mainfrom
codex/coengineer-autonomous-ownership-20260910

Conversation

@ajhcs

@ajhcs ajhcs commented Sep 10, 2026 •

Copy link
Copy Markdown
Owner

Merged as 2907aa41baa9bfbcc4c7342b83d929a4ae8544aa and published as v3.4.3, with partial qualification disclosed below. Global installation is verified against the exact release files. A setup-prefix defect required the direct-script workaround documented in the release and issue #44.

External agents previously handed work back to Codex after short implementation bursts, leaving Codex to perform corrections and reconstruct evidence. This change keeps bounded implementation and corrections with the external owner while Astra retains consequential decisions and final acceptance. Supervisor deadline extensions now govern active ACP turns.

The five public tools are preserved. Candidate-bound task.revision, explicit role preferences, durable lineage and duplicate admission checks use the existing ledger and decision-card components. Unknown usage remains unknown. The release also includes a one-provider first outcome, issue forms, a PR template, support links, contributor tasks, and a scoped roadmap.

Evaluation tooling includes three explicitly retrospective repository tasks, four offline fixtures, frozen checks, a seeded 24-trial protocol, and primary accounting for native helpers, reasoning, compaction, corrections, failures and cumulative snapshots. These facilities do not establish savings.

Final accepted candidate

  • Commit: c0aada7b76a1b4a42b5de20878293268d915321b
  • Tree: e8bf87014d372355329bf4a6386887a5fb542a54
  • Published 3.4.2 baseline: dede188029aff117c60e9a8c4299cc0ab0838be9
  • Exact local gate: 17 stages passed; 2,495 tests passed; one reproducible build; 174.265 seconds. Run 20260911T195734Z-f5c8b5e49e34, unchanged plan fingerprint a4bf902e9937e2e883c2a0c6592ac184078e0b92ee4302ec7be4d3be1048b660.
  • Independent candidate CI passed. Final branch CI and final PR CI also passed on this exact head.

Cursor-owned cleanup corrections went through Grok review, same-owner revisions, reviewer verification and Astra acceptance. Publication checks then exposed a further close/finalization race. A broad attempted fix was rejected. Astra froze a deterministic regression on 179e17f; Grok implemented the minimal correction in c0aada7; Cursor independently verified the exact commit, unchanged test, failing parent control, 10 runtime tests, 33 driver/bridge tests and five stress repetitions. Astra accepted the corrected candidate. Both provider tasks completed with normal cleanup and no active boundary or lock. Cursor's initial wait lost its connection after the report; durable inspection recovered completion without submitting another prompt.

The qualification report retains the ownership loop, restart/card/ledger evidence, comparison table and approved setup accounting. The publication correction record retains later failures and controls. Raw sessions, credentials and provider receipts stay private.

Publication with partial qualification

The maintainer explicitly requested merge to main, a new GitHub release and local global installation after the qualification gaps were reported. Automated correctness gates pass; full qualification is incomplete. The release notes disclose this disposition and preserve the existing release requirements.

There are zero valid matched comparison results, and no measured benefit is claimed. The first attempt remains invalid because inherited host instructions contaminated its context and normal CLI consent did not complete; no provider prompt was dispatched. Collection is paused pending context-isolation and consent preflight. Authenticated clean-environment agent onboarding and restart persistence, refreshed native-host provider acceptance, and Desktop wait/recovery checks remain pending. Paid API/Cloud spend is $0; those routes remain pending until an enforceable bound fits the approved $25 aggregate ceiling. Setup costs are separate from trial results.

Automatic balance routing, broader OS coverage and human usability validation remain roadmap work. Showcase preparation does not submit a listing or bypass the documented local-MCP limitation.

ajhcs and others added 30 commits September 10, 2026 21:26
…stay truthful.

Disable the fixed inner turn timer, race prompts against the worker AbortSignal, and stop promoting TimeoutError after partial output into end_turn success.

Co-authored-by: Cursor <cursoragent@cursor.com>
…cellation.

Replace the module-global active-turn signal with AsyncLocalStorage and harden late prompt settlement so cancel A leaves B live without unhandled rejections.

Co-authored-by: Cursor <cursoragent@cursor.com>
Simple run_request preferences fill omitted provider/model by role while
exact assignment selections win, and unknown preferred providers return
attention instead of a silent substitute. The existing task tool can
derive a fresh correction assignment from a completed, clean producer
without replaying an active or uncertain task. Compact receipts now
include a machine-derived coordination packet for the next action.
…visor dispatch.

Exact assignment selections now win over unused or unknown role preferences. Public admission receipts expose request_idempotency_key and per-assignment HEAD, completed candidates next-action to review, and reviseRun requires a fresh clean workspace plus original assignment context.
…rrection lineage.

Correction prompts keep original assignment constraints or reject overflow instead of clipping, reviseRun requires an authoritative completed task bound to the producer lane, admission persists compact correction lineage at creation, and coordination exposes revision only for proven clean writers without inferring it from provider prose.
Provide issue/PR templates, support and starter-task docs, a first-outcome example, and a compatibility-first quickstart so ownership work is easier to discover and improve.

Co-authored-by: Cursor <cursoragent@cursor.com>
…ling.

Reuse the usage ledger and decision card for shareable admission summaries, keep local completion distinct from Codex or PR acceptance, and add frozen benchmark cases with an offline trial comparison command.
Admission persists original/root lineage and one child follow per producer, identical retries stay idempotent, different feedback follows that child instead of branching, and an exhausted round rejects before provider dispatch so a new bounded assignment stays a deliberate later submit.
Replace the no-op version fixture with a failing summarize-checks stub, align chat/revision wording, and drop worker-facing prose from SUPPORT and the roadmap.

Co-authored-by: Cursor <cursoragent@cursor.com>
Keep local and run-result summaries coherent for failed, unfinal, and
unknown states, bind Codex acceptance to the exact candidate, and
degrade maximum-size usage and result projections inside existing
byte caps without dropping identity or uncertainty.
Materialize frozen cases to a real Git base, bind trial source identity
and input digests, separate wall elapsed from attempt sums, keep mixed
provider totals grouped, and label synthetic fixtures as unverified.
…roof.

Share one assignment-result reduction so failure and cancel outrank active
lanes, and remove the unused producer-from-receipt helper.
Remove contradictory in-place wording so revision children implement and
commit in the current assigned directory, treating original paths as
lineage rather than a navigation target.
…tation.

Bump Co-Engineer version surfaces, install commands, release notes, and evaluation
requirements ahead of the freeze, and wire provider-free trial-usage and
qualification unit stages into the gate and CI without claiming those suites ran.

Co-authored-by: Cursor <cursoragent@cursor.com>
Restore a usable published 3.4.2 install beside a concrete PR43 candidate path with a distinct marketplace so clean-agent reviewers can follow public commands without colliding identities or dead tag checkout.

Co-authored-by: Cursor <cursoragent@cursor.com>
…ording.

Align package/runtime exact-version checks with unreleased 3.4.3 metadata while keeping public install assertions on published 3.4.2, and clarify that [] is empty input while empty stdin is invalid JSON.

Co-authored-by: Cursor <cursoragent@cursor.com>
Keep codex-co-engineer as the shipped marketplace contract and document clean-install or supported remove/re-add paths instead of renaming the candidate marketplace.

Co-authored-by: Cursor <cursoragent@cursor.com>
Replace the 404 /build/cli link and document subscription login with optional --device-auth so first-outcome Grok setup needs no API key.

Co-authored-by: Cursor <cursoragent@cursor.com>
…r rules.

Measure Astra own-output versus published 3.4.2, restore the distinct local marketplace wrapper path without renaming the shipped manifest, and stop treating remove/re-add or historical notes as sufficient host-gate proof.

Co-authored-by: Cursor <cursoragent@cursor.com>
Emit analyzer-compatible trial records with breakdown digests while failing closed on incomplete attribution and privacy leaks.

Co-authored-by: Cursor <cursoragent@cursor.com>
Fix optional cache counters, phase endpoints, window reconciliation,
session binding, acceptance emission, CLI status, and path safety while
preserving prior synthetic coverage.

Co-authored-by: Cursor <cursoragent@cursor.com>
Bind helpers by thread id and parent graph, parse sub_agent_activity, keep
measured usage when acceptance is unknown, and match exact host settings.

Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Prepare three unrun retrospective tasks bound to real pre-fix SHAs, measured input digests, and deterministic materialized base commits, plus an operator schedule with seed 43.
Worker inputs now materialize pre-fix snapshots from dede188 and 3131f9a
with frozen independent checks. Direct delegation is required. Candidate
SHA, Astra model, and host settings bind only in an external execution
manifest. The offline evaluator uses task-level medians, not pooled ratios.
Bind per-case provider/model routes, require candidate tree and the Astra
host model, freeze hyphen-only trial ids, and reuse comparator aggregation.
Overhead is candidate/direct native output; turnaround is the median of
per-trial wall ratios. Astra output is counted once from host-usage-report
by_model rows without inventing provider tokens.
Reject reports whose bound trial, evidence digest, attempt coverage, or
per-model counters disagree with the measured trial, and keep partial
Astra numbers without treating corrupted counters as complete coverage.
ajhcs and others added 4 commits September 11, 2026 15:23
Skip by_model sum equality when a primary attempt counter is unknown so
importer-shaped incomplete reports keep observed model totals. Shuffle
approach positions within seed-43 case/rep groups and record the algorithm.
…URL.

Packaged mirrors cannot resolve repository-only ../release.md; keep requirements wording unchanged while matching the v3.4.2 packaging pattern.

Co-authored-by: Cursor <cursoragent@cursor.com>
Helpers keep an exact child thread_id while shared ancestor session_id is allowed only when manifest parent_id and session_meta parent_thread_id ancestry both prove the chain.

Co-authored-by: Cursor <cursoragent@cursor.com>
String source values like "cli" are not subagent parent links; reuse the binding prepass and drop unused ancestry rechecks.

Co-authored-by: Cursor <cursoragent@cursor.com>
@ajhcs

ajhcs commented Sep 11, 2026 •

Copy link
Copy Markdown
Owner Author

Qualification remains pending. This is an interim evidence report for public candidate b19c746c65f920e41d4e94a7a22444d523c063b0, tree 6ab71c23dd2131707e32e3248f0ea4b221bcbbad, against published 3.4.2 baseline dede188029aff117c60e9a8c4299cc0ab0838be9. No measured savings claim is supported. PR #43 stays in draft.

Independent ownership demonstration

Real accounting work exercised Cursor implementation → Grok independent review → same-owner task.revision → Grok verification → Astra acceptance for integration. The reviewed producer was c5ff185609e78d0dfa446686e65bd50a182dbd17; its corrected commit was b37ebc43cd3dd6a1d8152d752a1e9ada30e29c72. This was substantive accounting correction, not an intentionally inserted defect. Twenty-two focused regressions and a real native-accounting probe passed at that stage. Original and corrected commits, feedback, revision lineage, decisive checks, and terminal cleanup are retained privately.

Two fresh source-MCP processes returned identical lineage, coordination evidence, result card, and usage projection. The existing card correctly said Review needed, with codex_accepted: false; Astra's subsequent acceptance is a separate review decision. The ledger recorded one submission and 749 seconds for that revision, while provider tokens and subscription balance remained unknown. The worktree was clean and its owned process boundary was inactive and unlocked.

Further real-host verification exposed normal native helpers sharing the parent's session identity while retaining their own thread identity. Cursor corrected the importer through two more same-owner revisions, ending at 3324e6470ec8d842f3c5ff302b801354c1950a5c; Grok independently verified all 27 importer regressions, and Astra checked the actual private primary accounting before integration. Duplicate metadata traversal and unused state were removed. Grok-owned qualification accounting received Cursor review and same-owner corrections; the final review passed 39 focused checks.

A subsequent real cleanup failure exercised another complete correction loop: Cursor implementation, a failing Grok independent check, two same-owner revisions, final Grok verification, and Astra acceptance of b19c746. The correction evidence retains the rejected candidate, deterministic historical control, checks, cleanup and limits.

This is development source-MCP evidence. It does not satisfy the remaining demonstration on the fully qualified, refreshed native Desktop host.

Correctness and installation

Gate or observation Result Practical limit
Exact local release gate, candidate b19c746 Passed 17 stages, 2,494 tests, one build; 181.795 seconds; no first-pass failure Does not override a failing CI or live-host gate
PR CI run Passed Separate execution of the candidate
Push CI run Passed Earlier failed CI and the independently rejected first correction remain retained
Installed candidate inventory All 476 tracked plugin files match the exact candidate; two fresh Codex app-server discovery processes agree on the enabled installation The existing Desktop MCP connection still reports 3.4.2; a real Desktop restart and acceptance remain pending
Clean-environment agent prerequisites Ubuntu 24.04.4, system and user systemd, documented prerequisites installed in 6m34s Grok device sign-in expired without authentication; first outcome and restart persistence remain untested
Paid API/Cloud routes $0 dispatched No enforceable aggregate cap established; Cursor Cloud, Muse, and Ox Alpha remain pending under the $25 ceiling

The gate total includes Co-Engineer 2,297; comparison 16; importer 27; qualification 23; Cursor compatibility 122; and vendor provenance/reproducibility 9. Plan fingerprint: a4bf902e9937e2e883c2a0c6592ac184078e0b92ee4302ec7be4d3be1048b660.

Failures are retained. The earlier push CI run failed descendant cleanup and timed out; the first correction passed its local gate but failed independent stress checking and was rejected. The preceding candidate 3ab2cffe33f9056315ce664e8d60ef5c99710717 failed its first release gate because packaged release notes linked to a repository-only file (2,295/2,296 Co-Engineer tests passed). Cursor corrected the links before the new candidate's gate. Earlier failed review, consent-capability, wrong-workspace, and partial-output attempts remain recorded; none is counted as successful qualification. A live-host reload attempt also failed during client initialization before issuing any reload.

Comparison status and setup cost

The three retrospective repository inputs, independent acceptance checks, seed 43, matched provider/model settings, one-hour total deadlines, and three-round correction limit were frozen before collection. All three historical inputs fail their checks; private corrected-source controls pass. Reference solutions and prior trial outputs are excluded from worker contexts. The first four-arm matched group must validate collection before the remaining 20 trials.

Collection is paused and inconclusive: one invalid initial attempt, zero valid matched trials. The first published-3.4.2 execution read inherited host maintenance notes containing later implementation details, violating the required exclusion of reference context. It was stopped and retained. Normal CLI consent requests also returned cancelled; no external provider prompt was sent, and the planned run was cancelled. Frozen task/check files remained unchanged and failed as expected for the uncorrected input. Parent/coordinator accounting and the interrupted-generation limitation are retained privately. A separate context preflight showed that an instruction-byte-limit setting did not remove the global host guidance in the installed CLI. The remaining scheduled executions are paused until context isolation and real native consent are verified; no replacement attempt has been silently substituted.

Approach Valid completed / planned Independently accepted Native output per accepted result Turnaround
Native Codex, including permitted native helpers 0 / 6 Not measured Unknown Unknown
Published 3.4.2 0 / 6 Not measured Unknown Unknown
Candidate 3.4.3 0 / 6 Not measured Unknown Unknown
Direct supported-CLI delegation 0 / 6 Not measured Unknown Unknown

Thresholds remain fixed: candidate 6/6 acceptance; median task-level native output per accepted result at least 50% below native Codex and 25% below 3.4.2; Astra's own output below 3.4.2; median paired turnaround at most 2× native Codex; candidate native overhead at most 25% above direct delegation. Missing primary accounting makes the result inconclusive. Failed attempts, helpers, corrections, review, and coordination count; reasoning and compaction count once. Input/cache figures stay separate. Tokens are never translated into subscription percentages.

A setup-only snapshot, 2026-09-11 12:51:50.550–17:17:49.722 UTC, reconciled parent/helper responses against primary cumulative counters: 367,747 Astra output tokens, including 194,170 reasoning tokens, with four compactions. Input was 56,113,454, of which 54,417,408 was cached input. This is substantial evaluation preparation cost, not a benchmark result or savings evidence. It excludes subsequent preparation and a separate nine-output-token CLI probe. Accounting-only projections preserve source provenance; raw sessions and provider receipts remain private. A differing secondary legacy counter is retained as a diagnostic, not substituted for primary accounting.

Remaining gates include validated collection context and consent, the complete measured cohort, authenticated clean-environment onboarding with unchanged first-outcome checks and restart persistence, refreshed Grok/Cursor Local/Cloud/Muse/Ox Alpha acceptance, and documented Desktop wait/recovery checks. Clean-agent evidence will be labeled accordingly; human usability validation is later work.

The OpenAI showcase can describe externally owned work, independent correction, Astra acceptance, and transparent evidence limits. Preparing this story does not submit a listing or bypass the documented local-MCP limitation.

… CI.

Fail-soft /proc discovery, always signal the agent pid, and reap fixture
children after close so a missed process-group kill cannot strand the suite.

Co-authored-by: Cursor <cursoragent@cursor.com>
@ajhcs

ajhcs commented Sep 11, 2026 •

Copy link
Copy Markdown
Owner Author

The cleanup correction is now independently accepted and published as b19c746c65f920e41d4e94a7a22444d523c063b0 (tree 6ab71c23dd2131707e32e3248f0ea4b221bcbbad). PR #43 remains a draft because measured benefit and live-host gates are still outstanding.

The first cleanup correction was rejected during qualification. Candidate 2fec429191f12575b1aac34064e04b647d7d9f15 (tree 2904d627a96dced121bb8792ff985a63df4549a5) passed a new exact local gate: 17 stages, 2,493 tests, one build, 177.368 seconds. Cursor also reported 30 successful stress repetitions.

Independent Grok checking then reproduced the original failure on repetition 18: the existing descendant-cleanup assertion failed because a descendant remained alive after close() returned. The retained test log confirms the failure. Grok reached its review deadline before a final verdict and its owned boundary cleaned up normally. Both the passing gate and the failed independent check are retained.

Cursor corrected the actual ordering defect through same-owner task.revision: upstream resolved turn.result before retaining the persistent client, so an immediate close() could observe an empty client map. Commit 99d73d289272004d6d48bc3c4c52274637f66188 defers settlement until finalization and removes the speculative process-signaling changes. Astra then required the regression to assert in the synchronous continuation after the result; Cursor supplied b19c746.

Grok independently verified the complete corrected commit: 9 focused tests passed, all 12 stress repetitions passed (24 checks), and the identical committed regression failed against unchanged 2fec429 runtime assets. Astra checked the exact diff, retained logs, identical negative-control test bytes, and normal inactive/unlocked producer and reviewer boundaries before accepting. Normal same-run inspection recovered the terminal result after a transient completion observation; this remains source-MCP development evidence, not refreshed Desktop acceptance.

The new exact-candidate gate passed all 17 stages and 2,494 tests, with one build and no first-pass failures, in 181.795 seconds. Both PR CI and push CI passed. The earlier failed CI, rejected correction, and review timeout remain retained.

All 476 installed plugin files match the corrected candidate before and after two fresh project-scoped app-server discovery operations. The existing Desktop connection still reports 3.4.2; native acceptance remains pending.

The new candidate and evaluation inputs are frozen outside tracked files. A subsequent initial collection attempt was stopped and retained as invalid because inherited host notes contaminated its context and normal CLI consent did not complete. No external provider prompt or paid API/Cloud job was dispatched; zero valid matched comparisons are available, and further collection is paused. Trial accounting includes attributable outer-coordinator work as well as the measured lead, helpers, review, failures, and corrections. All thresholds remain fixed; no savings are claimed.

ajhcs and others added 4 commits September 11, 2026 17:25
Upstream resolved turn.result before finalize retained the persistent client, so runtime.close could observe an empty pending map and skip agent/descendant containment; defer settlement until after finalize and regress that ordering.

Co-authored-by: Cursor <cursoragent@cursor.com>
Post-result readFile awaits let finalize hide the ordering gap against the unrepaired runtime; keep the check in the same sync continuation and capture fixture PIDs with readFileSync for finally cleanup.

Co-authored-by: Cursor <cursoragent@cursor.com>
@ajhcs

ajhcs commented Sep 11, 2026 •

Copy link
Copy Markdown
Owner Author

The maintainer requested merging and publishing 3.4.3 with the previously disclosed qualification gaps. Publication remains queued while a newly observed automated-gate failure is corrected.

The documentation promotion exposed old branding assertions pinned to 3.4.2; those assertions were updated and their six focused checks pass. The next exact-candidate gate, on b7f97a2ad6a0b27bda88164e61d3e81ad8145d61, reached the existing 300-second unit deadline while the fake Grok agent for the unsupported-question bridge remained active. That candidate contains the same runtime as independently reviewed b19c746; earlier passing receipts do not replace this failure.

Cursor completed the third owned revision at public commit 2f191e5601c053de6995adf3c2834f009ec40720 on the review branch, with normal terminal cleanup. It reports 22 focused tests and five stress passes, but did not reproduce the original full-suite timeout in isolation. Grok is independently reviewing the exact commit, including whether the regression exercises the suspected failure and whether the broader cleanup changes are justified. The proposed correction has not been integrated into the release branch or accepted by Astra. Both failed gate receipts are retained. No merge or tag has been created yet; no measured savings claim is added.

Independent review outcome: rejected. Grok verified that the new regression also passes on unchanged b19, found a referenced cap timer delaying process exit by about five seconds, and found that an abort losing to a successful prompt can skip the normal drain of trailing output. Astra rejected 2f191e5; it remains on its review branch and is not integrated. The completed three-round correction chain is preserved.

The retained failed fixture adds a more specific clue: it had recorded failure and ACP closure while its fake agent remained alive. Astra explicitly authorized a separate bounded Grok implementation assignment from the unchanged PR43 preparation commit to investigate that mismatch and produce a deterministic prior-runtime failure with a minimal fix. Cursor will independently review the result. This is a new assignment, not a fourth revision or an automatic replay. Both provider runs completed normal cleanup; the release remains pending.

Deterministic failure now published: 179e17fab06a84a255cf7a96a576655ed5630fff adds an independent regression without changing the runtime. It pauses finalization after the closed-state check, completes close, then resumes finalization. The unchanged runtime retains the live agent and reopens the stored session. The test fails in about 1.3 seconds and cleans the fixture afterward. Its exact contents are frozen for the external implementation. The bounded Grok diagnostic attempt timed out with normal cleanup and is retained; a separately authorized Grok correction now has this concrete new acceptance check. Cursor will independently review the resulting public commit. No main merge, tag or global release replacement has occurred.

Final correction accepted — 2026-09-11

Grok produced c0aada7b76a1b4a42b5de20878293268d915321b, tree e8bf87014d372355329bf4a6386887a5fb542a54. The change preserves the independently frozen test on 179e17f and adds a minimal 25-line runtime overlay, with the bundled asset and manifest regenerated. It prevents retaining a client after close intent and restores the closed session record after a stale finalization save. The rejected broad proposal is not in its ancestry.

Cursor independently accepted this exact commit: the same frozen check fails on the parent with a live retained agent and a reopened record, then passes on the fix. Runtime tests passed 10/10, driver/bridge tests 33/33, and five repeated regression checks passed. Astra accepted the result. Both tasks completed with normal, inactive, unlocked cleanup. A dropped review connection was reconciled from durable completion without another prompt.

The exact release gate passed all 17 stages and 2,495 tests, with one reproducible build in 174.265 seconds, run 20260911T195734Z-f5c8b5e49e34. Candidate CI passed. Earlier failures and the timed-out diagnostic remain retained; the green result applies to this changed candidate.

Publication is proceeding at the maintainer's explicit direction with partial qualification. There are zero valid matched comparison results; clean-agent onboarding, refreshed native-host provider acceptance and Desktop wait/recovery remain incomplete. No savings claim is made, and the existing release requirements remain documented.

Published 3.4.3

PR #43 merged to main at 2907aa41baa9bfbcc4c7342b83d929a4ae8544aa. Its tree is exactly the independently accepted and locally gated e8bf87014d372355329bf4a6386887a5fb542a54. Both final branch CI and PR CI passed.

Codex-Co-Engineer 3.4.3 is released, with the partial qualification explicitly disclosed. The published GitHub source archive contains 664 tracked files; every byte and executable mode matched the tested tree. Archive SHA256: 190e89dd1dbd737388f673e4ac57ac8d13c68f4e37c98cd00e2932fd83befdf4. Global installation verification is complete; see the final record below. This publication does not turn the pending measurements, onboarding or native-host/Desktop qualification into passing evidence.

Global installation and final hosted checks

Main CI and tag CI both passed on released commit 2907aa41baa9bfbcc4c7342b83d929a4ae8544aa.

The global installation now contains exactly the 476 released plugin files, verified byte-for-byte with executable modes. A distinct local marketplace wrapper protects it from older same-name projects. Two fresh Codex app-server processes performed actual discovery for old and current project paths; the release remained enabled and unchanged, and older Co-Engineer identities remained uninstalled/disabled. The server launched from the installed package reported 3.4.3, healthy status, zero active jobs and the exact five-tool catalog. Setup checks passed.

The first installation inventory failed: npm --prefix leaked into the setup subprocess and created dependency files inside the local plugin source. Those extra files were copied into the plugin cache, although the released files were unchanged. The failed inventory was retained, the generated files were removed from the source, setup was invoked directly, and the plugin was removed/re-added through the normal CLI. Issue #44 tracks the defect, and the release notes publish the tested direct-setup workaround.

This installation needed maintainer correction and therefore is not clean-environment onboarding acceptance. Discovery-process persistence is also not a Desktop connection restart or the required native-host provider/wait-recovery qualification. Restart Codex to load the installed version. All previously stated measurement and live-qualification gaps remain open.

Upstream samples closed state then awaits session save before retain; close can finish in that gap, after which finalization overwrites the closed record and keeps the ACP client. Refuse retain after close intent and persist the closed snapshot.
@ajhcs ajhcs changed the title feat: retain external ownership and make Co-Engineer results demonstrable Release 3.4.3: external ownership, truthful evidence, and lifecycle fixes Sep 11, 2026
@ajhcs
ajhcs marked this pull request as ready for review September 11, 2026 20:11
@ajhcs
ajhcs merged commit 2907aa4 into main Sep 11, 2026
3 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant