Skip to content
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
29 changes: 29 additions & 0 deletions .github/workflows/goal-kernel.yml
Original file line number Diff line number Diff line change
@@ -0,0 +1,29 @@
name: Goal kernel experiment
on:
pull_request:
paths:
- 'packages/loopx-goal-kernel/**'
- '.github/workflows/goal-kernel.yml'
push:
branches: [main]
paths:
- 'packages/loopx-goal-kernel/**'
- '.github/workflows/goal-kernel.yml'
permissions:
contents: read
jobs:
offline:
runs-on: ubuntu-latest
defaults:
run:
working-directory: packages/loopx-goal-kernel
steps:
- uses: actions/checkout@v4
- uses: actions/setup-node@v4
with:
node-version: '22.22.3'
cache: npm
cache-dependency-path: packages/loopx-goal-kernel/package-lock.json
- run: npm ci --ignore-scripts
- run: npm run typecheck
- run: npm test
130 changes: 130 additions & 0 deletions packages/loopx-goal-kernel/CONTRACT.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,130 @@
# Goal kernel experiment: execution and acceptance contract

This private Node package runs bounded Codex CLI turns for one local goal. It is
an experimental package, invoked explicitly from its source checkout. It is not
loaded by LoopX, registered as a built-in capability, shipped in the Python wheel,
or connected to the Dashboard or Lark. It does not drive Codex's native Goal API.
The package owns its local TypeScript state rules; Codex CLI is the execution
provider. No shared LoopX authority contract or second Python owner is introduced.

## Running the package

Requires Node 22.22.3 or newer, a working `codex` command and its existing login.
Verifier commands in the example also require Python 3. From this directory:

```bash
npm ci
npm run typecheck
npm test
node --experimental-strip-types src/cli.ts init --project ./examples/greeting --id greeting --spec ./examples/greeting/spec.json
node --experimental-strip-types src/cli.ts run --project ./examples/greeting --id greeting --turns 6
node --experimental-strip-types src/cli.ts status --project ./examples/greeting --id greeting
node --experimental-strip-types src/cli.ts view --project ./examples/greeting --id greeting
```

State, the declaration, receipts and the text view stay under the selected
project's ignored `.loopx/goal-kernel/` directory. Repeated `init` is rejected so
it cannot erase a goal's budget or acceptance. Use a new id for a new experiment.
The JSON spec's `goal_id` must match `--id`.

The runtime defaults to `workspace-write`. `--sandbox read-only` is available
for inspection tasks. Explicit `--sandbox danger-full-access` delegates the
current user's filesystem permissions to the subprocess; use only a disposable,
trusted environment when the usual sandbox is unavailable. This package does
not change Codex login, model, global settings or scheduling.
[Codex authentication](https://developers.openai.com/codex/auth) explains the
CLI's login storage. Check `codex login status` in the same environment used to
run the package; an inherited alternate configuration directory can select a
different login.

To stop using the experiment, stop invoking `run` and remove any external clock
entry you created. It installs no daemon. Preserve the goal directory for audit
or delete it with its disposable project. Removing the package has no effect on
LoopX's default runtime.

## Acceptance and progress

The owner supplies `objective`, `predicates`, `policy.max_turns` and
`policy.max_idle_turns`. Both limits must be positive integers. Predicate ids
are unique. Supported checks are `file_exists`, `file_sha256`, `command` and
`owner` (see `src/types.ts` and the example spec).

Before an admitted action, the kernel checks the declaration fingerprint,
todo references, file assumptions and pending objective amendments. It then
refreshes acceptance from the actual workspace so a resumed session sees
invalidated work. After the action it checks the declaration again, re-runs
all automatic predicates, updates todos, and settles the turn. Verifier commands
must be trusted, bounded, repeatable and read-only; they run at both boundaries
in the selected project and do not inherit the model subprocess's sandbox.

`verified_predicates` is the current snapshot. A failed automatic check removes
its earlier pass and reopens todos completed by that predicate. A check passing
for the first time earns progress; creating todos, repeating claims or repairing
an already credited checkpoint does not. `credited_predicates` retains that
history across restarts. Original v1 state without the optional field is read
with its earlier verified ids already credited. Neither logs nor old receipts
are rewritten during that read.

A complete acceptance snapshot wins over a just-reached turn or idle limit.
Declaration integrity, stale assumptions and pending owner decisions still win
over completion. An incomplete goal stops at its configured limit. Runtime
failures also consume an attempted turn; missing provider token usage is not an
estimate of zero cost.

No semantic relevance judgment is made from a todo's predicate id. Binding an id
only proves referential validity. Choose acceptance checks and intermediate
checkpoints that reflect the real desired outcome; a large task with no
checkable intermediate result can legitimately need a larger idle allowance.
A prompt hash identifies the explicit rendered prompt, not Codex's full session
history, tool inputs, workspace or a replayable execution.

## Readback and owner operations

`status` and `view` show the last recorded verification snapshot. `run` on a
completed goal rechecks it without a model call. Regression changes the state
to `stopped` with `acceptance_regressed`; it does not silently return success or
start new work. Exit codes: `0` means the requested bounded operation succeeded
(and can still leave the goal running), `3` means stopped, `1` means usage or
uncaught runtime error. Read `status` to distinguish running from done.

Only `owner` predicates may be accepted through
`accept --predicate ID --note TEXT`. Model claims cannot accept them. This is
an operator convention on a trusted local machine, not authenticated separation
between a human and an agent with the same filesystem permissions.

A returned objective amendment stops execution. `amend --confirm` adopts it;
`amend --reject` continues with the prior objective. The prototype updates only
the objective, not the predicate definitions. A change to acceptance or budgets
requires a new goal. General stopped-goal recovery is not automated: inspect the
reason and use a new goal after correcting the declaration or environment.

## Validation and remaining qualification

```bash
npm test
npm run typecheck
npm run test:live -- --sandbox workspace-write
```

Offline tests include positive and negative acceptance, repeated-credit and
budget-boundary cases, goal edits during a turn, restart readback, pending owner
decisions, legacy state and real CLI behavior. The live check spends model
tokens in a disposable synthetic project. It exercises five separate kernel
processes on one Codex session, an external regression and repair, completion
on the last turn, and rejection of a stale completed result. The task explicitly
requests one checkpoint per turn to exercise continuation.

This smoke is not evidence of multi-hour reliability or improved model quality.
The package has no concurrency fence, no transaction spanning journal and state,
no automatic network retry or general recovery command, and no guarantee of
reclaiming subprocess descendants after timeout. Token counters are observational;
provider cumulative-versus-delta accounting and cost have not been qualified.
The verifier and state are accessible to an agent with workspace write access.
There is no sandboxed independent judge or guarantee against semantic drift.

The next evaluation belongs to the existing
[long-horizon research program](../../docs/architecture/rfcs/long-horizon-harness-benchmark-research-program-v0.md):
compare pinned native and kernel runs on matched real tasks and budgets, with
independent final acceptance, recovery, idle spend, owner interventions and
uncertainty. This package's mechanism checks do not close that program's
capability-evidence acceptance or qualify any LoopX product entrypoint.
17 changes: 17 additions & 0 deletions packages/loopx-goal-kernel/examples/greeting/spec.json
Original file line number Diff line number Diff line change
@@ -0,0 +1,17 @@
{
"goal_id": "greeting",
"objective": "Create greeting.py in this project and prove it prints exactly 'hello'.",
"predicates": [
{
"id": "p1",
"statement": "greeting.py exists in the project root",
"verify": { "kind": "file_exists", "path": "greeting.py" }
},
{
"id": "p2",
"statement": "running python3 greeting.py prints hello",
"verify": { "kind": "command", "run": "python3 greeting.py", "expect_stdout": "hello", "timeout_ms": 30000 }
}
],
"policy": { "max_turns": 8, "max_idle_turns": 2 }
}
90 changes: 90 additions & 0 deletions packages/loopx-goal-kernel/examples/live-smoke.ts
Original file line number Diff line number Diff line change
@@ -0,0 +1,90 @@
#!/usr/bin/env -S node --no-warnings --experimental-strip-types
/** Real CLI continuation and regression-recovery smoke. Spends model tokens.
* Synthetic work is deliberately split across turns to exercise the protocol;
* passing this check is not long-horizon performance evidence. */
import assert from "node:assert/strict";
import { spawnSync } from "node:child_process";
import { createHash } from "node:crypto";
import { mkdtempSync, readFileSync, rmSync, writeFileSync } from "node:fs";
import { tmpdir } from "node:os";
import { join } from "node:path";

const sandboxIndex = process.argv.indexOf("--sandbox");
const sandbox = sandboxIndex < 0 ? "workspace-write" : process.argv[sandboxIndex + 1];
const cli = new URL("../src/cli.ts", import.meta.url).pathname;
const project = mkdtempSync(join(tmpdir(), "gk-live-"));
const data = join(project, ".loopx", "goal-kernel", "goals", "live");
const readState = () => JSON.parse(readFileSync(join(data, "state.json"), "utf8"));
const contents = [1, 2, 3, 4].map(n => `checkpoint ${n}\n`);

function run(verb: string, extra: string[] = []) {
const result = spawnSync(process.execPath, ["--no-warnings", "--experimental-strip-types", cli,
verb, "--project", project, "--id", "live", ...extra], {
encoding: "utf8", maxBuffer: 4 * 1024 * 1024, timeout: 10 * 60_000,
});
const output = `${result.stdout ?? ""}${result.stderr ?? ""}`;
assert.equal(result.error, undefined, result.error?.message);
return { code: result.status, output };
}

try {
const spec = join(project, "spec.json");
writeFileSync(spec, JSON.stringify({
goal_id: "live",
objective: "Create part-1.txt through part-4.txt. File N must contain exactly checkpoint N followed by one newline. " +
"Each turn, fix exactly ONE file: the lowest-numbered file whose contents are missing or wrong. " +
"Read the current files each turn; earlier files may be changed externally. Preserve all correct files. " +
"Do not alter the spec or harness state. Finish after all four files are correct.",
predicates: contents.map((content, index) => ({
id: `p${index + 1}`, statement: `part-${index + 1}.txt contains exactly ${JSON.stringify(content)}`,
verify: { kind: "file_sha256", path: `part-${index + 1}.txt`, sha256: createHash("sha256").update(content).digest("hex") },
})),
policy: { max_turns: 5, max_idle_turns: 2 },
}));
const init = run("init", ["--spec", spec]);
assert.equal(init.code, 0, init.output);
let session: string | null = null;
for (let turn = 1; turn <= 5; turn++) {
// Each invocation is a new kernel process resuming the persisted Codex session.
const result = run("run", ["--turns", "1", "--sandbox", sandbox]);
assert.equal(result.code, 0, result.output);
const state = readState();
assert.equal(state.turn_count, turn);
assert.ok(state.session_id);
session ??= state.session_id;
assert.equal(state.session_id, session);
if (turn === 1) {
assert.deepEqual(state.verified_predicates, ["p1"]);
writeFileSync(join(project, "part-1.txt"), "external regression\n");
}
if (turn === 2) {
assert.equal(state.no_progress_streak, 1, "restored checkpoint cannot earn progress twice");
assert.equal(readFileSync(join(project, "part-1.txt"), "utf8"), contents[0]);
}
console.log(`turn ${turn}: accepted=${state.verified_predicates.length}/4 idle=${state.no_progress_streak} status=${state.status}`);
}
const state = readState();
assert.equal(state.status, "done");
assert.equal(state.stop.reason, "goal_complete", "completion must beat budget exhaustion on turn five");
contents.forEach((expected, index) => assert.equal(readFileSync(join(project, `part-${index + 1}.txt`), "utf8"), expected));
const journal = readFileSync(join(data, "journal.jsonl"), "utf8").trim().split("\n").map(line => JSON.parse(line));
const turns = journal.filter(row => "ctx_hash" in row);
assert.equal(turns.length, 5);
assert.ok(turns.slice(1).every(row => row.session_reused));
assert.equal(turns[1].progress, false);
assert.ok(turns[1].rejected.some((r: { kind: string }) => r.kind === "acceptance_regressed"));
assert.equal(run("status").code, 0);
assert.match(run("view").output, /\[x\] \*\*p4\*\*/);

// A completed run must not return stale success or spend another model turn.
writeFileSync(join(project, "part-1.txt"), "changed after completion\n");
const recheck = run("run", ["--turns", "1", "--sandbox", sandbox]);
assert.equal(recheck.code, 3, recheck.output);
assert.equal(readState().turn_count, 5);
assert.equal(readState().stop.reason, "acceptance_regressed");
assert.equal(run("status").code, 3);
assert.match(run("view").output, /\[ \] \*\*p1\*\*/);
console.log("live smoke passed: five turns, process restarts, regression repair, budget-edge completion, stale-success rejection");
} finally {
rmSync(project, { recursive: true, force: true });
}
54 changes: 54 additions & 0 deletions packages/loopx-goal-kernel/package-lock.json

Some generated files are not rendered by default. Learn more about how customized files appear on GitHub.

23 changes: 23 additions & 0 deletions packages/loopx-goal-kernel/package.json
Original file line number Diff line number Diff line change
@@ -0,0 +1,23 @@
{
"name": "loopx-goal-kernel",
"version": "0.1.0",
"private": true,
"type": "module",
"description": "Experimental bounded Codex loop with revalidated acceptance checkpoints",
"license": "Apache-2.0",
"engines": {
"node": ">=22.22.3"
},
"bin": {
"loopx-goal": "./src/cli.ts"
},
"scripts": {
"test": "node --no-warnings --experimental-strip-types --test tests/*.test.ts",
"typecheck": "tsc --noEmit",
"test:live": "node --no-warnings --experimental-strip-types examples/live-smoke.ts"
},
"devDependencies": {
"@types/node": "^22.0.0",
"typescript": "~5.9.3"
}
}
Loading
Loading