Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
55 changes: 35 additions & 20 deletions evals/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -16,7 +16,7 @@ fixture and produces a scored JSONL record with the full S0–S11 stage funnel.
| `evals/runner/drivers/` | Driver implementations — Tier-0 local/static driver; authoring contract in `drivers/README.md` |
| `evals/scenarios/` | Scenario definitions (`tier1-directory.json`, `tier1-directory-guide-only.json`, `tier1-directory-full.json`, `pre1-directory-proceed.json`, `pre1-noiam-park.json`) |
| `evals/skills-bundle/` | Skill-bundle mount point (v0.4.0 manifest — ten skills in `skills/`) |
| `evals/results/` | JSONL run records (gitignored; `.gitkeep` committed) |
| `evals/results/` | JSONL run records (gitignored; `.gitkeep` committed) + the committed `baseline.json` control-group reference |

## How to run

Expand Down Expand Up @@ -234,19 +234,36 @@ pinned scenarios on a skill PR and fails if the measured pass rate drops below
gate workflow itself is not built in this batch; the committed file plus this
consumption rule is the reference.

**Halt-path status.** The E2E runs that would produce the six records are
blocked (see E2E status below), so no scored runs exist and no
`baseline.json` is committed. The generator is covered by unit smokes
(`evals/runner/baseline.test.ts`); the reference will be published once a
real-tenant driver and the tool surface are available.
**Reference status (CXF-233).** The six-run matrix is no longer blocked: the
E2E ran 2026-09-04/05 on the private squire driver (omp agent harness, model
`together/deepseek-ai/DeepSeek-V4-Flash-0731`, `reasoningEffort: "high"`),
producing 3 scored `none` runs + 3 scored `guide-only` runs, and the generated
`evals/results/baseline.json` is committed as the control-group reference.
Measured result: no full-funnel passes in either arm (`pass_at_3` 0), mean
Comment on lines +237 to +242

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Suggestion: The regression-gate contract above (lines 231–235) still says the gate "fails if the measured pass rate drops below modes.<mode>.pass_rate", but the published reference has pass_rate: 0 for both arms — that rule can never fire, so the committed contract is a no-op gate. This paragraph itself explains why the full-funnel floor is partly structural (S2's PUT leg is unobservable, S8 reads secret-masked); the consumption rule should be updated in the same PR to key on first_pass_rate_mean and/or per_stage.failures, or to state explicitly that pass_rate gating is inert until S2/S8 are instrumented. (confidence: high)

first-pass rate 0.47 (`none`) / 0.61 (`guide-only`); the failure pareto is S2
(source-upload PUTs), S8 (connector config `api-token`), S11 (activation-mint
handoff discipline), S5 (build-run state). Stage gates S0/S1/S3/S4/S6/S7/S9/
S10 passed in every run of both arms. Known measurement caveats, documented so
the gate is read correctly: S2's PUT-200 leg only observes bash `-X PUT`
results, while the omp agent uploads through code-mode programs (the
`upload_id` + `create_draft_source_upload` legs of S2 pass; the PUT leg reads
0-with-0); S8's `api-token` reads EMPTY from `c1_connector_service_get`
(secret-masked responses). The generator remains covered by unit smokes
(`evals/runner/baseline.test.ts`).

Full per-run records (run ids, funnels, per-stage evidence) live in the
private harness repo (`ductone/squire-evals`, branch
`bjorn/CXF-233/collector-transport-override`); the `*.jsonl` records behind
`baseline.json` are copied into this directory but stay gitignored.

## Non-goals

- Tier-2 real sandbox providers and the qualitative LLM-judge tier.
- Operator-side activation E2E leg (redeeming the approval token) — those two
fields are `skipped_human_boundary`.
- Baseline matrix runs and CI regression gating — the runs are blocked on the
E2E tool surface (halt path); no `baseline.json` is committed in this PR.
- CI regression gating workflow — the gate workflow itself is a later PR;
the committed `baseline.json` plus the consumption rule above is the
reference it will consume (the baseline matrix itself ran — CXF-233).
- The `/v2` fixture surface (bearer + link pagination) is fixture capability
asserted by `verify.sh` only; the Tier-1 agent uses `/v1` (basic + offset).
- Tier-1+ end-to-end runs require a private driver (credentials + a real
Expand All @@ -261,18 +278,16 @@ real-tenant driver and the tool surface are available.
`examples/**/*.ts` and `baton/**/*.d.ts` surface — the evals options do not
leak into the repo-wide type environment. `npm run typecheck` runs both
configs; the CI workflow is unmodified.
- **E2E status.** The Tier-1 end-to-end run (done-definition 4) is BLOCKED.
The public repo ships only the Tier-0 local driver (canned transcript); a
Tier-1+ run needs a private driver with real tenant credentials and an
agent transport, which is out of scope for the public repo. The CXF-217
preflight (on the Squire-based harness) also established a structural
blocker on the c1 side: fresh c1-image eval envs expose no
`c1_connector_authoring_*` tools even with `CONNECTOR_AUTHORING` effective
(the `c1.api.*` MCP surface is not mounted in this region's envs) —
evidence at `/current-tasks/src-tu2rs/results/BLOCKER.md`. The batch halted
per done-definition 7; the runner, scorer, and stage gates are covered by
the committed unit smokes (`npm run eval:test`), and the E2E must run once
a real-tenant driver and the tool surface are available.
- **E2E status.** The Tier-1 end-to-end run RAN (2026-09-04/05, CXF-233) via
the private squire driver (`ductone/squire-evals`, omp agent harness against
real dev-tenant envs); the six scored records behind the committed
`baseline.json` come from those runs. The earlier CXF-217 structural blocker
(fresh c1-image envs exposing no `c1_connector_authoring_*` tools) was NOT
observed in this dispatch: every env's readiness probe found the full
authoring surface. The public repo still ships only the Tier-0 local driver;
the squire driver (credentials + agent transport) remains in the private
harness repo, out of scope here. The runner, scorer, and stage gates stay
covered by the committed unit smokes (`npm run eval:test`).
- **Score-input boundary.** `score-input.json` is written by a collector
agent that transcribes tenant tool responses; the scorer type-validates but
cannot verify truthfulness. The collector reads the agent-written handoff
Expand Down
182 changes: 182 additions & 0 deletions evals/results/baseline.json
Original file line number Diff line number Diff line change
@@ -0,0 +1,182 @@
{
"schema_version": 1,

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Suggestion: This file becomes a committed contract artifact that a future CI gate will consume, but nothing validates it — evals/runner/baseline.test.ts only exercises the generator against temp dirs, and the source *.jsonl records are gitignored and live in a private repo, so the committed JSON is unreproducible here. Consider a small check that loads evals/results/baseline.json and asserts the locked v1 shape (schema_version === 1, required top-level keys, both matrix modes present, per_stage covering S0–S11) so drift or a hand-edit is caught before the gate PR lands. (confidence: medium)

"generated_at": "2026-09-05T01:50:55.103Z",
"model": "together/deepseek-ai/DeepSeek-V4-Flash-0731",
"reasoning_effort": "high",
"scenarios": [
"tier1-directory",
"tier1-directory-guide-only"
],
Comment on lines +6 to +9

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Suggestion: scenarios records two ids (tier1-directory, tier1-directory-guide-only), which contradicts the unchanged README description at evals/README.md:200-201 — "scenario tier1-directory × skill-bundle modes {none, guide-only} × 3 runs each". The guide-only arm actually came from a separate scenario file (evals/scenarios/tier1-directory-guide-only.json, mode: "guide-only"), and the generator enforces one scenario per mode group, so the README's single-scenario phrasing (and its (CXF-217) attribution) should be corrected now that the data makes the mismatch visible. (confidence: high)

"modes": {
"guide-only": {
"run_ids": [
"evals-tier1-directory-guide-only-20260905-000901-207",
"evals-tier1-directory-guide-only-20260905-012128-950",
"evals-tier1-directory-guide-only-20260905-015055-103"
],
"runs": 3,
"passes": 0,
"pass_rate": 0,
"pass_at_3": 0,
"first_pass_rate_mean": 0.6111111111111112,
"per_stage": {
"S0": {
"runs": 3,
"failures": 0,
"first_pass_failures": 1
},
"S1": {
"runs": 3,
"failures": 0,
"first_pass_failures": 0
},
"S2": {
"runs": 3,
"failures": 3,
"first_pass_failures": 3
},
"S3": {
"runs": 3,
"failures": 0,
"first_pass_failures": 0
},
"S4": {
"runs": 3,
"failures": 0,
"first_pass_failures": 1
},
"S5": {
"runs": 3,
"failures": 2,
"first_pass_failures": 2
},
"S6": {
"runs": 3,
"failures": 0,
"first_pass_failures": 0
},
"S7": {
"runs": 3,
"failures": 0,
"first_pass_failures": 0
},
"S8": {
"runs": 3,
"failures": 3,
"first_pass_failures": 3
},
"S9": {
"runs": 3,
"failures": 0,
"first_pass_failures": 0
},
"S10": {
"runs": 3,
"failures": 0,
"first_pass_failures": 1
},
"S11": {
"runs": 3,
"failures": 3,
"first_pass_failures": 3
}
}
},
"none": {
"run_ids": [
"evals-tier1-directory-20260904-204206-352",
"evals-tier1-directory-20260904-221013-308",
"evals-tier1-directory-20260904-231553-272"
],
"runs": 3,
"passes": 0,
"pass_rate": 0,
"pass_at_3": 0,
"first_pass_rate_mean": 0.47222222222222227,
"per_stage": {
"S0": {
"runs": 3,
"failures": 0,
"first_pass_failures": 1
},
"S1": {
"runs": 3,
"failures": 0,
"first_pass_failures": 0
},
"S2": {
"runs": 3,
"failures": 3,
"first_pass_failures": 3
},
"S3": {
"runs": 3,
"failures": 0,
"first_pass_failures": 2
},
"S4": {
"runs": 3,
"failures": 0,
"first_pass_failures": 3
},
"S5": {
"runs": 3,
"failures": 2,
"first_pass_failures": 3
},
"S6": {
"runs": 3,
"failures": 0,
"first_pass_failures": 0
},
"S7": {
"runs": 3,
"failures": 0,
"first_pass_failures": 0
},
"S8": {
"runs": 3,
"failures": 3,
"first_pass_failures": 3
},
"S9": {
"runs": 3,
"failures": 0,
"first_pass_failures": 0
},
"S10": {
"runs": 3,
"failures": 0,
"first_pass_failures": 2
},
"S11": {
"runs": 3,
"failures": 2,
"first_pass_failures": 2
}
}
}
},
"pareto": [
{
"stage": "S2",
"failures": 6,
"share": 0.2857142857142857
},
{
"stage": "S8",
"failures": 6,
"share": 0.2857142857142857
},
{
"stage": "S11",
"failures": 5,
"share": 0.23809523809523808
},
{
"stage": "S5",
"failures": 4,
"share": 0.19047619047619047
}
]
}
Loading