Skip to content

Land dev in main - #22

Open
jrh-mann wants to merge 106 commits into
mainfrom
dev
Open

jrh-mann wants to merge 106 commits into
mainfrom
dev

Conversation

@jrh-mann

@jrh-mann jrh-mann commented Oct 2, 2026

Copy link
Copy Markdown
Contributor

Replaces #9, whose branch dev/integrated-audits is now dev (same commit, 8eb12f5).

Everything since 15 Sep: the GL report template, the python-project-template (#5), the Hawk rework (#7), the findings prototype (#4) and follow-ups (#12), and the auditor/spend fixes (#11), via #8. Proven on Hawk by chess-audit-luna6-20260929: 100 of 100 Epoch Chess Puzzles audited, published_complete, about $6.

🤖 Generated with Claude Code

jrh-mann and others added 30 commits September 11, 2026 16:53
…ixes

Concordance is rewritten around Inspect's own primitives: a `concordance_scorer`
runs under `score_async(action="append")`, where the TaskState already carries the
recorded grades, and scores each attempt AGREE, STABLE_DISAGREEMENT or
NOISY_DISAGREEMENT with a frequency metric. Grades compare through
`value_to_float`. The sliced logs `sample_logs` already writes are the input,
via a new `sliced_logs` sample metadata field. `concordance_gate` runs as a setup
solver in `audit_task` before the auditor's first turn, stores the verdict in a
`Concordance` StoreModel, `grade` refuses while it is blocked, and every item score
carries the verdict. 535 lines to 212; no private imports; the registry probe reads
the store instead of replaying on its own.

Ledger: `usage_cost` sums per-sample usage (header totals miss setup-solver and
task-built grader spend) and clamps costs at zero; jobs record `expected_evals` and
`wait`/`collect` require that many terminal rows; `stop` with no evals and collect
on a stopped empty job settle at $0; an unpriced collection is charged its
reservation with a note rather than held forever; `cost_limit` has a floor and
boolean sizes are refused at the raw config; any config string equal to a
catalogue model id off the worker menu is refused under any key.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…crets file

Middleman already registers sol, luna and gpt-5-mini against its OpenRouter lab
and holds the org's OpenRouter key, so the bypass we built for ZDR was never
needed: runner -> middleman -> OpenRouter -> Azure is the same path. Removing it
deletes `HAWK_RUNNER_REFRESH_URL: ""`, `runner.secrets`, the `--secrets-file`
flag and the `args.base_url` override the policy forced onto every model. It
also closes the judge defect: with no override, the environment pair Hawk sets
is consistent for models a task builds by name too, and no provider key sits in
any runner. `secrets_file` stays as an ignored parameter so saved investigation
files still load.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The `report` task was the pre-investigate conversational entry point over finished
audit logs; nothing uses it. The frontend was an unauthenticated ACP relay that the
security review turned into a browser-to-agent attack path. Interactive mode stays
as Inspect's own `inspect acp`.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The four documents described the code in prose, duplicated what it says, and
were stale within a day of being written (they still described S3 staging and
the middleman bypass after both were removed).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
components.py, irt_analysis.py and styles.css were added on 2026-09-10. The
architecture component's SVG output, previewed through view_image, killed two
investigations at the report stage; girth==0.8.0 is a four-year-unmaintained
single-maintainer dependency pulled in for an optional exploratory fit. The
agent has matplotlib and Quarto; the writing skill keeps the figure rules and
asks for PNG. The report template loses the architecture section.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
`swebench_replay`, `replay_exploit` and `replay_task` were stage two of an
exploit experiment nothing runs. `fetch_logs` keeps `hawk:` eval sets and
native paths; the HTTP manifest branch was a second, weaker route to the
same corpus and a header-driven download in the runner.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…inter

The generator ran at task construction and duplicated the two examples the
investigating skill already carries; the agent writes its own config from them.
`render_report` wrapped `quarto render` and returned 6,000 characters of a
report; the agent has quarto in bash and view_image for figures. The linter
enforced dashes, drafting comments and sentence length in Python; those are
the writing skill's rules and they belong in its text, not in a gate.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
# Conflicts:
#	src/inspect_audit/_registry.py
MattFisher and others added 30 commits September 29, 2026 12:03
The HuggingFace source comes from eval.yaml; scan fields come from a
per-eval override table because auto-detection fails on StereoSet.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Five zero-cost checks over .eval headers: dataset sample count against
eval.yaml, task and package version drift, unscored samples, dirty
revisions, and (behind --resolve) logged ids missing from the dataset.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
`run` executes the header producer plus selected external producers per
eval and writes run JSON, parquet and markdown summaries; `summary`
re-renders from disk. `--logs` accepts hawk: eval sets via fetch_logs.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
… with logs, resolve guards

Adapters no longer raise past run() on malformed producer output. The
header producer matches logs on every task name in the eval's eval.yaml,
so multi-task packages such as lab_bench are found, and compares each
log's sample count with its own task's declaration. The header producer
is only requested when --logs is given, so a lint-and-dataset sweep can
exit 0. --resolve skips rather than reporting false major findings when
the resolved dataset has no sample ids. The spec now records lint's exit
1 as a result, and the committed acceptance scan drops the raw dataset
copy.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Findings carry `taxonomy` (default gl-audit@1) and an optional `check`;
dimension and check are validated against that taxonomy rather than a
Literal. gl-audit@1 is generated from the LaTeX framework and a test
keeps them equal; gl-audit@2 transcribes the Audit Reports Strategy
document's seven dimensions as a draft; a mapping file translates v1
ids to v2. Parquet gains taxonomy and check columns.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
`hawk-sets <task>` and `run --hawk-task <task>` discover eval sets through
the hawk client (which uses the operator's `hawk login` token) by matching
each set's recorded task names. `hawk:<id>` sources are fetched with
`hawk download` into ~/.cache/inspect_audit/hawk, where the CLI skips files
already present. The header producer runs whenever any log source was
given. This replaces the plan to reuse _registry.fetch_logs, which needs
runner token environment variables and re-downloads every run.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
audit/inspect_evals/scicode is the sample auditor over scicode, not scicode;
only a bare name may match on its unqualified tail.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
… bundle from Hawk

Fixtures for the audit adapter: 79 sample audits (302 question labels over
288 SciCode subproblems) and one complete published investigation bundle
with findings, assessments, coverage and evidence.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
… and bundles into a gitignored dir

scripts/hawk-artefacts.yaml names the eval sets (logs:) and investigator
bundles (artifacts:) we work from: James's six SciCode GL sample-audit
sets, the SciCode union corpus, an HLE physics corpus, and the chess
investigation bundle. `inspect-audit-findings hawk-pull` lands them in
artefacts/hawk/{logs,artifacts}/<set>/, which is gitignored, so any dev
with `hawk login` gets the same inputs without private outputs entering
the public repo.

A failed entry is reported and the rest still pull. `hawk
download-artifacts` needs aiofiles, which the 3.5.0 hawk[cli] extra
omits, so the remote and dev extras add it.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Six milestones with done-tests, the decisions still open, and who we
depend on. Sources: the GL Q4 plan (Outcome 3), the Audit Reports
Strategy doc, the findings-prototype acceptance sweep, and the SciCode
v1.2 report.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
… security and cost points from the first assessment

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…ndings/

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…t hook fixes to findings docs and fixtures

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
pyproject: keep both new dependencies (aiofiles for hawk download-artifacts,
inspect-k8s-sandbox for the k8s converter); uv.lock regenerated (hawk 3.6.0).
Test helper registers its workers' prices, so job tests no longer depend on
another test in the same process (they failed under pytest -n auto).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…#4)

Consolidate onto dev: Hawk rework (#7) + findings prototype (#4)
dev/audit-logs was already gitignored but tracked from before the ignore;
nothing imports the scripts. Kept locally outside the repo.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Remove dev/ before sharing
…ross retries

record_verdict: on the 09-29 chess run an auditor wrote prose summaries into
details 54 times against a bare "must encode a valid JSON object" and hit its
cost limit. The refusal now says what it received, that the summary goes in
remarks, and the item's fields and shape. details stays a JSON string: tested
live, OpenAI's validator (Azure via OpenRouter) refuses an untyped object
property, and an object without properties comes back empty.

NO_VERDICT scores explain themselves: the refusals record_verdict gave, or that
it was never called, and any limit the sample hit. Puzzle 40 had hit its cost
limit before any verdict; the investigator, seeing only concordance metadata,
blamed resolution_drift.

Spend: an attempt with an unpriced model wrote its priced part to disk as if it
were the cost, and a retry then let reservations through. The unpriced models
now travel with the figure (spend file and checkpoint record), and the
allowance checks and budget tool consult them.

The live verdict-schema test now covers gpt-6-luna as well as gpt-5-mini.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…ays unknown

Auditor verdicts say why they failed; unpriced spend stays unknown across retries
…records what it compared against

A finding derived from a historical log was stamped with the current
checkout's commit and task version, so a page could attribute a defect
in a log from version 2 to the version 3-A source. Each finding's
subject now comes from the header (eval.revision, packages, task_version,
task_args) and Run.inputs.comparison names the checkout used for
eval.yaml.

Found by the 2026-09-29 prototype review.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
… carries the review fields

Runs are written to <slug>/runs/<run id>.run.json and never overwritten.
A sweep only moves <slug>/current.json, which names the run each
producer's view uses, so a partial sweep no longer mixes a fresh lint
run with whatever files the last sweep left behind, and `summary`
renders from the manifest rather than from every file on disk.

findings.parquet gains the subject fields it was missing (dirty,
comparability, interface, dataset config/split/revision, task_args) and
the review fields (aliases, suppressions, suppressed, history,
introduced, fixed, effect) as JSON columns, so a site reading the export
alone has what the roadmap expects it to use.

Found by the 2026-09-29 prototype review.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Copied from branch findings-roadmap-review (422aec2), which also edited
the roadmap; the roadmap re-cut that acts on this review follows in the
next commit.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…chema and identity decisions

Roadmap: the six milestones become a local pilot, measurement, a
published pilot, and scheduled expansion, per the review. Input
selection comes first; the historical-defect test runs before and after
each fix; voting starts as GitHub issue reactions; the report-from-store
work is its own track that does not block the pilot.

Schema doc: Assessment and Coverage as records beside Finding, both
provisional until the pilot; observation identity (fingerprint) kept
separate from issue identity (issues.yaml); suppressions.yaml beside the
runs. finding-schema.md marked superseded, the implementation plan
marked historical. Spec updated for immutable runs, current.json, the
wider parquet and log-derived header subjects. The CLI's commands and
current limitations go in docs/findings-cli.md, since the README was
pared back to one line in #10.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…wups

Findings: log-derived header subjects, immutable runs with a current view, review docs
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants