Repository navigation
Conversation
…ixes Concordance is rewritten around Inspect's own primitives: a `concordance_scorer` runs under `score_async(action="append")`, where the TaskState already carries the recorded grades, and scores each attempt AGREE, STABLE_DISAGREEMENT or NOISY_DISAGREEMENT with a frequency metric. Grades compare through `value_to_float`. The sliced logs `sample_logs` already writes are the input, via a new `sliced_logs` sample metadata field. `concordance_gate` runs as a setup solver in `audit_task` before the auditor's first turn, stores the verdict in a `Concordance` StoreModel, `grade` refuses while it is blocked, and every item score carries the verdict. 535 lines to 212; no private imports; the registry probe reads the store instead of replaying on its own. Ledger: `usage_cost` sums per-sample usage (header totals miss setup-solver and task-built grader spend) and clamps costs at zero; jobs record `expected_evals` and `wait`/`collect` require that many terminal rows; `stop` with no evals and collect on a stopped empty job settle at $0; an unpriced collection is charged its reservation with a note rather than held forever; `cost_limit` has a floor and boolean sizes are refused at the raw config; any config string equal to a catalogue model id off the worker menu is refused under any key. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…crets file Middleman already registers sol, luna and gpt-5-mini against its OpenRouter lab and holds the org's OpenRouter key, so the bypass we built for ZDR was never needed: runner -> middleman -> OpenRouter -> Azure is the same path. Removing it deletes `HAWK_RUNNER_REFRESH_URL: ""`, `runner.secrets`, the `--secrets-file` flag and the `args.base_url` override the policy forced onto every model. It also closes the judge defect: with no override, the environment pair Hawk sets is consistent for models a task builds by name too, and no provider key sits in any runner. `secrets_file` stays as an ignored parameter so saved investigation files still load. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The `report` task was the pre-investigate conversational entry point over finished audit logs; nothing uses it. The frontend was an unauthenticated ACP relay that the security review turned into a browser-to-agent attack path. Interactive mode stays as Inspect's own `inspect acp`. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The four documents described the code in prose, duplicated what it says, and were stale within a day of being written (they still described S3 staging and the middleman bypass after both were removed). Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
components.py, irt_analysis.py and styles.css were added on 2026-09-10. The architecture component's SVG output, previewed through view_image, killed two investigations at the report stage; girth==0.8.0 is a four-year-unmaintained single-maintainer dependency pulled in for an optional exploratory fit. The agent has matplotlib and Quarto; the writing skill keeps the figure rules and asks for PNG. The report template loses the architecture section. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
`swebench_replay`, `replay_exploit` and `replay_task` were stage two of an exploit experiment nothing runs. `fetch_logs` keeps `hawk:` eval sets and native paths; the HTTP manifest branch was a second, weaker route to the same corpus and a header-driven download in the runner. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…inter The generator ran at task construction and duplicated the two examples the investigating skill already carries; the agent writes its own config from them. `render_report` wrapped `quarto render` and returned 6,000 characters of a report; the agent has quarto in bash and view_image for figures. The linter enforced dashes, drafting comments and sentence length in Python; those are the writing skill's rules and they belong in its text, not in a gate. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
# Conflicts: # src/inspect_audit/_registry.py
The HuggingFace source comes from eval.yaml; scan fields come from a per-eval override table because auto-detection fails on StereoSet. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Five zero-cost checks over .eval headers: dataset sample count against eval.yaml, task and package version drift, unscored samples, dirty revisions, and (behind --resolve) logged ids missing from the dataset. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
`run` executes the header producer plus selected external producers per eval and writes run JSON, parquet and markdown summaries; `summary` re-renders from disk. `--logs` accepts hawk: eval sets via fetch_logs. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
… with logs, resolve guards Adapters no longer raise past run() on malformed producer output. The header producer matches logs on every task name in the eval's eval.yaml, so multi-task packages such as lab_bench are found, and compares each log's sample count with its own task's declaration. The header producer is only requested when --logs is given, so a lint-and-dataset sweep can exit 0. --resolve skips rather than reporting false major findings when the resolved dataset has no sample ids. The spec now records lint's exit 1 as a result, and the committed acceptance scan drops the raw dataset copy. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Findings carry `taxonomy` (default gl-audit@1) and an optional `check`; dimension and check are validated against that taxonomy rather than a Literal. gl-audit@1 is generated from the LaTeX framework and a test keeps them equal; gl-audit@2 transcribes the Audit Reports Strategy document's seven dimensions as a draft; a mapping file translates v1 ids to v2. Parquet gains taxonomy and check columns. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
`hawk-sets <task>` and `run --hawk-task <task>` discover eval sets through the hawk client (which uses the operator's `hawk login` token) by matching each set's recorded task names. `hawk:<id>` sources are fetched with `hawk download` into ~/.cache/inspect_audit/hawk, where the CLI skips files already present. The header producer runs whenever any log source was given. This replaces the plan to reuse _registry.fetch_logs, which needs runner token environment variables and re-downloads every run. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
audit/inspect_evals/scicode is the sample auditor over scicode, not scicode; only a bare name may match on its unqualified tail. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
… bundle from Hawk Fixtures for the audit adapter: 79 sample audits (302 question labels over 288 SciCode subproblems) and one complete published investigation bundle with findings, assessments, coverage and evidence. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
… and bundles into a gitignored dir
scripts/hawk-artefacts.yaml names the eval sets (logs:) and investigator
bundles (artifacts:) we work from: James's six SciCode GL sample-audit
sets, the SciCode union corpus, an HLE physics corpus, and the chess
investigation bundle. `inspect-audit-findings hawk-pull` lands them in
artefacts/hawk/{logs,artifacts}/<set>/, which is gitignored, so any dev
with `hawk login` gets the same inputs without private outputs entering
the public repo.
A failed entry is reported and the rest still pull. `hawk
download-artifacts` needs aiofiles, which the 3.5.0 hawk[cli] extra
omits, so the remote and dev extras add it.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Six milestones with done-tests, the decisions still open, and who we depend on. Sources: the GL Q4 plan (Outcome 3), the Audit Reports Strategy doc, the findings-prototype acceptance sweep, and the SciCode v1.2 report. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
… security and cost points from the first assessment Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…ndings/ Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…t hook fixes to findings docs and fixtures Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
pyproject: keep both new dependencies (aiofiles for hawk download-artifacts, inspect-k8s-sandbox for the k8s converter); uv.lock regenerated (hawk 3.6.0). Test helper registers its workers' prices, so job tests no longer depend on another test in the same process (they failed under pytest -n auto). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
dev/audit-logs was already gitignored but tracked from before the ignore; nothing imports the scripts. Kept locally outside the repo. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Remove dev/ before sharing
…ross retries record_verdict: on the 09-29 chess run an auditor wrote prose summaries into details 54 times against a bare "must encode a valid JSON object" and hit its cost limit. The refusal now says what it received, that the summary goes in remarks, and the item's fields and shape. details stays a JSON string: tested live, OpenAI's validator (Azure via OpenRouter) refuses an untyped object property, and an object without properties comes back empty. NO_VERDICT scores explain themselves: the refusals record_verdict gave, or that it was never called, and any limit the sample hit. Puzzle 40 had hit its cost limit before any verdict; the investigator, seeing only concordance metadata, blamed resolution_drift. Spend: an attempt with an unpriced model wrote its priced part to disk as if it were the cost, and a retry then let reservations through. The unpriced models now travel with the figure (spend file and checkpoint record), and the allowance checks and budget tool consult them. The live verdict-schema test now covers gpt-6-luna as well as gpt-5-mini. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…ays unknown Auditor verdicts say why they failed; unpriced spend stays unknown across retries
…records what it compared against A finding derived from a historical log was stamped with the current checkout's commit and task version, so a page could attribute a defect in a log from version 2 to the version 3-A source. Each finding's subject now comes from the header (eval.revision, packages, task_version, task_args) and Run.inputs.comparison names the checkout used for eval.yaml. Found by the 2026-09-29 prototype review. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
… carries the review fields Runs are written to <slug>/runs/<run id>.run.json and never overwritten. A sweep only moves <slug>/current.json, which names the run each producer's view uses, so a partial sweep no longer mixes a fresh lint run with whatever files the last sweep left behind, and `summary` renders from the manifest rather than from every file on disk. findings.parquet gains the subject fields it was missing (dirty, comparability, interface, dataset config/split/revision, task_args) and the review fields (aliases, suppressions, suppressed, history, introduced, fixed, effect) as JSON columns, so a site reading the export alone has what the roadmap expects it to use. Found by the 2026-09-29 prototype review. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Copied from branch findings-roadmap-review (422aec2), which also edited the roadmap; the roadmap re-cut that acts on this review follows in the next commit. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…chema and identity decisions Roadmap: the six milestones become a local pilot, measurement, a published pilot, and scheduled expansion, per the review. Input selection comes first; the historical-defect test runs before and after each fix; voting starts as GitHub issue reactions; the report-from-store work is its own track that does not block the pilot. Schema doc: Assessment and Coverage as records beside Finding, both provisional until the pilot; observation identity (fingerprint) kept separate from issue identity (issues.yaml); suppressions.yaml beside the runs. finding-schema.md marked superseded, the implementation plan marked historical. Spec updated for immutable runs, current.json, the wider parquet and log-derived header subjects. The CLI's commands and current limitations go in docs/findings-cli.md, since the README was pared back to one line in #10. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…wups Findings: log-derived header subjects, immutable runs with a current view, review docs
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Replaces #9, whose branch
dev/integrated-auditsis nowdev(same commit, 8eb12f5).Everything since 15 Sep: the GL report template, the python-project-template (#5), the Hawk rework (#7), the findings prototype (#4) and follow-ups (#12), and the auditor/spend fixes (#11), via #8. Proven on Hawk by
chess-audit-luna6-20260929: 100 of 100 Epoch Chess Puzzles audited,published_complete, about $6.🤖 Generated with Claude Code