Skip to content

Findings: leads, one eval's reviewed findings as hypotheses for the agents - #20

Merged
MattFisher merged 8 commits into
feat/findingsfrom
feat/leads
Oct 5, 2026
Merged

MattFisher merged 8 commits into
feat/findingsfrom
feat/leads

Conversation

@MattFisher

Copy link
Copy Markdown

Roadmap milestone 1, third slice, per the plan. Targets feat/findings.

What it is. inspect-audit-findings leads --out DIR [--sample ID] [--write PATH] EVAL renders one eval's current findings as LEADS.md: hypotheses with a location attached, for an agent to confirm or retire. Same current view as the summaries, review files applied. Suppressed observations dropped and counted; accepted issues listed first; groups capped at three examples with a pointer to the run file for the rest; skipped checks listed with one-line reasons. Every example carries its record id, which is what the agent cites when a register finding rests on a lead. --sample narrows to observations naming that sample (dataset and transcript locations) and says how many eval-wide observations were left out.

Staging contract. docs/leads.md describes the hooks: the investigator stages /inputs/findings/LEADS.md plus the run files and names it in seed.json; the sample auditor stages /audit/leads.md beside discrepancies.md; the chess import later reads cited record ids back into status history. Those hooks touch James's code and prompts and go as a separate PR. This module imports nothing from them.

Also. Two renderer helpers (inputs_lines, sorted_findings) are now public and shared. Under lint 0.9's rules mode, skipped rules keep their reason from the matching outcomes row; the trial had shown "Not examined" listing bare codes. The lint adapter tolerates outcomes: null.

Trial. StereoSet with lint and dataset producers and one accepted issue: preamble, Inputs, the issue, two duplicate-question groups capped at three with record ids, ten skipped lint rules with reasons. --sample <id> narrowed to the one observation on that sample.

Review. A fresh reviewer found one Important issue (--sample reported the eval's suppressed count as the sample's and hid eval-wide leads silently) and nine minors; the Important one and five minors that change what the agent reads are fixed with tests that failed first. The rest are in the first comment.

Gate: pre-commit clean, basedpyright 0, 535 passed.

🤖 Generated with Claude Code

MattFisher and others added 6 commits September 30, 2026 17:09
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…h record ids to cite

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…blic renderer helpers

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…e from packages[].rules

A rules row has no message, so the skip reason from the matching
outcomes row is carried over; LEADS.md's Not examined section was
listing bare rule codes.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
- --sample counts only that sample's suppressions and says how many
  eval-wide observations were left out, instead of reporting the whole
  eval's count as the sample's.
- The group count on the summary line matches the headings.
- Skip reasons render as one bounded line; an empty reason says so; a
  whole-producer skip is named as such, so a traceback cannot break the
  file or fill the agent's context.
- An eval where every producer skipped says so on the summary line
  rather than reading as "0 observations".
- The preamble no longer tells the sample auditor to treat leads like
  worker verdicts; it states the method directly and says "automated
  producer", since Scout-style producers are not deterministic.
- The lint adapter tolerates `outcomes: null` beside `rules`.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@MattFisher

Copy link
Copy Markdown
Author

Deferred minors from the branch review, none blocking:

  • Within a group, examples are ordered by location, so observations already linked to an issue can take all three example slots while new ones fall into "and N more". Sort unlinked observations first, or say "k of n linked to ISS-x" on the heading.
  • The run-file pointers (runs/<id>.run.json) are relative to <out>/<slug>/. The investigator staging copies that directory beside LEADS.md; the sample-auditor staging in docs/leads.md copies only leads.md. Either copy the runs too or mark the pointers as reference only in that contract.
  • --write pointed at a directory is a traceback rather than exit 2.
  • The no-runs error could list the evals present in the current view, since stereoset without the inspect_evals/ prefix is an easy mistake.

Left as they are, on purpose: --sample does not match LogLocation.location_hint (no producer puts a sample reference there); retracted findings are not filtered out yet (nothing produces that status; the loop-closing slice adds it); the per-sample auditor gets only leads that name its sample, now with a count of what was left out.

Posted by Claude Code on Matt's behalf.

MattFisher and others added 2 commits October 5, 2026 16:46
…ed file is the fallback and snapshot

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
# Conflicts:
#	tests/findings/test_lint_adapter.py
@MattFisher
MattFisher merged commit de7718a into feat/findings Oct 5, 2026
2 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants