Repository navigation
Findings: leads, one eval's reviewed findings as hypotheses for the agents - #20
Merged
Merged
Conversation
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…h record ids to cite Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…blic renderer helpers Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…e from packages[].rules A rules row has no message, so the skip reason from the matching outcomes row is carried over; LEADS.md's Not examined section was listing bare rule codes. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
- --sample counts only that sample's suppressions and says how many eval-wide observations were left out, instead of reporting the whole eval's count as the sample's. - The group count on the summary line matches the headings. - Skip reasons render as one bounded line; an empty reason says so; a whole-producer skip is named as such, so a traceback cannot break the file or fill the agent's context. - An eval where every producer skipped says so on the summary line rather than reading as "0 observations". - The preamble no longer tells the sample auditor to treat leads like worker verdicts; it states the method directly and says "automated producer", since Scout-style producers are not deterministic. - The lint adapter tolerates `outcomes: null` beside `rules`. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Author
|
Deferred minors from the branch review, none blocking:
Left as they are, on purpose: Posted by Claude Code on Matt's behalf. |
…ed file is the fallback and snapshot Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
# Conflicts: # tests/findings/test_lint_adapter.py
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Roadmap milestone 1, third slice, per the plan. Targets
feat/findings.What it is.
inspect-audit-findings leads --out DIR [--sample ID] [--write PATH] EVALrenders one eval's current findings asLEADS.md: hypotheses with a location attached, for an agent to confirm or retire. Same current view as the summaries, review files applied. Suppressed observations dropped and counted; accepted issues listed first; groups capped at three examples with a pointer to the run file for the rest; skipped checks listed with one-line reasons. Every example carries its record id, which is what the agent cites when a register finding rests on a lead.--samplenarrows to observations naming that sample (dataset and transcript locations) and says how many eval-wide observations were left out.Staging contract.
docs/leads.mddescribes the hooks: the investigator stages/inputs/findings/LEADS.mdplus the run files and names it inseed.json; the sample auditor stages/audit/leads.mdbesidediscrepancies.md; the chess import later reads cited record ids back into status history. Those hooks touch James's code and prompts and go as a separate PR. This module imports nothing from them.Also. Two renderer helpers (
inputs_lines,sorted_findings) are now public and shared. Under lint 0.9'srulesmode, skipped rules keep their reason from the matchingoutcomesrow; the trial had shown "Not examined" listing bare codes. The lint adapter toleratesoutcomes: null.Trial. StereoSet with lint and dataset producers and one accepted issue: preamble, Inputs, the issue, two duplicate-question groups capped at three with record ids, ten skipped lint rules with reasons.
--sample <id>narrowed to the one observation on that sample.Review. A fresh reviewer found one Important issue (
--samplereported the eval's suppressed count as the sample's and hid eval-wide leads silently) and nine minors; the Important one and five minors that change what the agent reads are fixed with tests that failed first. The rest are in the first comment.Gate: pre-commit clean, basedpyright 0, 535 passed.
🤖 Generated with Claude Code