Skip to content

fix(findings): replay task declares the eval's scorers so inspect-dataset's scorer-aware scanners judge applicability correctly - #28

Merged
MattFisher merged 1 commit into
feat/findingsfrom
fix/replay-scorers
Oct 7, 2026
Merged

MattFisher merged 1 commit into
feat/findingsfrom
fix/replay-scorers

Conversation

@MattFisher

@MattFisher MattFisher commented Oct 6, 2026 •

Copy link
Copy Markdown

Summary

Closes #27. A task-mode dataset scan replays dumped samples through a stand-in task that had no scorer, so inspect-dataset 0.5.0 marked answer_length and inconsistent_format not applicable with "this task has no scorer" for every eval.

  • The dump script records the task's scorer registry names in its meta, read the same way inspect-dataset reads them off a task.
  • The adapter passes them to the replay in INSPECT_AUDIT_SCORERS and records them under inputs.dataset.scorers.
  • The replay task declares one stub scorer per name, registered under the eval's exact scorer name. The stubs never score; the name is all inspect-dataset needs.
  • The Inputs section says which scorers the scan assumed.

Verified on StereoSet

The scan now reports the task's real scorers, inspect_evals/multiple_choice_scorer and inspect_evals/stereoset_scorer, and skips the two verbatim-text scanners with "this task scores with ..." rather than "no scorer". That skip is correct: StereoSet is not scored by verbatim text comparison, and the earlier 2169 findings were mostly answer_length noise on it.

The other drop noted on #27, duplicate_questions 18 to 0, is explained and by design: inspect-dataset 0.5.0 identifies a sample by its question and its choices, and StereoSet's eight repeated contexts carry different candidate sentences. No change needed.

Checks

  • uv run pre-commit run --all-files clean
  • uv run basedpyright src 0 errors
  • uv run pytest -q -p no:cacheprovider 575 passed, 5 skipped

🤖 Generated with Claude Code

…lay task so inspect-dataset judges applicability by them

Closes #27. The replay task declares stub scorers registered under the eval's scorer names; the run's inputs and the Inputs section record which scorers the scan assumed.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@MattFisher
MattFisher merged commit 9d54b71 into feat/findings Oct 7, 2026
2 checks passed
@MattFisher
MattFisher deleted the fix/replay-scorers branch October 7, 2026 02:45
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant