Skip to content

Findings: scan the dataset through the eval's task - #16

Merged
MattFisher merged 2 commits into
feat/findingsfrom
feat/dataset-task-mode
Sep 30, 2026
Merged

MattFisher merged 2 commits into
feat/findingsfrom
feat/dataset-task-mode

Conversation

@MattFisher

Copy link
Copy Markdown

Dataset scans now read the eval's samples through its task by default, instead of scanning eval.yaml's HuggingFace asset with auto-detected arguments.

Why. A survey of 32 inspect_evals tasks ran inspect-dataset scan <hf asset> exactly as the adapter does. Only 3 ran, and all three scanned train, which none of those evals uses. The other 29 failed on:

Cause Evals
Config name missing 12
Column names not auto-detected 7
No train split 4
google/ or openai/ owner mistaken for a Python package (inspect_dataset#31) 4
Script-based datasets 2

inspect-dataset scan inspect_evals/<task> worked for all 32.

Task scans. The adapter runs inspect-dataset scan inspect_evals/<task> for the first task in eval.yaml, or for a task declared in the pilot config. inspect-dataset then loads the samples through the eval's own loader, with its split, config, revision, field mapping and sample ids. Because that imports the eval, the command runs in the checkout's own environment: uv run --project <root> --frozen --with <inspect-dataset> inspect-dataset. --frozen means the checkout's uv.lock is never rewritten. INSPECT_AUDIT_DATASET_TASK_CMD overrides the command, with {ie_root} standing for the checkout path.

Declarations. A declaration with HuggingFace settings (path, config, split, revision, fields) still scans the HuggingFace dataset directly, as before. task cannot be combined with those settings; validation rejects it.

Recording. The run's subject dataset is eval.yaml's HuggingFace asset, or the task spec when there is none. The adapter does not use the summary's dataset fields in task mode, because a task scan's summary names the spec and a train split whatever the task loads (inspect_dataset#40). inputs.dataset gains mode and task. The Inputs section says "through the task's own loader".

StereoSet. Its HuggingFace declaration is removed from pilot.yaml in its own commit, so it scans through its task. The result is 18 findings (the real duplicate contexts) instead of 2,169. The sample ids are the same 2,123 in both modes, so those 18 findings keep their fingerprints. The config comment now shows how to declare a task.

Trial on real inputs (live inspect_evals checkout, pinned inspect-dataset afbc94c0b509):

Eval Before Now (task scan)
ARC skip 12 findings
GSM8K skip 4 findings
MMLU-Pro skip 850 findings
HumanEval skip 156 findings
StereoSet 2,169 findings 18 findings

The checkout had no tracked changes afterwards.

Limits.

  • One task per eval: ARC, WMDP and LAB-Bench cover their first task unless a task is declared.
  • A task scan imports the eval's code on the host, as running the eval would. That is fine for inspect_evals, but it needs a rethink before the adapter points at third-party repos.
  • Many task-scan findings are still scanner misfires (letter targets, choices ignored, whole prompts split on "or"). Those are tracked as inspect_dataset issues #30 and #32–#39.
  • The adapter does not yet read the scanner_status added by inspect_dataset#29. That can follow once #29 merges and the pin moves.

Gate: pre-commit clean, basedpyright 0, 501 passed.

🤖 Generated with Claude Code

MattFisher and others added 2 commits September 30, 2026 15:11
Scanning eval.yaml's HuggingFace asset with no other arguments failed
for 29 of 32 surveyed evals: missing configs and splits, column names
that auto-detection misses, script-based datasets, and HF owners such as
google/ and openai/ that inspect-dataset mistakes for Python packages.
The three that ran scanned train, which none of those evals uses.

The adapter now runs `inspect-dataset scan inspect_evals/<task>` for the
first task in eval.yaml, or a declared `task`. That loads the samples
through the eval's own loader, with its split, config, revision, field
mapping and sample ids. Because it imports the eval, it runs in the
checkout's own environment (`uv run --project <root> --frozen --with
<inspect-dataset>`); INSPECT_AUDIT_DATASET_TASK_CMD overrides it.

A declaration with HuggingFace settings still scans the HuggingFace
dataset directly, and a task cannot be combined with those settings.
The run's subject dataset is eval.yaml's HuggingFace asset, or the task
spec when there is none, because a task scan's summary names the spec
and a train split whatever the task loads. inputs.dataset gains `mode`
and `task`, and the Inputs section says when a scan went through the
task.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The pilot declared StereoSet's HuggingFace config, split and field roles
because auto-detection got them wrong. A task scan reads them from the
eval instead, and it sees the eval's samples: 18 findings (the real
duplicate contexts) instead of 2,169, most of which were length rules
applied to the repr of the struct-typed sentences column. The config
comment now shows how to declare a task.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@MattFisher
MattFisher merged commit f490ae2 into feat/findings Sep 30, 2026
2 checks passed
@MattFisher
MattFisher deleted the feat/dataset-task-mode branch September 30, 2026 05:23
MattFisher added a commit that referenced this pull request Sep 30, 2026
…id and issue columns; grouped summaries

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant