Repository navigation
Findings: scan the dataset through the eval's task - #16
Merged
Merged
Conversation
Scanning eval.yaml's HuggingFace asset with no other arguments failed for 29 of 32 surveyed evals: missing configs and splits, column names that auto-detection misses, script-based datasets, and HF owners such as google/ and openai/ that inspect-dataset mistakes for Python packages. The three that ran scanned train, which none of those evals uses. The adapter now runs `inspect-dataset scan inspect_evals/<task>` for the first task in eval.yaml, or a declared `task`. That loads the samples through the eval's own loader, with its split, config, revision, field mapping and sample ids. Because it imports the eval, it runs in the checkout's own environment (`uv run --project <root> --frozen --with <inspect-dataset>`); INSPECT_AUDIT_DATASET_TASK_CMD overrides it. A declaration with HuggingFace settings still scans the HuggingFace dataset directly, and a task cannot be combined with those settings. The run's subject dataset is eval.yaml's HuggingFace asset, or the task spec when there is none, because a task scan's summary names the spec and a train split whatever the task loads. inputs.dataset gains `mode` and `task`, and the Inputs section says when a scan went through the task. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The pilot declared StereoSet's HuggingFace config, split and field roles because auto-detection got them wrong. A task scan reads them from the eval instead, and it sees the eval's samples: 18 findings (the real duplicate contexts) instead of 2,169, most of which were length rules applied to the repr of the struct-typed sentences column. The config comment now shows how to declare a task. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
MattFisher
added a commit
that referenced
this pull request
Sep 30, 2026
…id and issue columns; grouped summaries Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Dataset scans now read the eval's samples through its task by default, instead of scanning
eval.yaml's HuggingFace asset with auto-detected arguments.Why. A survey of 32 inspect_evals tasks ran
inspect-dataset scan <hf asset>exactly as the adapter does. Only 3 ran, and all three scannedtrain, which none of those evals uses. The other 29 failed on:trainsplitgoogle/oropenai/owner mistaken for a Python package (inspect_dataset#31)inspect-dataset scan inspect_evals/<task>worked for all 32.Task scans. The adapter runs
inspect-dataset scan inspect_evals/<task>for the first task ineval.yaml, or for ataskdeclared in the pilot config. inspect-dataset then loads the samples through the eval's own loader, with its split, config, revision, field mapping and sample ids. Because that imports the eval, the command runs in the checkout's own environment:uv run --project <root> --frozen --with <inspect-dataset> inspect-dataset.--frozenmeans the checkout'suv.lockis never rewritten.INSPECT_AUDIT_DATASET_TASK_CMDoverrides the command, with{ie_root}standing for the checkout path.Declarations. A declaration with HuggingFace settings (path, config, split, revision, fields) still scans the HuggingFace dataset directly, as before.
taskcannot be combined with those settings; validation rejects it.Recording. The run's subject dataset is
eval.yaml's HuggingFace asset, or the task spec when there is none. The adapter does not use the summary's dataset fields in task mode, because a task scan's summary names the spec and atrainsplit whatever the task loads (inspect_dataset#40).inputs.datasetgainsmodeandtask. The Inputs section says "through the task's own loader".StereoSet. Its HuggingFace declaration is removed from
pilot.yamlin its own commit, so it scans through its task. The result is 18 findings (the real duplicate contexts) instead of 2,169. The sample ids are the same 2,123 in both modes, so those 18 findings keep their fingerprints. The config comment now shows how to declare a task.Trial on real inputs (live inspect_evals checkout, pinned inspect-dataset
afbc94c0b509):The checkout had no tracked changes afterwards.
Limits.
scanner_statusadded by inspect_dataset#29. That can follow once #29 merges and the pin moves.Gate: pre-commit clean, basedpyright 0, 501 passed.
🤖 Generated with Claude Code