Skip to content

About

Dataset quality scanner for AI evaluation benchmarks

Resources

Stars

0 stars

Watchers

0 watching

Forks

Repository files navigation

inspect-dataset

Dataset quality scanner for AI evaluation benchmarks. Companion to inspect-scout, which analyses agent trajectories — inspect-dataset audits the underlying datasets themselves.

Installation

pip install inspect-dataset

Or with uv:

uv add inspect-dataset

Usage

# Scan a HuggingFace dataset
inspect-dataset scan flaviagiammarino/vqa-rad --split test -o findings/

# Pick a config and split (needed when the dataset has several configs, or several splits and none is "train")
inspect-dataset scan allenai/ai2_arc --config ARC-Challenge --split test -o findings/

# Pin to a specific revision
inspect-dataset scan flaviagiammarino/vqa-rad --revision abc123 -o findings/

# Override auto-detected field names
inspect-dataset scan my-org/my-dataset \
  --question-field prompt \
  --answer-field label \
  -o findings/

# Run only specific scanners
inspect-dataset scan flaviagiammarino/vqa-rad \
  --scanners answer_length,duplicate_questions

# Adjust answer length threshold (default: 4 words)
inspect-dataset scan flaviagiammarino/vqa-rad --max-answer-words 6

# Measure a scalar inside a struct answer column (here, each of StereoSet's candidate sentences)
inspect-dataset scan McGill-NLP/stereoset --config intersentence --split validation \
  --question-field context --answer-field sentences --answer-subfield sentence

# Compare each sample with its own subset (task mode detects the subset key itself)
inspect-dataset scan my-org/my-dataset --group-by subject

# Limit samples loaded
inspect-dataset scan flaviagiammarino/vqa-rad --limit 500

# Scan a local annotation directory (JSON samples + sidecar markdown gold),
# cross-checking gold against cached extraction-tool outputs
inspect-dataset scan path/to/samples/ \
  --files-root path/to/extraction-cache/ \
  --scanner-module my_benchmark.audit.scanners

# Run LLM-powered scanners (requires --model)
inspect-dataset scan flaviagiammarino/vqa-rad \
  --model openai/gpt-4o-mini --split test -o findings/

# Run only specific LLM scanners
inspect-dataset scan flaviagiammarino/vqa-rad \
  --model openai/gpt-4o-mini \
  --scanners label_correctness,ambiguity

# View a saved report
inspect-dataset report findings/

## Interactive viewer

`inspect-dataset view` serves a local React app for browsing findings and
triaging issues.

Like `inspect_ai` and `inspect-scout`, the built frontend artifacts are
shipped in the repository and included in the package.

### Getting started

1. Install development dependencies:

```bash
uv sync                     # the dev group is installed by default
  1. Return to the repository root and generate a findings directory if you do not already have one:
uv run inspect-dataset scan flaviagiammarino/vqa-rad --split test -o findings/
  1. Launch the viewer:
uv run inspect-dataset view findings/
  1. Open the URL printed by the command, usually:
http://localhost:7576/

Rebuilding the frontend

You only need to rebuild the frontend if you change files in src/inspect_dataset/_view/www/:

cd src/inspect_dataset/_view/www
npm install
npm run build

The viewer accepts either a single findings directory, a parent directory containing multiple findings directories, or an explicit list of directories:

uv run inspect-dataset view findings/
uv run inspect-dataset view results/
uv run inspect-dataset view results/vqa-rad/ results/medqa/

Scanners

Scanner Severity What it flags
answer_length medium Answers longer than N words (default: 4). Long answers are unlikely to be reproduced verbatim by exact-match scorers.
duplicate_questions low to high Samples that appear more than once, by question, choices and image content. One finding per group. Duplicates inflate sample counts.
inconsistent_format low/medium Capitalisation, punctuation, or length deviations from the dataset majority (80%+ threshold).
answer_distribution high Datasets where a single answer accounts for ≥85% of samples — a model that always predicts that answer would score highly without any understanding.
encoding_issues low Questions or answers containing non-printable or control characters. Tabs inside fenced code and Asymptote [asy] blocks are ignored.
latex_escapes medium LaTeX commands inside math whose backslash was eaten by a Python string escape, such as \frac stored as a form feed followed by rac.
mojibake low/medium UTF-8 text decoded with the wrong codec (Windows-1252, Latin-1, Mac Roman), such as ‚Äì for –. Checks choices too and gives the repair.
binary_question_ratio low Datasets where a high proportion of questions are binary (yes/no).
markdown_integrity low/medium Structural problems in Markdown answers: table column-count mismatches, missing delimiter rows, heading jumps, empty image links.
extraction_artifacts low/medium Characters betraying un-cleaned PDF/OCR extraction: ligatures, soft hyphens, zero-width characters, U+FFFD.
text_layer_recall medium/high With --files-root: gold words no extraction tool found on the page (typo candidates); for full-page gold, words every tool found that gold omits.
numeric_provenance high With --files-root: numbers in the gold that no extraction tool extracted from the page — strong transcription-error candidates.

LLM Scanners (require --model)

Scanner Severity What it flags
ambiguity medium Questions that are ambiguous or underspecified — can be interpreted multiple ways.
label_correctness high Samples where the ground-truth answer appears to be factually incorrect.
answerability medium Questions that cannot be answered from the provided context (auto-detects context columns).

Output

When --output-dir is given, findings are written as:

findings/
    answer_length.json
    duplicate_questions.json
    inconsistent_format.json
    answer_distribution.json
    scan_summary.json    # counts by scanner/severity/category
    REPORT.md            # human-readable markdown

Each finding includes the scanner name, severity, category, explanation, sample index, sample ID (if available), and scanner-specific metadata.

scan_summary.json also has a scanner_status map with one entry per scanner that was run. A scanner that checked the dataset has {"status": "ran"}, whether or not it found anything. A scanner whose check does not fit the dataset has {"status": "not_applicable", "reason": "..."} and emits no findings. For example, answer_length and inconsistent_format measure answer text, so they do not apply to a list or struct answer column unless --answer-subfield selects a scalar inside it. The path is dotted, lists along it are measured element by element, and * measures each element of a list of strings. Plugin scanners can report the same status by raising inspect_dataset.ScannerNotApplicable(reason).

Most scanners also skip a dataset that lacks the input they check. A dataset with an empty answer on every row, such as IFEval, gets not_applicable from the scanners that read answers. image_mime_type needs an image field: --image-field for a HuggingFace or local scan, or at least one sample with an image in a task scan. text_layer_recall and numeric_provenance need --files-root. inspect-dataset scanners lists what each built-in scanner requires. A plugin scanner declares the same with requires, and the runner then records not_applicable with a standard reason without calling it:

from inspect_dataset import dataset_scanner


@dataset_scanner(description="Flag answers that repeat the question", requires=["answer"])
def answer_echoes_question(records, fields): ...

The names are "answer" (at least one row has a non-empty answer), "image" (an image field is set and at least one row has an image), and "artifacts" (at least one row has an extraction artifacts directory from --files-root). ScannerDef and LLMScannerDef take the same requires argument. An unknown name raises ValueError.

scan_summary.json also records how the population scanners (inconsistent_format, answer_distribution and binary_question_ratio) were grouped. On a benchmark made of subsets with different answer formats, one majority over the whole dataset is meaningless, and a balanced whole can hide one imbalanced subset. With a group field, these scanners compute their statistics per subset, and each finding names its group in metadata.group and in the explanation. Dataset-level findings become one per affected group. Groups with fewer than 20 answers are skipped by the two distribution scanners, where an imbalance is too likely to be chance. group_by is the field used, or null. group_by_source is "option" for --group-by FIELD, "auto" when task mode detected it, or null when there is no grouping or a Python caller did not pass group_by_source to run_scanners. Task mode groups automatically when exactly one of the Sample.metadata keys dataset_name, subset, subject or category has two or more values, for example dataset_name in BBH and subject in MMLU-Pro. --no-group-by turns that off. In Python, load_inspect_task sets FieldMap.group to the detected key, so set fields.group = None before run_scanners for pooled statistics. HuggingFace mode never groups automatically, because one config is already one subset.

scan_summary.json records the split and config that were scanned. For a HuggingFace dataset these are filled in when not given: the only split, or train when there are several, and the only config or the dataset's default one. split_defaulted and config_defaulted say whether each was filled in (true) or given (false). Both are null for task and local scans. When the dataset has several splits and none is train, or several configs and no default, the scan stops and lists the choices.

scan_summary.json records where the findings came from. version is the inspect-dataset version that ran the scan. When the dataset is an inspect_ai task, task is the task spec as given on the command line and scorers lists the registry names of the task's scorers, such as ["inspect_ai/choice"]. A task with no scorer has [], and scorers that are not inspect_ai registry objects are left out. Both are null for HuggingFace and local scans.

scan_summary.json records the split and config that were scanned. For a HuggingFace dataset these are filled in when not given: the only split, or train when there are several, and the only config or the dataset's default one. split_defaulted and config_defaulted say whether each was filled in (true) or given (false). Both are null for task and local scans. When the dataset has several splits and none is train, or several configs and no default, the scan stops and lists the choices.

inspect_ai tasks

A task spec such as inspect_evals/drop loads the task's samples. Each record has input (the text of the last user message), target (the first target string), targets (every target string, as a list) and id, plus choices when the sample has them and every key of the sample's metadata. input, target and id are the question, answer and id fields. When any sample has choices, scanners find them through FieldMap.choices, which lets them read a letter target as the choice it names. When any sample's input has images, images holds every image from all its messages and is the image field, so duplicate_questions compares images and image_mime_type checks them. Image files are read into {"bytes": ..., "path": ...} values, the shape of a HuggingFace image column. URLs keep only their path, and data URIs stay as strings. To measure every string of a list target, such as DROP's alternative answers, pass --answer-field targets --answer-subfield '*'.

answer_length and inconsistent_format assume a scorer that compares the answer text verbatim, so its length, capitalisation and punctuation affect the score. In task mode they run only when one of the task's scorers is in VERBATIM_SCORERS: inspect_ai/exact, inspect_ai/match, inspect_ai/includes and inspect_ai/pattern. Under any other scorer, such as inspect_ai/choice, a model grader or code execution, they are recorded as not applicable with the scorer names as the reason. A task with no scorer is not applicable too. HuggingFace and local scans have no scorer to check, so these scanners always run there. To run them under another text-comparing scorer, add its registry name to VERBATIM_SCORERS in inspect_dataset/scanners/_answers.py.

answer_distribution and binary_question_ratio measure a letter target as the text of the choice it names, so a multiple-choice task with the choices Yes and No counts yes and no answers rather than A and B. answer_distribution also checks the target letters, because one letter holding 85% of targets is an answer-position bias under choice(). Their findings carry measured in the metadata, "choice_text" or "letter", and say which one in the explanation. Without choices, or when no target is a letter that names a choice, both scanners measure the answer as it is and add no measured key.

Each task record also carries the raw dataset row its sample came from, under __source__. inspect-dataset records the rows while the task builds its dataset. Samples built by inspect_ai's hf_dataset, csv_dataset or json_dataset get their exact row, even when the dataset is shuffled. Samples an eval builds by hand from datasets.load_dataset are matched to a row by their id, using whichever column of the loaded tables holds unique values that match the most ids. An id match only counts when some text of the row also appears in the sample's input, targets or choices, so positional ids cannot attach the wrong row. In a field option or --group-by, source.<column> names a column of that row, so any scanner can check a field the eval never puts in the sample:

# Check MATH's worked solutions, which the task does not show the model
inspect-dataset scan inspect_evals/math --answer-field source.solution --scanners encoding_issues,latex_escapes

# Check MMLU-Pro's answer positions within each original source dataset (its raw src column)
inspect-dataset scan inspect_evals/mmlu_pro --group-by source.src --scanners answer_distribution

scan_summary.json has a source key for task scans: how many samples were joined (joined, total), by record and by id (joined_by_record, joined_by_id, id_column), and every datasets.load_dataset and hf_dataset call the task made, with its config, split and revision (loads). hf_dataset calls are recorded even when inspect_ai reads the dataset from its own cache. A dataset built with a custom builder script, as PIQA, MedQA and BBQ are, records no load. It is null for HuggingFace and local scans. Plugin scanners can read record["__source__"]. samples.json does not include the raw rows.

Integration with inspect-scout

inspect-scout tracks which samples models consistently fail or succeed on. inspect-dataset provides a complementary static pass before running evals. A future release will accept inspect-scout results directly to produce eval-informed findings and a clean_ids.txt export for quality-adjusted benchmark scores.

Releasing

Releases publish to PyPI via trusted publishing, with no API tokens. One-time setup: add a Trusted Publisher on PyPI for Generality-Labs/inspect_dataset, workflow publish.yml, environment pypi.

  1. Actions → Prepare release → Run workflow, choosing the bump (auto picks minor for new or changed features, patch for fixes only). It bumps version in pyproject.toml, collects changelog.d/ into CHANGELOG.md, and opens a Release vX.Y.Z pull request.
  2. On that pull request's Checks tab, click Approve workflows to run, then review it. Add a summary paragraph under the new heading if the release needs one.
  3. Merge it. Release on merge tags the merge commit, creates the GitHub release, and starts publish.yml, which checks the tag against the built wheel and uploads to PyPI.

The workflows are thin callers of python-project-template's reusable ones. By hand, the same is uv version --bump minor, uv run scriv collect, a pull request, and then git tag vX.Y.Z && git push origin vX.Y.Z on the merge commit, which starts publish.yml, plus gh release create vX.Y.Z for the GitHub release.

Development

uv sync                     # the dev group is installed by default
uv run pytest

Each pull request adds a changelog fragment rather than editing CHANGELOG.md, so concurrent PRs don't conflict. Run uv run scriv create, uncomment the sections that apply in the new file under changelog.d/, and commit it with the change, or delete it if the change needs no entry. To release, run the Prepare release workflow from the Actions tab. It bumps version in pyproject.toml, collects the fragments into CHANGELOG.md, and opens a release pull request, whose CI starts once you click Approve workflows to run on it; merging it tags the release. It needs Settings → Actions → General → Allow GitHub Actions to create and approve pull requests. By hand, the same is uv version --bump minor and uv run scriv collect.

If you are working on the interactive viewer itself, also install frontend dependencies and build the bundle:

cd src/inspect_dataset/_view/www
npm install
npm run build

About

Dataset quality scanner for AI evaluation benchmarks

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages