Shenxi Wu*, Yuhong Liu*, Haosong Zhang, Tongjin Zou, Yanxun Zhang,
Gaochang Chen, Dun Liang, Jiaqi Wang, Zhecan James Wang, Yuhang Zang, Dahua Lin†
* Equal contribution. † Corresponding author.
- 2026-09-28: We released the supporting code for benchmark authoring and review, SciDocIR construction, training-data generation, and evaluation, with an offline quick start and paper-to-code map.
- 2026-09-15: We released the complete SciDocBench benchmark and document assets on Hugging Face.
- 2026-09-12: We released the SciDocBench SFT and RL training data on Hugging Face.
- 2026-09-04: Our SciDocBench paper is available on arXiv.
- 2026-09-04: The official repository is initialized.
- 2026-09-04: SciDocBench is integrated into VLMEvalKit. We recommend using VLMEvalKit for standardized inference and evaluation.
Scientific papers combine text, equations, figures, tables, appendices, citations, code, and datasets. Reliable scientific-document assistants must therefore do more than retrieve visible text: they must locate evidence, verify numerical and logical relations, recover scientific structure, integrate information across documents, and produce reusable outputs.
SciDocBench evaluates selected evidence-grounded operations and their explicit outputs, from locating support to cross-paper synthesis. It contains:
- 124 manually designed and difficulty-screened questions;
- 7 scientific-document capability groups and 19 subtasks;
- 5 scientific domains;
- 2 question languages: English and Chinese;
- 2 document representations: All Images First and Markdown Interleaved;
- 496 matched evaluation instances in total;
- 3 evaluator families: rule-based, LLM-as-a-judge, and execution-based evaluation.
Each question is instantiated under four matched settings while preserving its task semantics and evaluation criteria:
| Setting | Question language | Document representation |
|---|---|---|
| EN-AF | English | All Images First |
| EN-IL | English | Markdown Interleaved |
| ZH-AF | Chinese | All Images First |
| ZH-IL | Chinese | Markdown Interleaved |
| Group | Capability |
|---|---|
| A | Document Perception and Structure |
| B | Scientific Information Extraction |
| C | Evidence Alignment and Verification |
| D | Cross-Document Understanding |
| E | Reconstruction and Execution |
| F | Paper-Code Alignment |
| G | Dataset Understanding |
Overall scores are computed over all 496 evaluation instances on a 0-100 scale. Failed or unusable responses receive zero.
| Rank | Model | Overall | EN-AF | EN-IL | ZH-AF | ZH-IL |
|---|---|---|---|---|---|---|
| 1 | Claude Opus 5 | 62.60 | 63.19 | 61.34 | 65.21 | 60.68 |
| 2 | GPT-5.6-Sol | 61.00 | 61.07 | 60.32 | 61.32 | 61.31 |
| 3 | Gemini 3.6 Flash | 59.87 | 59.19 | 62.98 | 59.68 | 57.61 |
| 4 | Qwen3.8-Max | 57.15 | 59.55 | 55.33 | 58.13 | 55.57 |
| 5 | Qwen3.7-Plus | 55.88 | 57.17 | 53.34 | 57.57 | 55.42 |
| 6 | GPT-5.6-Terra | 54.19 | 51.88 | 55.69 | 52.83 | 56.37 |
| 7 | Claude Opus 4.8 | 53.41 | 53.63 | 52.43 | 54.13 | 53.46 |
| 8 | Qwen3.8-27B | 52.89 | 58.65 | 47.25 | 54.21 | 51.46 |
| 9 | GPT-5.6-Luna | 50.58 | 47.55 | 53.99 | 46.65 | 54.13 |
| 10 | Kimi K2.5 | 50.38 | 54.74 | 45.42 | 55.04 | 46.33 |
| 11 | Claude Sonnet 4.6 | 49.57 | 52.25 | 51.18 | 53.62 | 41.21 |
| 12 | GLM-4.6V | 42.44 | 43.30 | 46.70 | 38.21 | 41.56 |
| 13 | MiMo-V2.5 | 40.16 | 42.26 | 39.26 | 44.66 | 34.46 |
Claude Opus 5 currently ranks first with 62.60, followed by GPT-5.6-Sol with 61.00 and Gemini 3.6 Flash with 59.87. No evaluated model reaches 63. Capability leaders are distributed across model families, and overall scores can conceal substantial differences between scientific-document skills.
The complete SciDocBench benchmark is publicly available on Hugging Face. The release contains 496 evaluation instances derived from 124 semantic questions, with 124 instances in each language-representation setting. It includes 7,052 image references and 2,758 unique document images after SHA-256 deduplication.
The repository provides a Hugging Face Parquet test split, a portable VLMEvalKit-compatible SciDocBench.tsv, content-addressed image assets, and integrity manifests with file and image checksums. Download it with:
hf download HenryExcellent/SciDocBench \
--repo-type dataset \
--local-dir SciDocBenchFor an offline VLMEvalKit installation, materialize the benchmark under LMUData:
export LMU_DATA=/path/to/VLMEvalKit/LMUData
mkdir -p "$LMU_DATA"
cp SciDocBench/SciDocBench.tsv "$LMU_DATA/SciDocBench.tsv"
mkdir -p "$LMU_DATA/images/SciDocBench"
cp -a SciDocBench/images/SciDocBench/. "$LMU_DATA/images/SciDocBench/"The SciDocBench Training Data release provides benchmark-oriented multimodal data for supervised fine-tuning and reinforcement learning:
| Dataset | Train | Validation | Total |
|---|---|---|---|
| SFT | 3,844 | 80 | 3,924 |
| RL | 10,056 | 87 | 10,143 |
| Total | 13,900 | 167 | 14,067 |
The SFT data contain 981 seeds, each expanded into four language-representation settings. Across SFT and RL, the release contains 55,603 unique images and 120,489 image references. It preserves ordered multimodal conversations, bilingual questions, document-input variants, SFT reasoning traces, and RL solutions used by the reward implementation. JSONL files use portable, content-addressed image paths; images are distributed in tar shards.
Download and materialize the dataset with:
hf download HenryExcellent/SciDocBench-Training-Data \
--repo-type dataset \
--local-dir SciDocBench-Training-Data
cd SciDocBench-Training-Data
for shard in image_shards/images-*.tar; do
tar -xf "${shard}"
done
python prepare_remote.py --check-only --hash-images
python prepare_remote.pyThe preparation script verifies assets and writes machine-local training files under runtime/, including LLaMA-Factory registration and ms-swift RL inputs. The portable source files remain under data/sft/ and data/rl/. See the dataset card for field definitions and licensing notes, and experimental settings for the training mixtures used in the reported experiments.
The supporting materials provide benchmark authoring and review interfaces, document processing, task generation, training-data assembly, and the SciDocBench evaluation adapter.
| Component | Code and documentation |
|---|---|
| Expert question authoring and review | Annotation UI, blind grading review |
| arXiv sources and SciDocIR | Construction runbook, portable IR implementation |
| Task generation and four input settings | Task instructions, source pipeline index |
| SFT and GRPO | Experiment settings, GRPO preparation and reward |
| Benchmark evaluation | VLMEvalKit adapter, question-level evaluator inventory |
The paper-to-code map identifies the implementation and required inputs for each pipeline stage. To run the portable construction demo without an API or GPU:
git clone https://github.com/InternLM/SciDocBench.git
cd SciDocBench/supporting_materials
python3 -m venv .venv
source .venv/bin/activate
python -m pip install .
python -B -m unittest discover -s tests -v
python -B -m scidoc_support.cli demo --output ../scidoc_synthetic_demoThe demo uses synthetic fixtures. Document assets, model weights, and API credentials are configured separately for full construction and training workflows.
SciDocBench is available in VLMEvalKit, which is the recommended evaluation entry point. After downloading the benchmark automatically or materializing the offline release as shown above, run:
python run.py --data SciDocBench --model <MODEL_NAME> --verboseReplace <MODEL_NAME> with a model registered in VLMEvalKit. The command performs inference and evaluation. Configure the credentials required by the selected model and semantic judge according to the VLMEvalKit documentation.
The evaluation guide documents final-answer extraction, question-specific rule and judge routing, and the execution evaluator. The supporting adapter pins the public benchmark revision and checksum. Review local changes before copying it into an existing VLMEvalKit checkout.
SciDocBench: A Workflow-Centered Benchmark and Data Pipeline for Scientific Document Understanding
Shenxi Wu, Yuhong Liu, Haosong Zhang, Tongjin Zou, Yanxun Zhang, Gaochang Chen, Dun Liang, Jiaqi Wang, Zhecan James Wang, Yuhang Zang, and Dahua Lin.
| Component | Status |
|---|---|
| Paper and benchmark description | Available on arXiv |
| Standardized evaluation | Available through VLMEvalKit |
| SciDocBench data and document assets | Available on Hugging Face |
| SciDocBench SFT and RL training data | Available on Hugging Face |
| SciDocIR preprocessing and data-generation code | Supporting code |
| Reproduction documentation | Runbook, paper-to-code map |
Please cite our paper as follows:
@article{wu2026scidocbench,
title = {SciDocBench: A Workflow-Centered Benchmark and Data Pipeline for Scientific Document Understanding},
author = {Wu, Shenxi and Liu, Yuhong and Zhang, Haosong and Zou, Tongjin and Zhang, Yanxun and Chen, Gaochang and Liang, Dun and Wang, Jiaqi and Wang, Zhecan James and Zang, Yuhang and Lin, Dahua},
journal = {arXiv preprint arXiv:2609.05141},
year = {2026},
eprint = {2609.05141},
archivePrefix = {arXiv},
primaryClass = {cs.AI},
doi = {10.48550/arXiv.2609.05141},
url = {https://arxiv.org/abs/2609.05141}
}We thank the VLMEvalKit team for providing the standardized evaluation framework.
The released benchmark and training annotations are governed by the terms described in the benchmark dataset card and training dataset card. Paper pages and figures may retain source-specific copyright or license terms. Users are responsible for checking the applicable source terms before redistribution or commercial use.
See the supporting-code notices for third-party components and attribution.
