Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
3 changes: 3 additions & 0 deletions .github/workflows/test.yml
Original file line number Diff line number Diff line change
Expand Up @@ -10,8 +10,11 @@ jobs:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- name: Fetch pinned rubric baseline
run: git fetch --no-tags --depth=1 origin 421390ea7fa4708f3d89d7695f9a16debb861daf
- uses: astral-sh/setup-uv@v6
- run: uv sync --frozen --python 3.12
- run: uv run python review_rubric_revision.py
- run: BCODE_TEST_CHROME="$(command -v google-chrome)" uv run python -m unittest discover -s tests
- run: uv run ruff check bcode_runner.py bcode_eval.py bcode_results.py tests/test_bcode_runner.py tests/test_bcode_binary.py
- run: go install github.com/rhysd/actionlint/cmd/actionlint@v1.7.7
Expand Down
64 changes: 0 additions & 64 deletions BU_Bench_V2_55.json

This file was deleted.

57 changes: 18 additions & 39 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -32,27 +32,27 @@

<br/>

## BU Bench V2
## BU Bench V2.1

**200 web tasks scored against weighted findings rubrics**
**200 web tasks scored against weighted findings rubrics — the default task set.**

| Evaluation set | Tasks | File |
| --- | ---: | --- |
| Full BU Bench V2 | 200 | [BU_Bench_V2.enc](BU_Bench_V2.enc) |
| BU Bench V2 — 55-task subset | 55 | [BU_Bench_V2_55.json](BU_Bench_V2_55.json) (task IDs only) |

The 55-task subset selects public IDs `bu2-001` through `bu2-055` from the full release. It is the public overlap with the original 60-task benchmark; five original tasks were omitted and the remaining tasks were renumbered. Use the public IDs in the subset file, not the first 55 historical IDs.

Tasks, rubrics and weights live once, in the encrypted 200-task release. `BU_Bench_V2_55.json` pins that file's SHA-256 and lists the subset IDs; it does not contain a second dataset. Label scores with the set used: **200 tasks**, **55-task subset**, or **legacy 60 tasks**.
| File | Purpose |
| --- | --- |
| [BU_Bench_V2.enc](BU_Bench_V2.enc) | All 200 current V2.1 tasks, rubrics and weights |
| [BU_Bench_V2_review_cases.enc](BU_Bench_V2_review_cases.enc) | Judge review scenarios; not a runnable task set |
| [rubric_revision.json](rubric_revision.json) | Revision metadata and integrity hashes |

**Dataset version: BU Bench V2.1.**
Main includes the September 25 CAPTCHA-alignment update
(`2026-09-25-captcha-alignment`), with eight further task/rubric corrections on
top of the nine in the [original V2.1 snapshot](https://github.com/browser-use/benchmark/releases/tag/v2.1).
All 200 task IDs, item IDs, weights and 55-task subset membership are unchanged.
The original `v2.1` tag and its download assets retain the September 24 snapshot;
use this revision's commit and encrypted-file checksum to identify the updated
V2.1 data. The filename remains `BU_Bench_V2.enc`.
All 200 task IDs, item IDs and weights are unchanged. The filename remains
`BU_Bench_V2.enc` for compatibility; its contents are V2.1.

The old 55-task selectors and original 200-task snapshot are available through
Git history, rather than as extra files in the current checkout. The original
`v2.1` tag and its download assets retain the September 24 snapshot; use main
for the latest V2.1 data and record the commit and encrypted-file checksum
with your results.

See [changes, validation and remaining work](RUBRIC_REVISION.md). The update
requires successful Walmart source verification and gives no item credit to a
Expand All @@ -67,34 +67,13 @@ saved-evidence semantic validation remains pending.

The tasks are encrypted to keep their text out of web crawlers and model training data. Please do not publish decrypted tasks or rubrics in plaintext or use them for model training.

<details>
<summary>Select the 55-task subset</summary>

After decrypting `BU_Bench_V2.enc` into a Python object named `benchmark`:

```python
import hashlib
import json
from pathlib import Path

subset = json.loads(Path("BU_Bench_V2_55.json").read_text())
assert hashlib.sha256(Path(subset["source"]).read_bytes()).hexdigest() == subset["source_sha256"]
by_id = {task["id"]: task for task in benchmark["tasks"]}
tasks = [by_id[task_id] for task_id in subset["task_ids"]]
assert len(tasks) == subset["task_count"]
```

This selects evaluation records, including the judge's rubric and weights. Pass only the task instructions to the evaluated agent; keep the rubric and weights for the judge.

</details>

### Historical results — 60 tasks

<img alt="Legacy 60-task BU Bench V2 results - Mean rubric score by model and cost per task, including GPT-6 Astra" src="official_plots/bu_bench_v2_astra.jpg" width="100%">

These results use the earlier 60-task set, not the full 200-task release or the 55-task subset. Compare model scores only on the same task set.
These results use the earlier 60-task set. They are not results for the current 200-task V2.1 dataset. Compare model scores only on the same task set and revision.

### Running BU Bench V2 (default)
### Running BU Bench V2.1 (default)

The default entry point runs all 200 tasks with [BrowserCode](https://bcode.sh/)
and the existing [findings judge](findings_judge.py). Anyone can clone this public
Expand Down Expand Up @@ -139,7 +118,7 @@ Defaults:
| Setting | Default |
| --- | --- |
| Executor | BrowserCode 0.1.20, `openai/gpt-6-luna`, xhigh reasoning |
| Tasks | All 200 in `BU_Bench_V2.enc` |
| Tasks | All 200 V2.1 tasks in `BU_Bench_V2.enc` |
| Browser | Browser Use Cloud, one session per task |
| Limits | Up to 100 concurrent tasks; 3,600 seconds per task |
| Judge | `gpt-5.6-luna`, xhigh reasoning |
Expand Down
32 changes: 26 additions & 6 deletions RUBRIC_REVISION.md
Original file line number Diff line number Diff line change
Expand Up @@ -6,8 +6,8 @@ The stable [V2.1 release](https://github.com/browser-use/benchmark/releases/tag/
and its downloaded assets are unchanged.

This V2.1 update changes eight task/rubric contracts relative to the original tag and retains
the nine earlier corrections. All 200 IDs, every item ID/weight/canary and the
55-task subset membership remain unchanged. Seven task instructions change;
the nine earlier corrections. All 200 IDs and every item ID/weight/canary
remain unchanged. Seven task instructions change;
016 changes only its rubric. The deterministic scorer, timeouts, screenshot
handling, reward-hacking calculation and historical scores are unchanged.
The judge system prompt gains one clarification, recorded as adapter 2.1.3:
Expand Down Expand Up @@ -43,7 +43,7 @@ the agent is not required to use an undocumented solver or retry path.
## Validation

- Artifact validation checks the 17 declared revisions against the original
snapshot, preserving all task IDs, weights, canaries and subset membership.
snapshot from Git history, preserving all task IDs, weights and canaries.
- Thirteen new encrypted semantic review scenarios supplement the existing
eleven. They cover all four fully blocked Walmart tasks, incomplete product
verification, valid grocery shortfalls, the observation-record distinction,
Expand Down Expand Up @@ -138,6 +138,26 @@ uv run python review_rubric_revision.py --write-private-diffs
```

The second command writes plaintext task/rubric diffs into ignored
`run_data/rubric-review/`. Keep those private. The original encrypted August 25
dataset remains in `snapshots/BU_Bench_V2_2026-08-25.enc`. This release adds no
score-cap machinery and makes no changes to historical results.
`run_data/rubric-review/`. Keep those private.

The validator reads the original encrypted August 25 dataset from
`BU_Bench_V2.enc` at the `base_commit` pinned in `rubric_revision.json` and
verifies its SHA-256. The historical dataset and 55-task selectors are no longer
duplicated in the checkout. Use a full Git clone for revision checks. For a
shallow clone, fetch the pinned commit; for a source archive, supply the original
encrypted baseline explicitly:

```bash
uv run python review_rubric_revision.py --base-artifact /path/to/original/BU_Bench_V2.enc
Comment thread
MagMueller marked this conversation as resolved.
```

Download the [original encrypted baseline at the pinned commit](https://github.com/browser-use/benchmark/raw/421390ea7fa4708f3d89d7695f9a16debb861daf/BU_Bench_V2.enc),
or extract it from a full clone with `git show <base_commit>:BU_Bench_V2.enc`.
Both paths verify `base_encrypted_sha256` from `rubric_revision.json`.

Without the pinned Git history, the unittest suite skips historical-comparison
checks with retrieval instructions; the standalone validator fails until you
supply the baseline. CI explicitly fetches that commit and runs the validator,
so those checks remain required there. Normal benchmark execution needs only
the current dataset and does not need Git history or the original baseline.
This cleanup changes no tasks, rubrics, scores or runtime defaults.
61 changes: 50 additions & 11 deletions review_rubric_revision.py
Original file line number Diff line number Diff line change
Expand Up @@ -6,6 +6,7 @@
import hashlib
import json
import re
import subprocess
from pathlib import Path

from cryptography.fernet import Fernet
Expand All @@ -17,9 +18,41 @@ def digest(text: str) -> str:
return hashlib.sha256(text.encode()).hexdigest()


def decrypt(path: Path, key_name: str) -> dict:
def decrypt_bytes(artifact: bytes, key_name: str) -> dict:
key = base64.urlsafe_b64encode(hashlib.sha256(key_name.encode()).digest())
return json.loads(Fernet(key).decrypt(base64.b64decode(path.read_bytes())))
return json.loads(Fernet(key).decrypt(base64.b64decode(artifact)))


def decrypt(path: Path, key_name: str) -> dict:
return decrypt_bytes(path.read_bytes(), key_name)


def read_base_artifact(
root: Path, manifest: dict, base_artifact: Path | None = None
) -> bytes:
"""Read the pinned baseline without keeping an obsolete dataset in the checkout."""
if base_artifact is not None:
artifact = base_artifact.read_bytes()
else:
commit = manifest["base_commit"]
if not re.fullmatch(r"[0-9a-f]{40}", commit):
raise ValueError("Base commit must be a full lowercase Git SHA")
try:
artifact = subprocess.run(
["git", "-C", str(root), "show", f"{commit}:BU_Bench_V2.enc"],
check=True,
capture_output=True,
).stdout
except (OSError, subprocess.CalledProcessError) as exc:
raise ValueError(
"Pinned baseline is absent from local Git history. "
"Use a full clone, fetch the base_commit from rubric_revision.json, "
"or pass --base-artifact /path/to/original/BU_Bench_V2.enc. "
"Normal benchmark execution does not need the historical baseline."
) from exc
if hashlib.sha256(artifact).hexdigest() != manifest["base_encrypted_sha256"]:
raise ValueError("Encrypted artifact hash mismatch: historical baseline")
return artifact


def index_tasks(tasks: list[dict], label: str) -> dict[str, dict]:
Expand All @@ -34,17 +67,18 @@ def index_tasks(tasks: list[dict], label: str) -> dict[str, dict]:
return {task["id"]: task for task in tasks}


def validate_revision(root: Path = ROOT) -> tuple[dict, dict, dict, dict]:
def validate_revision(
root: Path = ROOT, *, base_artifact: Path | None = None
) -> tuple[dict, dict, dict, dict]:
manifest = json.loads((root / "rubric_revision.json").read_text())
base_path = root / "snapshots/BU_Bench_V2_2026-08-25.enc"
candidate_path = root / "BU_Bench_V2.enc"
for path, field in [
(base_path, "base_encrypted_sha256"),
(candidate_path, "candidate_encrypted_sha256"),
if hashlib.sha256(candidate_path.read_bytes()).hexdigest() != manifest[
"candidate_encrypted_sha256"
]:
if hashlib.sha256(path.read_bytes()).hexdigest() != manifest[field]:
raise ValueError(f"Encrypted artifact hash mismatch: {path.name}")
original = decrypt(base_path, "BU_Bench_V2")
raise ValueError(f"Encrypted artifact hash mismatch: {candidate_path.name}")
original = decrypt_bytes(
read_base_artifact(root, manifest, base_artifact), "BU_Bench_V2"
)
candidate = decrypt(candidate_path, "BU_Bench_V2")
cases = decrypt(root / "BU_Bench_V2_review_cases.enc", "BU_Bench_V2_review_cases")
before = index_tasks(original["tasks"], "Base")
Expand Down Expand Up @@ -108,8 +142,13 @@ def validate_revision(root: Path = ROOT) -> tuple[dict, dict, dict, dict]:
def main():
parser = argparse.ArgumentParser(description=__doc__)
parser.add_argument("--write-private-diffs", action="store_true")
parser.add_argument(
"--base-artifact",
type=Path,
help="Original encrypted baseline for source archives or shallow clones; hash-verified.",
)
args = parser.parse_args()
manifest, before, after, cases = validate_revision()
manifest, before, after, cases = validate_revision(base_artifact=args.base_artifact)
print(
f"Validated {len(manifest['changes'])} revised tasks; other tasks and all weights unchanged."
)
Expand Down
4 changes: 2 additions & 2 deletions rubric_revision.json
Original file line number Diff line number Diff line change
Expand Up @@ -199,14 +199,14 @@
{
"task_id": "bu2-171",
"reverted_from_pr": 32,
"restored_from": "snapshots/BU_Bench_V2_2026-08-25.enc",
"restored_from": "421390ea7fa4708f3d89d7695f9a16debb861daf:BU_Bench_V2.enc",
"rubric_sha256": "8b855fdf62d3e3d31935ac2bc978cc86dc4386af4b8bc1d1d1437f81cf3a20d5",
"reason": "Restore acquisition requirements; reporting an access failure is not completion."
},
{
"task_id": "bu2-185",
"reverted_from_pr": 32,
"restored_from": "snapshots/BU_Bench_V2_2026-08-25.enc",
"restored_from": "421390ea7fa4708f3d89d7695f9a16debb861daf:BU_Bench_V2.enc",
"rubric_sha256": "2820f3c00db846389b72b90b45785d364dcd15b532aea05dd07151ae4d1b0a4f",
"reason": "Restore acquisition requirements; reporting an access failure is not completion."
}
Expand Down
1 change: 0 additions & 1 deletion snapshots/BU_Bench_V2_2026-08-25.enc

This file was deleted.

64 changes: 0 additions & 64 deletions snapshots/BU_Bench_V2_55_2026-08-25.json

This file was deleted.

Loading
Loading