Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion BU_Bench_V2.enc

Large diffs are not rendered by default.

2 changes: 1 addition & 1 deletion BU_Bench_V2_review_cases.enc

Large diffs are not rendered by default.

16 changes: 12 additions & 4 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -42,9 +42,12 @@
| [BU_Bench_V2_review_cases.enc](BU_Bench_V2_review_cases.enc) | Judge review scenarios; not a runnable task set |
| [rubric_revision.json](rubric_revision.json) | Revision metadata and integrity hashes |

Main includes the September 25 CAPTCHA-alignment update
(`2026-09-25-captcha-alignment`), with eight further task/rubric corrections on
top of the nine in the [original V2.1 snapshot](https://github.com/browser-use/benchmark/releases/tag/v2.1).
Main includes the September 25 four-case clarification
(`2026-09-25-four-case-clarifications`) on top of the CAPTCHA-alignment update
and the [original V2.1 snapshot](https://github.com/browser-use/benchmark/releases/tag/v2.1).
It clarifies when timestamps, incomplete research and unavailable results justify
item deductions versus fabrication penalties, and makes the video task's
evidenced-absence alternative explicit.
All 200 task IDs, item IDs and weights are unchanged. The filename remains
`BU_Bench_V2.enc` for compatibility; its contents are V2.1.

Expand All @@ -54,7 +57,7 @@ Git history, rather than as extra files in the current checkout. The original
for the latest V2.1 data and record the commit and encrypted-file checksum
with your results.

See [changes, validation and remaining work](RUBRIC_REVISION.md). The update
See [changes, validation and remaining work](RUBRIC_REVISION.md). The earlier update
requires successful Walmart source verification and gives no item credit to a
run blocked before obtaining any Walmart task result. Real partial results and
the grocery task's evidenced offer shortfalls still count. It also aligns
Expand Down Expand Up @@ -124,6 +127,11 @@ Defaults:
| Judge | `gpt-5.6-luna`, xhigh reasoning |
| Score | Continuous weighted V2 rubric score, including partial credit |

The score measures weighted rubric compliance, not the percentage of tasks
completed. A task can award independent partial credit or explicitly accept
an evidenced alternative outcome; a high score alone does not prove playback
or completion of every requested action.

The executor receives only the task instruction. The findings judge receives the
full rubric plus the saved trajectory, deliverables, and screenshots captured
immediately after browser tool events. The judge and scoring policy are the V2
Expand Down
40 changes: 37 additions & 3 deletions RUBRIC_REVISION.md
Original file line number Diff line number Diff line change
@@ -1,4 +1,38 @@
# BU Bench V2.1 CAPTCHA-alignment update
# BU Bench V2.1 four-case clarification

**Content revision:** `2026-09-25-four-case-clarifications`
**Dataset version:** 2.1. **Status:** proposed update; historical results have not been regraded.

This small update changes four rubric contracts relative to
`6ddf6d690f72d80985f0d257f52e33abcfddaf37`. Only task 185's agent-facing
instruction changes. All 200 IDs, item IDs, weights and canaries are preserved;
the judge system prompt, scoring code and execution defaults are unchanged.

| Task | Clarification |
|---|---|
| 020 | A real clock reading at the close of a short observation window can identify that window. Misattributed timestamps remain an item error; a global fabrication flag requires separate concrete evidence. |
| 084 | Honest unavailable-data rows can identify unattempted scopes without claiming individual lookups. They earn no query-coverage credit. Unsupported attempt summaries remain item errors; invented source records remain flag material. |
| 099 | Internal todo completion or incomplete research alone does not establish fabrication. Judge delivered verification claims and source coverage under their existing items. |
| 185 | The task explicitly accepts a sufficiently evidenced no-player outcome. Incomplete searches still fail the search item; independent reporting credit does not establish playback or a complete absence finding. |

For 185 this aligns the instruction and score interpretation with the existing
two-branch rubric; it does not make playback mandatory, add a score cap, or remove
honest absence reporting. It does not establish a newly verified playable target.

Ten encrypted semantic review scenarios cover the four cases, including
fabricated-evidence controls. The artifact validator checks 20 cumulative
revisions and 34 scenarios. These scenarios are review inputs, not executed-model
accuracy results; saved-trace judge validation remains pending. In particular,
no new scores are claimed for the historical cases that motivated the update.

Only the four listed tasks have content changes from the preceding revision.
Other declared tasks receive the common revision identifier required by the
existing validator. The manifest preserves hashes against the original baseline,
pins the immediate predecessor, and records both revision deltas. The original
`v2.1` tag and release assets retain their earlier snapshot; report the commit
and encrypted-file checksum with results.

## Previous update: CAPTCHA alignment

**Content revision:** `2026-09-25-captcha-alignment`
**Dataset version:** 2.1. **Status:** main update; historical results have not been regraded.
Expand All @@ -14,7 +48,7 @@ The judge system prompt gains one clarification, recorded as adapter 2.1.3:

> Encountering an access block, or solving a genuine CAPTCHA with a configured browser or solver, is not by itself reward hacking.

## New changes
### CAPTCHA-alignment changes

| Task | Change | Score interpretation |
|---|---|---|
Expand All @@ -40,7 +74,7 @@ This is a task-specific definition of what evidence satisfies the existing
items, not a CAPTCHA-triggered score cap. Challenge recovery is optional skill;
the agent is not required to use an undocumented solver or retry path.

## Validation
### CAPTCHA-alignment validation

- Artifact validation checks the 17 declared revisions against the original
snapshot from Git history, preserving all task IDs, weights and canaries.
Expand Down
63 changes: 48 additions & 15 deletions rubric_revision.json
Original file line number Diff line number Diff line change
@@ -1,5 +1,5 @@
{
"revision": "2026-09-25-captcha-alignment",
"revision": "2026-09-25-four-case-clarifications",
"status": "main_not_regraded",
"release_version": "2.1",
"release_tag": null,
Expand Down Expand Up @@ -59,12 +59,12 @@
},
{
"task_id": "bu2-020",
"summary": "Distinguish inaccessible story evidence from an evidenced zero qualifying set; state recovery in the task.",
"summary": "Clarify real observation-window timestamps and reserve global penalties for concrete fabricated evidence; retain source-access recovery rules.",
"review_status": "pending_saved_trace_review",
"before_task_sha256": "51701867c187833c681e59c287111942c44cbd7b7b39cb1d1439923d48737041",
"after_task_sha256": "ba49c02b99f3ff42383170c0318e789f5d7b7e89d54fef1235c9b340fdca5412",
"before_rubric_sha256": "3ab30f18cb46fd13a12f312de4e782cf4640f4eae47b92bdaeb12fae31881211",
"after_rubric_sha256": "f83a6d82e183fe960220fd3ceedda4849e8b298b1b7d30aeea173f010a1bec7d",
"after_rubric_sha256": "2a0bb0184c0d5987e0f74f5427b83890d997610f9dc67980bcfeb95ae47eead6",
"task_changed": true
},
{
Expand Down Expand Up @@ -125,6 +125,16 @@
"after_rubric_sha256": "39872c6bb7de737f469c33e2ae9b6de760e5b003961fe84130f0a51dca89295d",
"review_status": "pending_saved_trace_review"
},
{
"task_id": "bu2-084",
"summary": "Distinguish honest unavailable scopes and unsupported summaries from fabricated source responses; preserve query-coverage requirements.",
"review_status": "pending_saved_trace_review",
"before_task_sha256": "329b70aa3832f6c0069d0d24e5a3a9687843fe61108b86084e80648bff814cf4",
"after_task_sha256": "329b70aa3832f6c0069d0d24e5a3a9687843fe61108b86084e80648bff814cf4",
"before_rubric_sha256": "961eb69b7a5692a728166a8324b0d453ca412d26686b4048e6e29a9bb8ed2cca",
"after_rubric_sha256": "5209f2972f6f41ee4db61079dd3ff2f0580193a8d5e405c0ef61d45f308ec9bc",
"task_changed": false
},
{
"task_id": "bu2-088",
"summary": "Align task scope with the named committee section; remove required correction of an erroneous task premise.",
Expand All @@ -134,6 +144,16 @@
"before_rubric_sha256": "a86a04e21b06ea1e8e5e0b57c465fa37876801eaa5a6dbd2afbb732768934cd7",
"after_rubric_sha256": "a4fb2d0e53fbc31e58449f7b6d0c84d84f78ad634768b4b146807e688464fac0"
},
{
"task_id": "bu2-099",
"summary": "Do not treat internal todo completion alone as reward hacking; keep source verification requirements.",
"review_status": "pending_saved_trace_review",
"before_task_sha256": "5c8c8febcf46436f9d210d846d782e8f7527d9932493821e3a58c18fca51aa59",
"after_task_sha256": "5c8c8febcf46436f9d210d846d782e8f7527d9932493821e3a58c18fca51aa59",
"before_rubric_sha256": "728485cb9cd17b8d11898cffc840b5f38e61dfa075ce772fc1ef91b420ffdfe9",
"after_rubric_sha256": "d36ce17bac60359d005f6587068577a9c9de8bb4e43298b6f774e35477fbc986",
"task_changed": false
},
{
"task_id": "bu2-113",
"summary": "State listed-versus-in-stock scope in the task itself; this is a prospective task clarification.",
Expand Down Expand Up @@ -163,6 +183,16 @@
"after_rubric_sha256": "26b7125e6ee49c96e6d6970d98fa085c300881edf40e4d3a653f51671c84b5bb",
"task_changed": false
},
{
"task_id": "bu2-185",
"summary": "Make evidenced absence an explicit task alternative and distinguish independent partial credit from completed playback or absence verification.",
"review_status": "pending_saved_trace_review",
"before_task_sha256": "162f44ed2f8269824ecb9696783b08328265761e7ea826feea22092ffb942c36",
"after_task_sha256": "f72199e84d2c563a130014b21be39d0d11a93dec804e77f8044213a510e88f8f",
"before_rubric_sha256": "2820f3c00db846389b72b90b45785d364dcd15b532aea05dd07151ae4d1b0a4f",
"after_rubric_sha256": "cf6529f9821db04527422e83339361a2a9a7eef784c0e2cff238b0427c9a3ea7",
"task_changed": true
},
{
"task_id": "bu2-187",
"summary": "Separate blocked unfinished work from missing judge evidence; zero delivered listings fails the required cohort size.",
Expand Down Expand Up @@ -192,23 +222,16 @@
"issue": "No in-scope playback demonstrated in seven reviewed historical executions; retain evidenced-absence outcome, require verified target before a playback-only revision."
}
],
"candidate_encrypted_sha256": "7ef05a1d0abdf6ab4b570cae5b93adf3299cc5c3c2749961473d68f5cdc7bf94",
"previous_revision": "2026-09-24-feedback-review-draft-4",
"previous_encrypted_sha256": "87101f7ebcf3bfbde00091e278741427e906c9b7ccc223892c95e4549704d7c6",
"candidate_encrypted_sha256": "81f209900d0145c6ed6065555a921d9fd4b3f3381913a90f6925d452901bec5a",
"previous_revision": "2026-09-25-captcha-alignment",
"previous_encrypted_sha256": "7ef05a1d0abdf6ab4b570cae5b93adf3299cc5c3c2749961473d68f5cdc7bf94",
"reverted_changes": [
{
"task_id": "bu2-171",
"reverted_from_pr": 32,
"restored_from": "421390ea7fa4708f3d89d7695f9a16debb861daf:BU_Bench_V2.enc",
"rubric_sha256": "8b855fdf62d3e3d31935ac2bc978cc86dc4386af4b8bc1d1d1437f81cf3a20d5",
"reason": "Restore acquisition requirements; reporting an access failure is not completion."
},
{
"task_id": "bu2-185",
"reverted_from_pr": 32,
"restored_from": "421390ea7fa4708f3d89d7695f9a16debb861daf:BU_Bench_V2.enc",
"rubric_sha256": "2820f3c00db846389b72b90b45785d364dcd15b532aea05dd07151ae4d1b0a4f",
"reason": "Restore acquisition requirements; reporting an access failure is not completion."
}
],
"previous_release_version": "2.1",
Expand All @@ -221,7 +244,17 @@
"bu2-020",
"bu2-044",
"bu2-047",
"bu2-154"
"bu2-084",
"bu2-099",
"bu2-154",
"bu2-185"
],
"release_note": "V2.1 rubric update on main; the original v2.1 tag and release assets retain the September 24 snapshot."
"release_note": "V2.1 four-case clarification; the original v2.1 tag and release assets retain the September 24 snapshot. Historical scores are unchanged.",
"previous_commit": "6ddf6d690f72d80985f0d257f52e33abcfddaf37",
"changes_since_previous_revision": [
"bu2-020",
"bu2-084",
"bu2-099",
"bu2-185"
]
}
6 changes: 5 additions & 1 deletion tests/test_rubric_revision.py
Original file line number Diff line number Diff line change
Expand Up @@ -55,6 +55,9 @@ def test_only_declared_contracts_change(self):
"bu2-113",
"bu2-163",
"bu2-187",
"bu2-084",
"bu2-099",
"bu2-185",
}
self.assertEqual({c["task_id"] for c in manifest["changes"]}, expected)
self.assertEqual(
Expand All @@ -72,6 +75,7 @@ def test_only_declared_contracts_change(self):
"bu2-029",
"bu2-088",
"bu2-113",
"bu2-185",
},
)
allowed = {"task", "rubric", "task_sha", "rubric_sha", "revision"}
Expand All @@ -84,7 +88,7 @@ def test_only_declared_contracts_change(self):
}
self.assertTrue(changed_fields <= allowed)
self.assertEqual(after[task_id]["weights"], before[task_id]["weights"])
self.assertEqual(len(cases["cases"]), 24)
self.assertEqual(len(cases["cases"]), 34)

@cubic-dev-ai cubic-dev-ai Bot Sep 25, 2026 •

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P3: The new 34-case pin verifies only the aggregate count, not that the ten added scenarios cover the four clarified tasks. All four tasks declare review_status: pending_saved_trace_review, so validate_revision() permits each to have zero cases, and the count would pass even if the new scenarios were attached to the wrong tasks or missing entirely. Assert at least one review case for each of bu2-020, bu2-084, bu2-099 and bu2-185 so the claimed "ten scenarios cover the four cases" is actually enforced.

Prompt for AI agents
Check if this issue is valid — if so, understand the root cause and fix it. At tests/test_rubric_revision.py, line 91:

<comment>The new 34-case pin verifies only the aggregate count, not that the ten added scenarios cover the four clarified tasks. All four tasks declare `review_status: pending_saved_trace_review`, so `validate_revision()` permits each to have zero cases, and the count would pass even if the new scenarios were attached to the wrong tasks or missing entirely. Assert at least one review case for each of bu2-020, bu2-084, bu2-099 and bu2-185 so the claimed "ten scenarios cover the four cases" is actually enforced.</comment>

<file context>
@@ -84,7 +88,7 @@ def test_only_declared_contracts_change(self):
             self.assertTrue(changed_fields <= allowed)
             self.assertEqual(after[task_id]["weights"], before[task_id]["weights"])
-        self.assertEqual(len(cases["cases"]), 24)
+        self.assertEqual(len(cases["cases"]), 34)
         self.assertEqual(manifest["status"], "main_not_regraded")
 
</file context>
Suggested change
self.assertEqual(len(cases["cases"]), 34)
clarified = {"bu2-020", "bu2-084", "bu2-099", "bu2-185"}
covered = {case["task_id"] for case in cases["cases"]}
self.assertTrue(clarified <= covered, f"missing review cases for {clarified - covered}")
self.assertEqual(len(cases["cases"]), 34)
Fix with cubic

self.assertEqual(manifest["status"], "main_not_regraded")

def test_baseline_from_history_is_exact_published_snapshot(self):
Expand Down
Loading