From 8a04671a6b742681d2500a8912cdcf360861a5d5 Mon Sep 17 00:00:00 2001 From: MagMueller Date: Fri, 25 Sep 2026 07:03:33 +0000 Subject: [PATCH] Restore prominent benchmark scores in README --- README.md | 12 ++++++------ 1 file changed, 6 insertions(+), 6 deletions(-) diff --git a/README.md b/README.md index e88dbe4..7aaae60 100644 --- a/README.md +++ b/README.md @@ -32,6 +32,12 @@
+## Benchmark scores — historical 60-task results + +Legacy 60-task BU Bench V2 results - Mean rubric score by model and cost per task, including GPT-6 Astra + +These results use the earlier 60-task set. They are not results for the current 200-task V2.1 dataset. Compare model scores only on the same task set and revision. + ## BU Bench V2.1 **200 web tasks scored against weighted findings rubrics — the default task set.** @@ -91,12 +97,6 @@ saved-evidence semantic validation remains pending. The tasks are encrypted to keep their text out of web crawlers and model training data. Please do not publish decrypted tasks or rubrics in plaintext or use them for model training. -### Historical results — 60 tasks - -Legacy 60-task BU Bench V2 results - Mean rubric score by model and cost per task, including GPT-6 Astra - -These results use the earlier 60-task set. They are not results for the current 200-task V2.1 dataset. Compare model scores only on the same task set and revision. - ### Running BU Bench V2.1 (default) The default entry point runs all 200 tasks with [BrowserCode](https://bcode.sh/)