Skip to content

refactor(benchmarks): retire Harbor-backed Kepler implementations (1/2) - #113

Open
abhinav-pola wants to merge 2 commits into
mainfrom
devin/1790205187-remove-harbor-benchmarks
Open

abhinav-pola wants to merge 2 commits into
mainfrom
devin/1790205187-remove-harbor-benchmarks

Conversation

@abhinav-pola

@abhinav-pola abhinav-pola commented Sep 23, 2026 •

Copy link
Copy Markdown
Contributor

TL;DR

Deletes the six native benchmark implementations that are covered by Harbor (upstream registry or a canonical Harbor-format task repo run by gcp-harbor-trial), plus the harbor/agent-cli/sandbox support code that only they used. gpqa_diamond stays.

What changed?

  • Removed benchmark ids terminal_bench, deep_swe, wandr, swe_atlas_qa, swe_atlas_tw, swe_atlas_rf from the registry, meta table, config union, options map, and CLI dispatch.
  • Deleted src/benchmarks/{terminal-bench,deep-swe,wandr,swe-atlas,harbor,agent-cli}, src/sandbox, scripts/agent-runtime, and test/helpers/terminal-bench-sandbox.ts. Every remaining import was checked; none referenced these.
  • Dropped the agentReasoningEffort derivation in buildBenchmarkConfig (no remaining options schema has that field).
  • The BENCH_CHILD_WORKFLOW_ID control-character check previously imported from agent-cli/runner now lives in src/cli/index.ts (resolveSessionId), since the id still becomes the x-session-id header for every run. Covered by new tests.
  • package.json: removed the corresponding exports entries, the package-agent-runtime script, and the modal / smol-toml dependencies. bun.lock regenerated via bun install.
  • Removed tests whose subject no longer exists (17 in registry.test.ts, 11 in cli/index.test.ts).

Rebased onto main (bd5c6d0)

toolcall_formats, max reasoning effort, the Kepler rename and other main changes are kept. These main PRs only touched the deleted dirs, so their changes are removed along with them: #121 (src/sandbox/modal-schema.ts Modal region validation), #124 (pi billed cost in agent-cli/runner.ts), #126 (auto-router cost tier forwarding, incl. new agent-cli/request-plugins.ts), #127 (switchyard-router plugin forwarding), #128 (server-tool child cost in pi billed cost), #129 (Switchyard candidate models to pi), #131 (pi-reported cost without resolver). Nothing outside the deleted dirs imports any of that code.

Why?

These benchmarks are run through the Harbor pipeline in openrouter-web (services/gcp-harbor-trial/.../benchmarks.py), pointing at the canonical task repos (harbor-framework/terminal-bench*, datacurve-ai/deep-swe, perplexityai/wandr, scaleapi/SWE-Atlas). Maintaining a second TypeScript implementation duplicates that. Per Slack discussion, gpqa_diamond is kept even though it is in the upstream registry.

Stack

  1. This PR removes the original six Harbor-backed implementations.
  2. #140 removes Tau3 Banking only after native replacement validation. MMMU Pro Vision and GPQA remain in Kepler.

Native adapter work: Harbor #3595. The follow-up is draft and blocked on parity, dataset publication, and runtime integration; it is not permission to remove Tau3 now.

How to test

bun install && bun run format:check && bun run check && bun run typecheck && bun test && bun run build

Expected: all pass (locally 1356 tests, 0 failures). Running the CLI with --benchmark terminal_bench now errors with an unknown benchmark id.

Benchmark impact

No scoring changes for remaining benchmarks. The six removed ids can no longer be run from this harness.

Reviewer focus

Checklist

  • Tests cover changed behavior
  • Public API or configuration changes are backward compatible, or the break is documented (break: removed package exports and benchmark ids, listed above)
  • Benchmark changes document dataset provenance and licensing (n/a, removals only)
  • No credentials, private results, or restricted dataset contents are included
  • Documentation is updated where needed (no docs referenced the removed ids; .agents/skills/add-benchmark/SKILL.md still cites terminal-bench/deep-swe as examples)

Link to Devin session: https://openrouter.devinenterprise.com/sessions/488f5985ba004657be7909e99df2729e
Open in Devin Desktop: https://openrouter.devinenterprise.com/desktop/session/488f5985ba004657be7909e99df2729e?variant=devin


Devin Review

@devin-ai-integration

Copy link
Copy Markdown
Contributor

I'll fix CI failures and address comments from users with write access that start with 'DevinAI' or '@devin'.

  • Disable automatic comment, CI, and merge conflict monitoring

Original prompt from Abhinav

SYSTEM:
<latest_message>
Abhinav Pola (U090K0G7JF3) [ts=1790204148.773039]: @Devin every benchmark that already exists in harbor should be deleted from <https://github.com/OpenRouterTeam/benchmark-harness|github.com/OpenRouterTeam/benchmark-harness>. whats the full list?
</latest_message>

=== BEGIN THREAD HISTORY (in #intern-benchy) ===
Abhinav Pola (U090K0G7JF3) [ts=1790204148.773039]: @Devin every benchmark that already exists in harbor should be deleted from <https://github.com/OpenRouterTeam/benchmark-harness|github.com/OpenRouterTeam/benchmark-harness>. whats the full list?
=== END THREAD HISTORY ===
Channel ID: C0BAKP8P5C3
Thread URL: https://openrouter.slack.com/archives/C0BAKP8P5C3/p1790204148773039?thread_ts=1790204148.773039&amp;cid=C0BAKP8P5C3

The <latest_message> is the message that you should use to guide your goals + task for this session, and you should use the rest of the slack thread as context.
A [ts=...] marker on a Slack message is that message's timestamp. To act on a specific message with the slack tool (e.g. adding an emoji reaction via the reaction command), pass that value as timestamp along with the Channel ID — no extra lookup call is needed.

devin-ai-integration[bot]

This comment was marked as resolved.

@devin-ai-integration
devin-ai-integration Bot force-pushed the devin/1790205187-remove-harbor-benchmarks branch from 6f323eb to c635123 Compare October 8, 2026 20:34
@ayush-or ayush-or changed the title Remove benchmarks now run through Harbor (terminal_bench, deep_swe, wandr, swe_atlas_*) refactor(benchmarks): retire Harbor-backed Kepler implementations (1/2) Oct 9, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant