Repository navigation
refactor(benchmarks): retire Harbor-backed Kepler implementations (1/2) - #113
abhinav-pola wants to merge 2 commits into
Conversation
|
I'll fix CI failures and address comments from users with write access that start with 'DevinAI' or '@devin'.
Original prompt from Abhinav
|
…e_atlas_*) and the harbor/agent-cli/sandbox support code
6f323eb to
c635123
Compare
TL;DR
Deletes the six native benchmark implementations that are covered by Harbor (upstream registry or a canonical Harbor-format task repo run by gcp-harbor-trial), plus the harbor/agent-cli/sandbox support code that only they used.
gpqa_diamondstays.What changed?
terminal_bench,deep_swe,wandr,swe_atlas_qa,swe_atlas_tw,swe_atlas_rffrom the registry, meta table, config union, options map, and CLI dispatch.src/benchmarks/{terminal-bench,deep-swe,wandr,swe-atlas,harbor,agent-cli},src/sandbox,scripts/agent-runtime, andtest/helpers/terminal-bench-sandbox.ts. Every remaining import was checked; none referenced these.agentReasoningEffortderivation inbuildBenchmarkConfig(no remaining options schema has that field).BENCH_CHILD_WORKFLOW_IDcontrol-character check previously imported fromagent-cli/runnernow lives insrc/cli/index.ts(resolveSessionId), since the id still becomes thex-session-idheader for every run. Covered by new tests.package.json: removed the correspondingexportsentries, thepackage-agent-runtimescript, and themodal/smol-tomldependencies.bun.lockregenerated viabun install.registry.test.ts, 11 incli/index.test.ts).Rebased onto main (bd5c6d0)
toolcall_formats,maxreasoning effort, the Kepler rename and other main changes are kept. These main PRs only touched the deleted dirs, so their changes are removed along with them: #121 (src/sandbox/modal-schema.tsModal region validation), #124 (pi billed cost inagent-cli/runner.ts), #126 (auto-router cost tier forwarding, incl. newagent-cli/request-plugins.ts), #127 (switchyard-router plugin forwarding), #128 (server-tool child cost in pi billed cost), #129 (Switchyard candidate models to pi), #131 (pi-reported cost without resolver). Nothing outside the deleted dirs imports any of that code.Why?
These benchmarks are run through the Harbor pipeline in openrouter-web (
services/gcp-harbor-trial/.../benchmarks.py), pointing at the canonical task repos (harbor-framework/terminal-bench*, datacurve-ai/deep-swe, perplexityai/wandr, scaleapi/SWE-Atlas). Maintaining a second TypeScript implementation duplicates that. Per Slack discussion,gpqa_diamondis kept even though it is in the upstream registry.Stack
Native adapter work: Harbor #3595. The follow-up is draft and blocked on parity, dataset publication, and runtime integration; it is not permission to remove Tau3 now.
How to test
Expected: all pass (locally 1356 tests, 0 failures). Running the CLI with
--benchmark terminal_benchnow errors with an unknown benchmark id.Benchmark impact
No scoring changes for remaining benchmarks. The six removed ids can no longer be run from this harness.
Reviewer focus
src/cli/index.ts: the removedagentReasoningEffortderivation is dead without the deleted benchmarks.package.jsonexports (./sandbox*,./benchmarks/harbor/*,./benchmarks/agent-cli/*,./benchmarks/{deep-swe,swe-atlas,terminal-bench}/schema).Checklist
.agents/skills/add-benchmark/SKILL.mdstill cites terminal-bench/deep-swe as examples)Link to Devin session: https://openrouter.devinenterprise.com/sessions/488f5985ba004657be7909e99df2729e
Open in Devin Desktop: https://openrouter.devinenterprise.com/desktop/session/488f5985ba004657be7909e99df2729e?variant=devin