Skip to content

Support direct remote logs and native Hawk auth in investigations - #23

Draft
jrh-mann wants to merge 8 commits into
devfrom
codex/dtbench-s3-reader
Draft

jrh-mann wants to merge 8 commits into
devfrom
codex/dtbench-s3-reader

Conversation

@jrh-mann

@jrh-mann jrh-mann commented Oct 2, 2026 •

Copy link
Copy Markdown
Contributor

Changes

Make explicit object-store and signed HTTP .eval sources usable by the investigation reader and child-audit staging. Cache direct inputs atomically, validate by parsing, namespace repeated sample IDs by source/item/epoch, and sanitize credentials in download errors.

Preserve the direct OpenRouter routing policy while refreshing Hawk-native credentials for control-plane operations when the generic provider hook is disabled. Bind the qualified openrouter/openai/... catalog entry through Inspect's builtin model configuration; this fixes child submission without changing the allowed model/route.

Document read_eval_log_samples(..., all_samples_required=False) for successful imported archives with unequal per-item trial counts. Such archives are sparse by design; strict rectangular streaming is inappropriate.

Verification

  • Twenty targeted reader, staging, auth and policy tests pass on the exact Hawk Inspect runtime.
  • Seven real remote source logs stage successfully into the registered audit task.
  • An actual Hawk investigator submitted, ran and collected its Sol child after these fixes; control-plane authentication and model binding are verified end to end.
  • A four-item focused Sol audit completed with substantive verdicts after adjusting its independent execution limits. This does not establish whole-benchmark semantic correctness.

Draft against dev. Existing development checkout and running jobs remain isolated from this worktree. No keys, presigned URLs or evaluation logs are included.

Per-attempt replay fix

DTBench's native logs exposed a second replay problem: score_async inherited the benchmark's custom epoch reducer, and overwrite mode also inherited its metric definitions. The auditor's categorical concordance scorer and numeric collector are different scorers and must use their own aggregates. The replay now uses mode_score() for concordance and mean() / mean_score() for the collector, while retaining the actual benchmark scorer and individual grades.

Verified on Hawk's exact Inspect runtime: 29 grader/concordance tests pass. A regression with unavailable custom metric/reducer names and three epochs fails without the overrides and passes with them. Native Nano replay reproduces all nine stored grades across three selected items, including invalid responses; authored correct/wrong/malformed outputs yield 1/0/NaN. Original source log hashes remain unchanged.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants