Repository navigation
Conversation
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Changes
Make explicit object-store and signed HTTP
.evalsources usable by the investigation reader and child-audit staging. Cache direct inputs atomically, validate by parsing, namespace repeated sample IDs by source/item/epoch, and sanitize credentials in download errors.Preserve the direct OpenRouter routing policy while refreshing Hawk-native credentials for control-plane operations when the generic provider hook is disabled. Bind the qualified
openrouter/openai/...catalog entry through Inspect's builtin model configuration; this fixes child submission without changing the allowed model/route.Document
read_eval_log_samples(..., all_samples_required=False)for successful imported archives with unequal per-item trial counts. Such archives are sparse by design; strict rectangular streaming is inappropriate.Verification
Draft against
dev. Existing development checkout and running jobs remain isolated from this worktree. No keys, presigned URLs or evaluation logs are included.Per-attempt replay fix
DTBench's native logs exposed a second replay problem:
score_asyncinherited the benchmark's custom epoch reducer, and overwrite mode also inherited its metric definitions. The auditor's categorical concordance scorer and numeric collector are different scorers and must use their own aggregates. The replay now usesmode_score()for concordance andmean()/mean_score()for the collector, while retaining the actual benchmark scorer and individual grades.Verified on Hawk's exact Inspect runtime: 29 grader/concordance tests pass. A regression with unavailable custom metric/reducer names and three epochs fails without the overrides and passes with them. Native Nano replay reproduces all nine stored grades across three selected items, including invalid responses; authored correct/wrong/malformed outputs yield 1/0/NaN. Original source log hashes remain unchanged.