Repository navigation
fix(evaluation): score CAR scenarios on the mode key only - #627
Merged
Merged
Conversation
A CAR answer now passes when it is a single-key object whose key matches the gold mode (response / clarification / abstain). The value under the key no longer affects pass or score; required-term coverage stays in the mode_* fields as a diagnostic. Replaces the CAR tests that targeted the removed groundtruth_eval.json term scoring and updates docs/static-json-evaluation.md to match. Signed-off-by: Shuxin Lin <linshuhsin@gmail.com>
DhavalRepo18
approved these changes
Oct 8, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Description
Score clarification-abstain-response (CAR) scenarios only on the top-level mode key of the answer (
response,clarificationorabstain). The text under the key no longer affects pass or score.Fix Details
src/evaluation/scorers/static_json.py,_evaluate_mode_json: an answer passes when it is a JSON object with exactly one top-level key and that key matches the gold mode (case-insensitive).score,f1andstrict_exact_match_accuracyare 1.0 on a pass and 0.0 otherwise.mode_required_terms,mode_matched_terms,mode_term_coverage) is still reported as a diagnostic, but no longer feeds pass, score or the precision/recall/F1 key counts. The per-keydetailsnow hold only the mode-key comparison.tests/test.sh, so Harbor rewards change too.aafeedback_changes. They targeted thegroundtruth_eval.jsonterm scoring andcar_scorethat were removed in 8f1e538. I replaced them with key-only tests: a correct key with no matching terms passes, a wrong key with matching terms fails, the key is case-insensitive, an answer that isn't an object fails, and theStaticJsonScorerwrapper gives 1.0/0.0.docs/static-json-evaluation.md: the CAR pass criteria and fields now match the code, and the docs saystatic_jsondoes not readgroundtruth_eval.json.Impact on Benchmarking
CAR pass rates will go up for answers that pick the right mode but word it differently from the gold answer. Trade-off: an agent that always picks the same mode gets credit on every scenario expecting that mode. This PR has no before/after agent-run comparison.
Conflicts with
origin/eval-scoring-relaxation(d1a668c), which rewrites the same function so that every required term is needed to pass.Related Issues
Verification Steps
uv run pytest src/evaluation -q -k "not integration": 110 passed (6 failed before this change).uv run pytest src/ -q -k "not integration": 752 passed, 2 failed insrc/observability/tests/test_file_exporter.py. Those 2 also fail on the base branch without this change.benchmarks/harbor/datasets/passes when scored against itself.Checklist