Skip to content

fix: remove the default defect-group cap - #87

Merged
svozza merged 4 commits into
mainfrom
fix/no-defect-group-cap
Oct 1, 2026
Merged

svozza merged 4 commits into
mainfrom
fix/no-defect-group-cap

Conversation

@svozza

@svozza svozza commented Oct 1, 2026 •

Copy link
Copy Markdown
Owner

A review can have fewer than ten finding entries and still be rejected by the shipped four-defect-group cap. On Cal.com, that cap pushed confirmed concerns into residual risk and repeatedly missed the Office365 destination defect.

Ship review.max_distinct_groups: null and state the effective group policy in the generated constraints. Keep the ten-entry maximum, required group field, numeric group range and all provenance/output checks. Numeric consumer caps remain supported and are disclosed in the prompt. Runtime files match the evaluated no-cap candidate exactly; the additional changes update the default-policy tests and architecture/ADR documentation. An evaluation-only correction also addresses the grouping fixture’s documented rejection of an optional root-cause finding; its exact path/line/diagnosis matcher preserves required callers and rejects unrelated or split groups. No reviewer input changes accompany that correction.

The completed direct comparison against shipped 8b91cb5 used three paired repetitions across six frozen source snapshots (36 reviews). Known target delivery improved from 18/24 to 21/24, recovering Office365 in all three candidate runs versus none in control. Summed review time fell 5.57% and submission rejections fell 5 to 0. Both arms still had one false structured finding and incomplete repair advice; repeated nonexhaustive targets do not establish general accuracy. Raw source assessments remain private/local.

Validation:

  • Full deterministic suite: 2,400 passed, 2 fixture-dependent skips.
  • Type checking and diff checks passed.
  • Boundary tests accept ten distinct groups, reject eleven entries, and preserve configured-cap enforcement.
  • Initial CI: two raw passes, then a false grouping failure documented by the fixture itself. All three runs remain retained. Uniform regrading corrects the false failure and exposes an incorrectly split root in the second run.
  • Final-head CI is green: deterministic tests/type checking and three fresh full evaluation passes, totaling 126/126 review scenarios and 18/18 plan scenarios, with zero API errors. Earlier runs remain diagnostic evidence and do not count toward this gate.

Final-head evaluation runs: 1, 2, 3. Deterministic CI: Quality check.

@svozza
svozza deployed to ai-pr-review-runtime October 1, 2026 15:23 — with GitHub Actions Active
@svozza svozza added the run-evals Run the eval suite against a live model on this PR label Oct 1, 2026
@svozza
svozza deployed to ai-pr-review-runtime October 1, 2026 15:29 — with GitHub Actions Active
@svozza svozza added run-evals Run the eval suite against a live model on this PR and removed run-evals Run the eval suite against a live model on this PR labels Oct 1, 2026
@svozza
svozza had a problem deploying to ai-pr-review-runtime October 1, 2026 22:54 — with GitHub Actions Failure
@svozza
svozza deployed to ai-pr-review-runtime October 1, 2026 23:03 — with GitHub Actions Active
@svozza svozza added run-evals Run the eval suite against a live model on this PR and removed run-evals Run the eval suite against a live model on this PR labels Oct 1, 2026
@svozza
svozza deployed to ai-pr-review-runtime October 1, 2026 23:11 — with GitHub Actions Active
@svozza svozza added run-evals Run the eval suite against a live model on this PR and removed run-evals Run the eval suite against a live model on this PR labels Oct 1, 2026
@svozza
svozza deployed to ai-pr-review-runtime October 1, 2026 23:18 — with GitHub Actions Active
@svozza
svozza merged commit 5391472 into main Oct 1, 2026
14 checks passed

This branch was successfully deployed

1 active deployment
ai-pr-review-runtime — 153e9c9d Deployed Oct 1, 2026 by svozza via evals #197
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

run-evals Run the eval suite against a live model on this PR

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant