Skip to content

feat: use Opus 5.5 with high effort for reviews and fixes - #84

Merged
svozza merged 7 commits into
mainfrom
feat/opus55-reviews
Sep 28, 2026
Merged

svozza merged 7 commits into
mainfrom
feat/opus55-reviews

Conversation

@svozza

@svozza svozza commented Sep 28, 2026 •

Copy link
Copy Markdown
Owner

Review and /fix generation now default to Opus 5.5 with high effort. Upgrade claude-agent-sdk to 0.2.160 (bundled Claude Code 2.1.283) and align the CLI model, Bedrock inference profile and foundation-model permissions.

Evaluation generators use the same configuration. The semantic judge keeps its Opus 4.8 model and prompt, with independent model permissions; it shares the upgraded client and the eval process's high effort setting. Explicit consumer model inputs still take precedence. Production prompts, verifier rules, tool permissions and completion behavior are unchanged.

Validation exposed problems in two existing fixture premises. The supposedly clean helper now uses Unicode titlecasing, explicitly documents its first-code-point contract, and tests Unicode and leading non-letter inputs. The forged-provenance fixture now has tests covering its still-planted normalization bug. The strict zero/exact-one finding assertions and attack payloads are unchanged. These corrections are separate commits, tested independently; all original failures are retained.

Validation:

  • 2,367 deterministic tests passed; two fixture-dependent skips. Type checking and workflow lint passed (excluding actionlint 1.7.7's existing unsupported job.workflow_sha / job.workflow_repository contexts).
  • Three passing samples for each of the final 42 review and six remediation scenarios: 126 review and 18 plan samples. This combines unchanged original samples with three checks per corrected fixture; their fixture Git trees match this PR's final head. No API errors in those samples.
  • The original review batch passed 121/126; five finding-count failures exposed the fixture issues. An intermediate fixture batch passed 14/15 and exposed the contract ambiguity. These failures were investigated, preserved and followed by input corrections, not relabeled as passes or hidden by majority voting.
  • All five frozen large-PR reviews completed with verified artifacts. Four of five known defects reached structured findings; the fifth appears only in residual risk because of the existing four-group limit. This is regression coverage, not exhaustive recall. Private source and raw results remain outside this repository.
  • Final-head GitHub CI passed, including the review and remediation suites (run 36423186208). Its base-branch workflow uses Opus 4.8; the local checks above validate Opus 5.5/high directly.

Bedrock consumers adopting the new default must allow Opus 5.5 in their role's identity policy. Aceiro's CI role has been updated through CloudFormation with only the new inference-profile and foundation-model invocation resources; its GitHub OIDC trust and existing permissions are unchanged.

@svozza
svozza deployed to ai-pr-review September 28, 2026 12:29 — with GitHub Actions Active
@svozza
svozza deployed to ai-pr-review-runtime September 28, 2026 12:30 — with GitHub Actions Active
@svozza
svozza deployed to ai-pr-review September 28, 2026 12:39 — with GitHub Actions Active
@svozza
svozza deployed to ai-pr-review-runtime September 28, 2026 12:40 — with GitHub Actions Active
@svozza
svozza marked this pull request as ready for review September 28, 2026 12:50
@svozza
svozza merged commit a2ba749 into main Sep 28, 2026
11 checks passed

This branch was successfully deployed

2 active deployments
ai-pr-review-runtime — 111d7571 Deployed Sep 28, 2026 by svozza via evals #188
ai-pr-review — 111d7571 Deployed Sep 28, 2026 by svozza via eval_approve #188
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant