Skip to content

Support Claude Opus 5.5 in the Anthropic backend - #20

Open
MrJoy wants to merge 4 commits into
wandercom:mainfrom
MrJoy:feat/opus-5-5-tool-choice
Open

MrJoy wants to merge 4 commits into
wandercom:mainfrom
MrJoy:feat/opus-5-5-tool-choice

Conversation

@MrJoy

@MrJoy MrJoy commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor

Opus 5.5 rejects forced tool use (tool_choice: type "tool" and "any" are not supported for this model.), and AnthropicBackend forced it on every call. Every role on Opus 5.5 failed on its first request.

The backend now starts forced. On that specific 400 it flips the instance to tool_choice: auto (with disable_parallel_tool_use), adds one system-prompt line naming the tool, and retries. set_model() resets it. Opus 4.8 and older never see a difference. Under auto the model can answer in text, so a response without a tool call now gets a correction and a retry through the existing 3-attempt loop instead of an immediate RuntimeError.

Reading guide: _stream_tool_call in src/pact/backends/anthropic.py is the whole mechanism. Both stream paths now go through it.

Why not strict: true or structured outputs: both require additionalProperties: false on every object. ComponentContract, ContractTestSuite, InterviewResult and ShapingPitch all have dict[str, X] fields, which that can't express. Going strict means reworking those schemas first. Pydantic validation plus the correction loop already covers schema validity.

Also in here:

  • claude-opus-5-5 priced at $4/$20. The fuzzy prefix match in pricing_for_model was resolving it to claude-opus-4 at $15/$75. claude-opus-4-8 still hits the same wrong match; I left that alone.
  • claude-opus-5-5 max-tokens cap raised to 128K. Thinking can't be turned off on 5.5 and it counts against max_tokens.
  • stop_reason: "refusal" raises right away instead of burning two retries.

Verified live against the API with the new code, one run per model:

INFO pact.backends.anthropic: claude-opus-5-5 rejected forced tool_choice; using tool_choice auto
claude-opus-5-5 register ok: _RegisterResponse forced= False 2s
claude-opus-5-5 interview ok: risks 9 questions 7 21s
claude-opus-5-5 contract ok: duration_parser 2 functions 34s
claude-opus-4-8 register ok: _RegisterResponse forced= True 2s
claude-opus-4-8 interview ok: risks 6 questions 7 37s
claude-opus-4-8 contract ok: duration_parser 2 functions 23s

No missed-tool-call retries fired on 5.5 in that run. One run each, so that's a sample and not a reliability number.

Test suite: 27 failures, and main fails the same 27 in the same venv (missing optional deps: openai, tree-sitter, Slack/Linear). The new code adds no failures.

🤖 Generated with Claude Code

https://claude.ai/code/session_011FeZUmAZbyQWDJNAPkUA3S


View with [code]smith Autofix with [code]smith
Need help on this PR? Tag @codesmith-bot with what you need. Autofix is disabled.

Claude Opus 5.5 (and Sonnet 5.5 / Fable 5.1) 400 on `tool_choice` of type `tool` or `any`, so every Pact role on Opus 5.5 died on its first request.  Also price Opus 5.5 at $4/$20; the fuzzy prefix match was billing it at the Opus 4 rate.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011FeZUmAZbyQWDJNAPkUA3S
@MrJoy

MrJoy commented Sep 28, 2026

Copy link
Copy Markdown
Contributor Author

Does this maintain backwards compatibility with Opus 5 / 4.8?

@MrJoy

MrJoy commented Sep 28, 2026

Copy link
Copy Markdown
Contributor Author

Yes. The request those models get is byte-for-byte the same as before: forced tool_choice, same system prompt. The switch to auto only happens on the specific 400 that 5.5 returns, and neither Opus 5 nor 4.8 ever returns it.

Checked live against both with this branch:

claude-opus-5 register ok: _RegisterResponse forced= True 3s
claude-opus-5 interview ok: risks 10 questions 10 48s
claude-opus-5 contract ok: duration_parser 2 functions 49s
claude-opus-4-8 register ok: _RegisterResponse forced= True 2s
claude-opus-4-8 interview ok: risks 6 questions 7 37s
claude-opus-4-8 contract ok: duration_parser 2 functions 23s

forced= True after the first call means the switch never happened.

Three retry-loop changes do apply to every model, but none of them can fire on a normal forced-mode response:

  1. No tool call in the response: retry with a correction instead of an immediate RuntimeError. Under forced tool_choice that only happens when output got cut off.
  2. max_tokens with no tool call: this used to raise, and now it bumps max_tokens and retries. Before, the bump only happened when a truncated tool call was present.
  3. stop_reason: "refusal" raises right away instead of retrying.

N.B. Opus 5 and 4.8 are both still mispriced at the Opus 4 rate ($15/$75) through the fuzzy match in pricing_for_model. That was already broken before this PR and isn't fixed here. Can add the entries if you want them in this one.

The fuzzy prefix match in `pricing_for_model` was resolving both to `claude-opus-4` at $15/$75.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011FeZUmAZbyQWDJNAPkUA3S
@MrJoy

MrJoy commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor Author

Added in 1d2b7fa: claude-opus-5 and claude-opus-4-8 at $5/$25, per the current Anthropic pricing table. claude-opus-4-7 goes through the same fuzzy match and still lands at $15/$75. Left it out since nobody asked, but it's a one-liner if you want it.

Checked every entry against the vendor pricing pages as of 2026-09-28.  `gemini-3-pro-preview` was shut down 2026-03-09; `gemini-3.1-pro-preview` replaces it.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011FeZUmAZbyQWDJNAPkUA3S
@MrJoy

MrJoy commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor Author

9b28e00 goes through the whole DEFAULT_MODEL_PRICING table. Every entry was checked against the vendor's own pricing page today.

Wrong before, fixed now:

Model Was Now
o3 $10 / $40 $2 / $8
gemini-2.5-flash $0.15 / $0.60 $0.30 / $2.50
gemini-2.5-flash-lite $0.075 / $0.30 $0.10 / $0.40
gemini-3-flash-preview $0.15 / $0.60 $0.50 / $3.00

Missing before, and the fuzzy match was resolving them to the wrong rate:

Model Resolved to Now
claude-opus-4-7 $15 / $75 (Opus 4) $5 / $25
claude-sonnet-5, claude-sonnet-5-5 $3 / $15 (Sonnet 4) $2 / $10
claude-fable-5, claude-fable-5-1, claude-mythos-5, claude-mythos-5-1 $1 / $5 (Haiku fallback) $10 / $50

gemini-3-pro-preview is gone. Google shut it down on 2026-03-09 and gemini-3.1-pro-preview ($2 / $12) replaces it. The old entry was wrong anyway, at $1.25 / $10.

claude-haiku-3 is retired and off Anthropic's pricing page. $0.25 / $1.25 was its last published price, so I left it in.

test_gemini_backend.py asserted the old 2.5 Flash price. Nobody noticed because it skips when google-genai isn't installed. Fixed it, and ran it with .[all-backends] installed.

@MrJoy
MrJoy marked this pull request as ready for review September 28, 2026 19:03

@adaptcom adaptcom Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Confidence Score: 2/5

Summary

Adds automatic tool-choice fallback and refreshes model pricing. A reproduced concurrency bug can abort parallel adoption, so the fallback needs correction before merge.

Important Files Changed

File Overview
src/pact/backends/anthropic.py Adds automatic tool-choice fallback, correction retries, refusal handling, and a higher token cap.
src/pact/budget.py Updates model pricing and adds newer model entries.
tests/test_anthropic_tool_choice.py Exercises sequential fallback, model reset, caching, and missing-tool retries.
tests/test_budget.py Adds exact-match pricing coverage.
tests/test_gemini_backend.py Updates Flash pricing expectations.
tests/test_validation_retry.py Expects three attempts when tool calls are absent.

↻ Re-run review · View in Adapt

Comment thread src/pact/backends/anthropic.py Outdated
Two forced requests in flight at once would both get the 400.  The first flipped the instance flag to auto, and then the second saw auto in its guard and re-raised the 400 instead of retrying, which aborted parallel runs.  Each request now tracks whether it was sent forced, and falls back on its own 400 whatever the instance flag says.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MXo6tGa9GdEhd8bZuX6cuo

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Development

Successfully merging this pull request may close these issues.

1 participant