Conversation
Claude Opus 5.5 (and Sonnet 5.5 / Fable 5.1) 400 on `tool_choice` of type `tool` or `any`, so every Pact role on Opus 5.5 died on its first request. Also price Opus 5.5 at $4/$20; the fuzzy prefix match was billing it at the Opus 4 rate. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_011FeZUmAZbyQWDJNAPkUA3S
|
Does this maintain backwards compatibility with Opus 5 / 4.8? |
|
Yes. The request those models get is byte-for-byte the same as before: forced Checked live against both with this branch:
Three retry-loop changes do apply to every model, but none of them can fire on a normal forced-mode response:
N.B. Opus 5 and 4.8 are both still mispriced at the Opus 4 rate ($15/$75) through the fuzzy match in |
The fuzzy prefix match in `pricing_for_model` was resolving both to `claude-opus-4` at $15/$75. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_011FeZUmAZbyQWDJNAPkUA3S
|
Added in |
Checked every entry against the vendor pricing pages as of 2026-09-28. `gemini-3-pro-preview` was shut down 2026-03-09; `gemini-3.1-pro-preview` replaces it. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_011FeZUmAZbyQWDJNAPkUA3S
|
Wrong before, fixed now:
Missing before, and the fuzzy match was resolving them to the wrong rate:
|
There was a problem hiding this comment.
Confidence Score: 2/5
Summary
Adds automatic tool-choice fallback and refreshes model pricing. A reproduced concurrency bug can abort parallel adoption, so the fallback needs correction before merge.
Important Files Changed
| File | Overview |
|---|---|
| src/pact/backends/anthropic.py | Adds automatic tool-choice fallback, correction retries, refusal handling, and a higher token cap. |
| src/pact/budget.py | Updates model pricing and adds newer model entries. |
| tests/test_anthropic_tool_choice.py | Exercises sequential fallback, model reset, caching, and missing-tool retries. |
| tests/test_budget.py | Adds exact-match pricing coverage. |
| tests/test_gemini_backend.py | Updates Flash pricing expectations. |
| tests/test_validation_retry.py | Expects three attempts when tool calls are absent. |
Two forced requests in flight at once would both get the 400. The first flipped the instance flag to auto, and then the second saw auto in its guard and re-raised the 400 instead of retrying, which aborted parallel runs. Each request now tracks whether it was sent forced, and falls back on its own 400 whatever the instance flag says. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01MXo6tGa9GdEhd8bZuX6cuo
Opus 5.5 rejects forced tool use (
tool_choice: type "tool" and "any" are not supported for this model.), andAnthropicBackendforced it on every call. Every role on Opus 5.5 failed on its first request.The backend now starts forced. On that specific 400 it flips the instance to
tool_choice: auto(withdisable_parallel_tool_use), adds one system-prompt line naming the tool, and retries.set_model()resets it. Opus 4.8 and older never see a difference. Underautothe model can answer in text, so a response without a tool call now gets a correction and a retry through the existing 3-attempt loop instead of an immediateRuntimeError.Reading guide:
_stream_tool_callinsrc/pact/backends/anthropic.pyis the whole mechanism. Both stream paths now go through it.Why not
strict: trueor structured outputs: both requireadditionalProperties: falseon every object.ComponentContract,ContractTestSuite,InterviewResultandShapingPitchall havedict[str, X]fields, which that can't express. Going strict means reworking those schemas first. Pydantic validation plus the correction loop already covers schema validity.Also in here:
claude-opus-5-5priced at $4/$20. The fuzzy prefix match inpricing_for_modelwas resolving it toclaude-opus-4at $15/$75.claude-opus-4-8still hits the same wrong match; I left that alone.claude-opus-5-5max-tokens cap raised to 128K. Thinking can't be turned off on 5.5 and it counts againstmax_tokens.stop_reason: "refusal"raises right away instead of burning two retries.Verified live against the API with the new code, one run per model:
No missed-tool-call retries fired on 5.5 in that run. One run each, so that's a sample and not a reliability number.
Test suite: 27 failures, and
mainfails the same 27 in the same venv (missing optional deps: openai, tree-sitter, Slack/Linear). The new code adds no failures.🤖 Generated with Claude Code
https://claude.ai/code/session_011FeZUmAZbyQWDJNAPkUA3S
Need help on this PR? Tag
@codesmith-botwith what you need. Autofix is disabled.