Summary
When a Virtual Model in GoModel is configured to round-robin across upstreams that expose different API protocols (OpenAI Chat Completions vs. OpenAI Responses API), a chat session cannot be continued once the upstream swap crosses a protocol boundary.
- ✅ Requests that land on the same protocol type work fine.
- ❌ As soon as round-robin switches from one protocol to another, continuing the conversation fails — the previous conversation state is not portable between the two APIs.
Use Case / Motivation
Virtual Models are a great way to combine capacity from multiple providers behind a single model name. A typical setup:
| Upstream |
Protocol |
Example model |
kimicode |
Chat Completion |
k3 |
chatgpt |
Responses API |
gpt-5.6-terra |
zai |
Chat Completion |
glm-5.3-flash |
Scenario: A client (e.g. Kimi Code CLI) sends an OpenAI Chat Completion request to the Virtual Model. GoModel round-robins between the three upstreams above. Requests hitting kimi-code or zai work as expected — but as soon as the router lands on chatgpt (Responses API), the conversation can no longer be continued.
Observed behavior
- Send request N → routed to Chat Completion upstream → response OK.
- Send request N+1 → routed to Responses API upstream → failure / broken context.
- From then on, the session is effectively unusable for that virtual model.
Root cause (assumption): Chat Completions and the Responses API represent conversation state differently (messages array vs. response/previous_response_id chain). When GoModel swaps upstreams mid-conversation, state from one protocol cannot be replayed into the other, and the client has no way to know a protocol switch occurred.
Expected behavior (options)
- 🅰️ GoModel translates conversation state between protocols on upstream swap (messages ↔ response chain), so mixed-protocol round-robin works transparently.
- 🅱️ GoModel detects protocol incompatibility and pins a conversation (by session/thread identifier) to a protocol-compatible upstream subset.
- 🅲️ Document this as a known limitation and allow Virtual Models to be restricted to a single protocol type across all upstreams (validation/warning at config time).
Test plan
Detailed test matrix will follow later today — mapping client type × client protocol × upstream protocol and whether the conversation survives upstream swaps. Planned coverage:
- Kimi Code CLI → Virtual Model (mixed round-robin)
- Raw OpenAI Chat Completions client → Virtual Model
- Responses API client → Virtual Model
- Behavior per load-balancing strategy (
round_robin, least_latency, …)
Environment
- GoModel version: (to be added)
- Deployment: (to be added)
- Config snippet (virtual model + upstreams): (to be added)
Root cause (confirmed by code walk on main @ 2c0f662)
The router is protocol-blind: viableTargets (internal/virtualmodels/chain.go:49) only checks catalog availability, never the upstream's API protocol. Session affinity is default-on (internal/virtualmodels/types.go:65-67), but only engages if a session id is detected (internal/virtualmodels/balancer.go:119-121) — Kimi Code CLI sends none of the built-in session headers (internal/session/session.go:47), and content auto-detection (internal/session/detect.go:59,128, default on per config/config.go:212) needs a request body ≤ 1 MiB (internal/auditlog/constants.go:7). No detected id → per-request rotation crosses the protocol boundary.
The boundary fails in both directions:
- Chat client → chatgpt: no Chat→Responses request translator exists anywhere in the codebase.
chatgpt.ChatCompletion/StreamChatCompletion return 501 unsupported_provider_operation (internal/providers/chatgpt/chatgpt.go:159-166).
- Responses client with
previous_response_id → chat-translated upstream: hard 400 in the shared adapter (internal/providers/responses_adapter.go:91-93). The local response store is never consulted on the request path to expand previous_response_id into replayable input.
- Failover worsens it: default retry statuses
429,5xx (internal/gateway/failover_policy.go:11) — the 501 above is a 5xx, so ShouldRetry sweeps the failover chain into the incompatible target.
The only portable state model is the gateway-managed conversation field (internal/server/conversation_responses.go:15-21), which can't combine with previous_response_id (:53). Full write-up and option analysis in the investigation comment below.
Suggested direction (from investigation)
Option B (pin conversations to a protocol-compatible subset) is the most contained fix in this codebase: the pinning machinery (stickySessions) and provider-protocol info (catalog.GetProviderType) already exist; the gap is precisely that viableTargets/balancedResolution ignore protocol — filtering the viable pool by client endpoint protocol before strategy selection is a contained change at internal/virtualmodels/chain.go:49 / internal/virtualmodels/balancer.go:56. Option A (state translation on swap) is feasible but requires a brand-new Chat→Responses request translator plus wiring the response store into the request path — much larger surface. Option C is cheap and worth shipping alongside B as a config-time warning.
Wayfinder map (the wayfinding view of this issue)
Destination
Any client protocol (OpenAI Chat Completions, OpenAI Responses, Anthropic Messages) can reach any upstream protocol through GoModel, so virtual-model round-robin and failover work across protocol boundaries. This issue's symptom — a conversation breaking when the router crosses the Chat Completions / Responses boundary — disappears because the protocol matrix is complete and translation is loss-bounded and explicit.
Notes
- Domain: protocol translation providers (
internal/providers), the dispatch path (internal/gateway/inference_execute.go), and the virtual-model balancer (internal/virtualmodels).
- Every session consults the existing translator inventory in
internal/providers/responses_adapter.go and the streaming converters.
- Architecture decisions are settled; open decisions are ticketed below as child issues.
- Work only in this repo (
weselben/GoModel); no PRs to any other upstream.
Decisions so far
- Investigation: confirmed root cause (protocol-blind router + missing chat→responses translation + previous_response_id rejection) and the failure shapes per direction — see the investigation comment for the full trace and file:line citations.
- Architecture: extend the existing provider-adapts pattern rather than add a central gateway gate. Providers already translate their own dialect; the chat↔responses pair is the only incomplete pair. No new dispatch layer.
- Topology: chat completions is the hub (star). Not a canonical internal message representation.
- Slice shape: translators first (pure additions), then chatgpt wiring (replaces its 501s), then
previous_response_id expansion from the local response store (orthogonal, separate decision).
- Verify chatgpt stream-parser reuse and lock translator signatures —
collapseResponsesStream not reusable for streaming (non-streaming collapse); five new functions locked with exact signatures (ConvertChatRequestToResponses, ConvertResponsesResponseToChat, ChatResponsesStreamConverter, ChatViaResponses, StreamChatViaResponses); responses→chat needs index-tracked call_id patching (id arrives late on response.output_item.done).
- Decide follow-up scope: protocol-aware pool filtering and config-time warning — defer same-protocol filtering (pure optimization after the matrix); no config-time mixed-protocol warning, warn only on missing directed edges.
- Lock the chat↔responses translation-loss policy — the ADR-0011 field-forwarding ladder already locks it; initial per-field rows for the chat↔responses pair recorded in the ticket.
- kimicode speaks the Responses API natively (draft PR ENTERPILOT/GoModel#916); live probe confirms kimi's native
/responses is stateless (store:true and previous_response_id both 400), so gateway-managed state is the only chaining path even against responses-speaking upstreams.
Not yet specified
- Documentation updates (
docs/features/virtual-models.mdx, docs/features/failover.mdx) once the loss policy is locked.
Out of scope
- A canonical internal message model / full protocol IR redesign.
- Migrating the ~10 provider-local
ResponsesViaChat wrappers onto shared translators (cleanup only; behavior-preserving, deferred).
- Decide scope: session-detection gap for bodies over 1 MiB — affinity-only concern once the matrix is complete; separate effort.
Summary
When a Virtual Model in GoModel is configured to round-robin across upstreams that expose different API protocols (OpenAI Chat Completions vs. OpenAI Responses API), a chat session cannot be continued once the upstream swap crosses a protocol boundary.
Use Case / Motivation
Virtual Models are a great way to combine capacity from multiple providers behind a single model name. A typical setup:
kimicodek3chatgptgpt-5.6-terrazaiglm-5.3-flashScenario: A client (e.g. Kimi Code CLI) sends an OpenAI Chat Completion request to the Virtual Model. GoModel round-robins between the three upstreams above. Requests hitting
kimi-codeorzaiwork as expected — but as soon as the router lands onchatgpt(Responses API), the conversation can no longer be continued.Observed behavior
Root cause (assumption): Chat Completions and the Responses API represent conversation state differently (messages array vs. response/
previous_response_idchain). When GoModel swaps upstreams mid-conversation, state from one protocol cannot be replayed into the other, and the client has no way to know a protocol switch occurred.Expected behavior (options)
Test plan
Detailed test matrix will follow later today — mapping client type × client protocol × upstream protocol and whether the conversation survives upstream swaps. Planned coverage:
round_robin,least_latency, …)Environment
Root cause (confirmed by code walk on
main@ 2c0f662)The router is protocol-blind:
viableTargets(internal/virtualmodels/chain.go:49) only checks catalog availability, never the upstream's API protocol. Session affinity is default-on (internal/virtualmodels/types.go:65-67), but only engages if a session id is detected (internal/virtualmodels/balancer.go:119-121) — Kimi Code CLI sends none of the built-in session headers (internal/session/session.go:47), and content auto-detection (internal/session/detect.go:59,128, default on perconfig/config.go:212) needs a request body ≤ 1 MiB (internal/auditlog/constants.go:7). No detected id → per-request rotation crosses the protocol boundary.The boundary fails in both directions:
chatgpt.ChatCompletion/StreamChatCompletionreturn 501unsupported_provider_operation(internal/providers/chatgpt/chatgpt.go:159-166).previous_response_id→ chat-translated upstream: hard 400 in the shared adapter (internal/providers/responses_adapter.go:91-93). The local response store is never consulted on the request path to expandprevious_response_idinto replayable input.429,5xx(internal/gateway/failover_policy.go:11) — the 501 above is a 5xx, soShouldRetrysweeps the failover chain into the incompatible target.The only portable state model is the gateway-managed
conversationfield (internal/server/conversation_responses.go:15-21), which can't combine withprevious_response_id(:53). Full write-up and option analysis in the investigation comment below.Suggested direction (from investigation)
Option B (pin conversations to a protocol-compatible subset) is the most contained fix in this codebase: the pinning machinery (
stickySessions) and provider-protocol info (catalog.GetProviderType) already exist; the gap is precisely thatviableTargets/balancedResolutionignore protocol — filtering the viable pool by client endpoint protocol before strategy selection is a contained change atinternal/virtualmodels/chain.go:49/internal/virtualmodels/balancer.go:56. Option A (state translation on swap) is feasible but requires a brand-new Chat→Responses request translator plus wiring the response store into the request path — much larger surface. Option C is cheap and worth shipping alongside B as a config-time warning.Wayfinder map (the wayfinding view of this issue)
Destination
Any client protocol (OpenAI Chat Completions, OpenAI Responses, Anthropic Messages) can reach any upstream protocol through GoModel, so virtual-model round-robin and failover work across protocol boundaries. This issue's symptom — a conversation breaking when the router crosses the Chat Completions / Responses boundary — disappears because the protocol matrix is complete and translation is loss-bounded and explicit.
Notes
internal/providers), the dispatch path (internal/gateway/inference_execute.go), and the virtual-model balancer (internal/virtualmodels).internal/providers/responses_adapter.goand the streaming converters.weselben/GoModel); no PRs to any other upstream.Decisions so far
previous_response_idexpansion from the local response store (orthogonal, separate decision).collapseResponsesStreamnot reusable for streaming (non-streaming collapse); five new functions locked with exact signatures (ConvertChatRequestToResponses,ConvertResponsesResponseToChat,ChatResponsesStreamConverter,ChatViaResponses,StreamChatViaResponses); responses→chat needs index-trackedcall_idpatching (id arrives late onresponse.output_item.done)./responsesis stateless (store:trueandprevious_response_idboth 400), so gateway-managed state is the only chaining path even against responses-speaking upstreams.Not yet specified
docs/features/virtual-models.mdx,docs/features/failover.mdx) once the loss policy is locked.Out of scope
ResponsesViaChatwrappers onto shared translators (cleanup only; behavior-preserving, deferred).