Skip to content

API - Translation issues / Round-Robin between models from different Upstream API Styles #84

Description

@weselben

Summary

When a Virtual Model in GoModel is configured to round-robin across upstreams that expose different API protocols (OpenAI Chat Completions vs. OpenAI Responses API), a chat session cannot be continued once the upstream swap crosses a protocol boundary.

  • ✅ Requests that land on the same protocol type work fine.
  • ❌ As soon as round-robin switches from one protocol to another, continuing the conversation fails — the previous conversation state is not portable between the two APIs.

Use Case / Motivation

Virtual Models are a great way to combine capacity from multiple providers behind a single model name. A typical setup:

Upstream Protocol Example model
kimicode Chat Completion k3
chatgpt Responses API gpt-5.6-terra
zai Chat Completion glm-5.3-flash

Scenario: A client (e.g. Kimi Code CLI) sends an OpenAI Chat Completion request to the Virtual Model. GoModel round-robins between the three upstreams above. Requests hitting kimi-code or zai work as expected — but as soon as the router lands on chatgpt (Responses API), the conversation can no longer be continued.

Observed behavior

  1. Send request N → routed to Chat Completion upstream → response OK.
  2. Send request N+1 → routed to Responses API upstream → failure / broken context.
  3. From then on, the session is effectively unusable for that virtual model.

Root cause (assumption): Chat Completions and the Responses API represent conversation state differently (messages array vs. response/previous_response_id chain). When GoModel swaps upstreams mid-conversation, state from one protocol cannot be replayed into the other, and the client has no way to know a protocol switch occurred.

Expected behavior (options)

  • 🅰️ GoModel translates conversation state between protocols on upstream swap (messages ↔ response chain), so mixed-protocol round-robin works transparently.
  • 🅱️ GoModel detects protocol incompatibility and pins a conversation (by session/thread identifier) to a protocol-compatible upstream subset.
  • 🅲️ Document this as a known limitation and allow Virtual Models to be restricted to a single protocol type across all upstreams (validation/warning at config time).

Test plan

Detailed test matrix will follow later today — mapping client type × client protocol × upstream protocol and whether the conversation survives upstream swaps. Planned coverage:

  • Kimi Code CLI → Virtual Model (mixed round-robin)
  • Raw OpenAI Chat Completions client → Virtual Model
  • Responses API client → Virtual Model
  • Behavior per load-balancing strategy (round_robin, least_latency, …)

Environment

  • GoModel version: (to be added)
  • Deployment: (to be added)
  • Config snippet (virtual model + upstreams): (to be added)

Root cause (confirmed by code walk on main @ 2c0f662)

The router is protocol-blind: viableTargets (internal/virtualmodels/chain.go:49) only checks catalog availability, never the upstream's API protocol. Session affinity is default-on (internal/virtualmodels/types.go:65-67), but only engages if a session id is detected (internal/virtualmodels/balancer.go:119-121) — Kimi Code CLI sends none of the built-in session headers (internal/session/session.go:47), and content auto-detection (internal/session/detect.go:59,128, default on per config/config.go:212) needs a request body ≤ 1 MiB (internal/auditlog/constants.go:7). No detected id → per-request rotation crosses the protocol boundary.

The boundary fails in both directions:

  • Chat client → chatgpt: no Chat→Responses request translator exists anywhere in the codebase. chatgpt.ChatCompletion/StreamChatCompletion return 501 unsupported_provider_operation (internal/providers/chatgpt/chatgpt.go:159-166).
  • Responses client with previous_response_id → chat-translated upstream: hard 400 in the shared adapter (internal/providers/responses_adapter.go:91-93). The local response store is never consulted on the request path to expand previous_response_id into replayable input.
  • Failover worsens it: default retry statuses 429,5xx (internal/gateway/failover_policy.go:11) — the 501 above is a 5xx, so ShouldRetry sweeps the failover chain into the incompatible target.

The only portable state model is the gateway-managed conversation field (internal/server/conversation_responses.go:15-21), which can't combine with previous_response_id (:53). Full write-up and option analysis in the investigation comment below.

Suggested direction (from investigation)

Option B (pin conversations to a protocol-compatible subset) is the most contained fix in this codebase: the pinning machinery (stickySessions) and provider-protocol info (catalog.GetProviderType) already exist; the gap is precisely that viableTargets/balancedResolution ignore protocol — filtering the viable pool by client endpoint protocol before strategy selection is a contained change at internal/virtualmodels/chain.go:49 / internal/virtualmodels/balancer.go:56. Option A (state translation on swap) is feasible but requires a brand-new Chat→Responses request translator plus wiring the response store into the request path — much larger surface. Option C is cheap and worth shipping alongside B as a config-time warning.


Wayfinder map (the wayfinding view of this issue)

Destination

Any client protocol (OpenAI Chat Completions, OpenAI Responses, Anthropic Messages) can reach any upstream protocol through GoModel, so virtual-model round-robin and failover work across protocol boundaries. This issue's symptom — a conversation breaking when the router crosses the Chat Completions / Responses boundary — disappears because the protocol matrix is complete and translation is loss-bounded and explicit.

Notes

  • Domain: protocol translation providers (internal/providers), the dispatch path (internal/gateway/inference_execute.go), and the virtual-model balancer (internal/virtualmodels).
  • Every session consults the existing translator inventory in internal/providers/responses_adapter.go and the streaming converters.
  • Architecture decisions are settled; open decisions are ticketed below as child issues.
  • Work only in this repo (weselben/GoModel); no PRs to any other upstream.

Decisions so far

  • Investigation: confirmed root cause (protocol-blind router + missing chat→responses translation + previous_response_id rejection) and the failure shapes per direction — see the investigation comment for the full trace and file:line citations.
  • Architecture: extend the existing provider-adapts pattern rather than add a central gateway gate. Providers already translate their own dialect; the chat↔responses pair is the only incomplete pair. No new dispatch layer.
  • Topology: chat completions is the hub (star). Not a canonical internal message representation.
  • Slice shape: translators first (pure additions), then chatgpt wiring (replaces its 501s), then previous_response_id expansion from the local response store (orthogonal, separate decision).
  • Verify chatgpt stream-parser reuse and lock translator signatures — collapseResponsesStream not reusable for streaming (non-streaming collapse); five new functions locked with exact signatures (ConvertChatRequestToResponses, ConvertResponsesResponseToChat, ChatResponsesStreamConverter, ChatViaResponses, StreamChatViaResponses); responses→chat needs index-tracked call_id patching (id arrives late on response.output_item.done).
  • Decide follow-up scope: protocol-aware pool filtering and config-time warning — defer same-protocol filtering (pure optimization after the matrix); no config-time mixed-protocol warning, warn only on missing directed edges.
  • Lock the chat↔responses translation-loss policy — the ADR-0011 field-forwarding ladder already locks it; initial per-field rows for the chat↔responses pair recorded in the ticket.
  • kimicode speaks the Responses API natively (draft PR ENTERPILOT/GoModel#916); live probe confirms kimi's native /responses is stateless (store:true and previous_response_id both 400), so gateway-managed state is the only chaining path even against responses-speaking upstreams.

Not yet specified

  • Documentation updates (docs/features/virtual-models.mdx, docs/features/failover.mdx) once the loss policy is locked.

Out of scope

  • A canonical internal message model / full protocol IR redesign.
  • Migrating the ~10 provider-local ResponsesViaChat wrappers onto shared translators (cleanup only; behavior-preserving, deferred).
  • Decide scope: session-detection gap for bodies over 1 MiB — affinity-only concern once the matrix is complete; separate effort.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions