Skip to content
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
11 changes: 11 additions & 0 deletions cmd/gomodel/docs/docs.go

Some generated files are not rendered by default. Learn more about how customized files appear on GitHub.

11 changes: 11 additions & 0 deletions docs/openapi.json

Some generated files are not rendered by default. Learn more about how customized files appear on GitHub.

8 changes: 6 additions & 2 deletions docs/providers/kimicode.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -7,8 +7,12 @@ keywords: ["Kimi Code", "Moonshot", "quota", "provider setup"]

Kimi Code is an OpenAI-compatible coding assistant served at `https://api.kimi.com/coding/v1`.
GoModel routes chat, model listing, embeddings, and passthrough requests through the shared
OpenAI adapter. The `/v1/responses` endpoint is translated through chat completions, while
files and batches are not supported by the upstream endpoint.
OpenAI adapter. The `/v1/responses` endpoint is forwarded natively to the upstream Responses
API, while files and batches are not supported by the upstream endpoint.

The upstream Responses API keeps no state: requests with `store: true` are rewritten to
`store: false`, and requests with a `previous_response_id` are rejected with an
invalid-request error instead of being answered statelessly.

## Configure

Expand Down
22 changes: 13 additions & 9 deletions internal/core/responses.go
Original file line number Diff line number Diff line change
Expand Up @@ -184,15 +184,19 @@ type ResponsesInputElement struct {

// ResponsesResponse represents the response from the Responses API.
type ResponsesResponse struct {
ID string `json:"id"`
Object string `json:"object"` // "response"
CreatedAt int64 `json:"created_at"`
Model string `json:"model"`
Provider string `json:"provider"`
Status string `json:"status"` // "completed", "incomplete", "failed", "in_progress"
Output []ResponsesOutputItem `json:"output"`
Usage *ResponsesUsage `json:"usage,omitempty"`
Error *ResponsesError `json:"error,omitempty"`
ID string `json:"id"`
Object string `json:"object"` // "response"
CreatedAt int64 `json:"created_at"`
// CompletedAt is the upstream completion timestamp; zero while the
// response is still in progress.
CompletedAt int64 `json:"completed_at,omitempty"`
Store *bool `json:"store,omitempty"`
Model string `json:"model"`
Provider string `json:"provider"`
Status string `json:"status"` // "completed", "incomplete", "failed", "in_progress"
Output []ResponsesOutputItem `json:"output"`
Usage *ResponsesUsage `json:"usage,omitempty"`
Error *ResponsesError `json:"error,omitempty"`
// IncompleteDetails explains a status of "incomplete": the model hit
// max_output_tokens, was stopped by a content filter, or the upstream
// stream was interrupted.
Expand Down
49 changes: 49 additions & 0 deletions internal/core/responses_json_test.go
Original file line number Diff line number Diff line change
Expand Up @@ -938,3 +938,52 @@ func TestResponsesBlocksFromContentPartsKeepsLargeIntegersExact(t *testing.T) {
require.Contains(t, string(encoded), want)
}
}

func TestResponsesResponseJSON_CompletedAtAndStoreRoundTrip(t *testing.T) {
cases := []struct {
name string
storeJSON string
wantStore *bool
wantStored bool
}{
{name: "store false", storeJSON: `,"store":false`, wantStore: new(bool), wantStored: true},
{name: "store absent", storeJSON: ``, wantStore: nil, wantStored: false},
}

for _, tc := range cases {
t.Run(tc.name, func(t *testing.T) {
var resp ResponsesResponse
err := json.Unmarshal([]byte(`{
"id":"resp_123",
"object":"response",
"created_at":1677652288,
"completed_at":1677652299,
"model":"gpt-4o-mini",
"provider":"openai",
"status":"completed",
"output":[]`+tc.storeJSON+`
}`), &resp)
require.NoError(t, err)
require.Equal(t, int64(1677652299), resp.CompletedAt)
if tc.wantStore == nil {
require.Nil(t, resp.Store)
} else {
require.NotNil(t, resp.Store)
require.Equal(t, *tc.wantStore, *resp.Store)
}

body, err := json.Marshal(resp)
require.NoError(t, err)

var decoded map[string]any
err = json.Unmarshal(body, &decoded)
require.NoError(t, err)
require.Equal(t, float64(1677652299), decoded["completed_at"])
store, present := decoded["store"]
require.Equal(t, tc.wantStored, present, "store presence mismatch in marshaled payload: %s", string(body))
if tc.wantStore != nil {
require.Equal(t, *tc.wantStore, store)
}
})
}
}
9 changes: 6 additions & 3 deletions internal/core/types.go
Original file line number Diff line number Diff line change
Expand Up @@ -155,9 +155,12 @@ type ResponseMessage struct {
// PromptTokensDetails holds extended input token breakdown (OpenAI/xAI).
type PromptTokensDetails struct {
CachedTokens int `json:"cached_tokens"`
AudioTokens int `json:"audio_tokens"`
TextTokens int `json:"text_tokens"`
ImageTokens int `json:"image_tokens"`
// CacheWriteTokens counts tokens written to the provider's prompt cache
// (Kimi Code / Anthropic-style cache creation).
CacheWriteTokens int `json:"cache_write_tokens,omitempty"`

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Regenerate OpenAPI schemas

This adds cache_write_tokens to the public usage model, but the checked-in schemas in docs/openapi.json and cmd/gomodel/docs/docs.go still omit it. Generated clients and API consumers therefore cannot discover or model a value that runtime Responses payloads now expose. This is non-blocking, but the generated API documentation should be refreshed.

Note: If this suggestion doesn't match your team's coding style, reply to this and let me know. I'll remember it for next time!

Artifacts

Evidence from the check

  • Python checker authored for this validation; it parses the checked-in JSON schema and inspects the generated Go schema definition, showing whether cache_write_tokens is publicly exposed.

Command output from the check

  • Captured output of the executed checker from /home/user/repo with exit code 0; it confirms the source field exists while both public schemas omit it.

View artifacts

T-Rex Ran code and verified through T-Rex

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Should be fixed in 001ddeb

AudioTokens int `json:"audio_tokens"`
TextTokens int `json:"text_tokens"`
ImageTokens int `json:"image_tokens"`
}

// CompletionTokensDetails holds extended output token breakdown (OpenAI/xAI).
Expand Down
31 changes: 31 additions & 0 deletions internal/core/usage_json_test.go
Original file line number Diff line number Diff line change
Expand Up @@ -118,3 +118,34 @@ func TestResponsesUsageMarshalJSON_UsesResponsesDetailFieldNames(t *testing.T) {
_, exists = payload["raw_usage"]
require.False(t, exists, "did not expect raw_usage in marshaled responses payload: %s", string(body))
}

func TestResponsesUsageJSON_CacheWriteTokensRoundTrip(t *testing.T) {
var usage ResponsesUsage
err := json.Unmarshal([]byte(`{
"input_tokens": 100,
"output_tokens": 20,
"total_tokens": 120,
"input_tokens_details": {
"cached_tokens": 88,
"cache_write_tokens": 12
}
}`), &usage)
require.NoError(t, err)
require.NotNil(t, usage.PromptTokensDetails)
require.Equal(t, 88, usage.PromptTokensDetails.CachedTokens)
require.Equal(t, 12, usage.PromptTokensDetails.CacheWriteTokens)

body, err := json.Marshal(usage)
require.NoError(t, err)

var payload map[string]any
err = json.Unmarshal(body, &payload)
require.NoError(t, err)

inputDetails, ok := payload["input_tokens_details"].(map[string]any)
require.True(t, ok, "decoded input_tokens_details = %#v, want object", payload["input_tokens_details"])
require.Equal(t, float64(88), inputDetails["cached_tokens"])
require.Equal(t, float64(12), inputDetails["cache_write_tokens"])
_, exists := payload["raw_usage"]
require.False(t, exists, "did not expect cache_write_tokens to leak into raw_usage: %s", string(body))
}
97 changes: 89 additions & 8 deletions internal/providers/kimicode/kimicode.go
Original file line number Diff line number Diff line change
@@ -1,11 +1,18 @@
// Package kimicode provides Kimi Code API integration for the LLM gateway.
//
// The "kimicode" provider routes to Kimi Code's OpenAI-compatible chat
// completions endpoint, so all transport goes through the shared chat-centric
// adapter and model IDs are forwarded unchanged.
// The "kimicode" provider routes to Kimi Code's OpenAI-compatible API: chat
// completions, model listing, embeddings, and passthrough go through the
// shared chat-centric adapter, while the Responses API is served natively by
// the upstream /responses endpoint. Kimi Code retains no responses, so
// previous_response_id is rejected with an invalid-request error and
// store=true is pinned to false.
package kimicode

import (
"context"
"io"
"net/http"

"github.com/enterpilot/gomodel/internal/core"
"github.com/enterpilot/gomodel/internal/providers"
"github.com/enterpilot/gomodel/internal/providers/openai"
Expand All @@ -23,19 +30,93 @@ var Registration = providers.Registration{
}

// Provider implements the core.Provider interface for Kimi Code. Kimi Code is
// OpenAI-compatible, so all transport goes through the shared chat-centric
// OpenAI-compatible, so most transport goes through the shared chat-centric
// adapter: chat completions, model listing, embeddings, and passthrough are
// exposed via the embedded *openai.ChatCompatible.
// exposed via the embedded *openai.ChatCompatible. The Responses API is
// forwarded natively to the upstream /responses endpoint (rejecting
// previous_response_id and pinning store to false — see Responses below).
type Provider struct {
*openai.ChatCompatible
responses *openai.CompatibleProvider
}

var _ core.Provider = (*Provider)(nil)

// New creates a new Kimi Code provider.
func New(cfg providers.ProviderConfig, opts providers.ProviderOptions) core.Provider {
return &Provider{openai.NewChatCompatible(cfg.APIKey, opts, openai.CompatibleProviderConfig{
baseURL := providers.ResolveBaseURL(cfg.BaseURL, defaultBaseURL)
return &Provider{
ChatCompatible: openai.NewChatCompatible(cfg.APIKey, opts, compatibleConfig(baseURL)),
responses: openai.NewCompatibleProvider(cfg.APIKey, opts, compatibleConfig(baseURL)),
}
}

// compatibleConfig is the shared OpenAI-compatible configuration for both
// adapter instances. SetHeaders defaults to plain Bearer auth because
// NewCompatibleProvider (unlike NewChatCompatible) applies no default.
func compatibleConfig(baseURL string) openai.CompatibleProviderConfig {
return openai.CompatibleProviderConfig{
ProviderName: "kimicode",
BaseURL: providers.ResolveBaseURL(cfg.BaseURL, defaultBaseURL),
})}
BaseURL: baseURL,
SetHeaders: func(req *http.Request, apiKey string) {
providers.SetAuthHeaders(req, apiKey, providers.AuthHeaderConfig{AuthScheme: "Bearer "})
},
}
}

// SetBaseURL overrides the upstream endpoint for both the chat-centric
// adapter and the native Responses adapter.
func (p *Provider) SetBaseURL(url string) {
p.ChatCompatible.SetBaseURL(url)
p.responses.SetBaseURL(url)
}

// Responses serves the Responses API natively through the upstream /responses
// endpoint. Kimi Code retains no responses, so a non-empty
// previous_response_id is rejected with an invalid-request error before any
// upstream call; store=true is pinned to false by adaptResponsesRequest.
func (p *Provider) Responses(ctx context.Context, req *core.ResponsesRequest) (*core.ResponsesResponse, error) {
if err := rejectPreviousResponseID(req); err != nil {
return nil, err
}
return p.responses.Responses(ctx, adaptResponsesRequest(req))
Comment thread
coderabbitai[bot] marked this conversation as resolved.
}

// StreamResponses forwards the request to the upstream /responses endpoint
// with stream enabled, returning its Responses SSE stream. Like Responses, it
// rejects a non-empty previous_response_id before any upstream call.
func (p *Provider) StreamResponses(ctx context.Context, req *core.ResponsesRequest) (io.ReadCloser, error) {
if err := rejectPreviousResponseID(req); err != nil {
return nil, err
}
return p.responses.StreamResponses(ctx, adaptResponsesRequest(req))
}

// rejectPreviousResponseID fails requests chaining from an earlier response:
// Kimi Code does not retain responses, so a previous response ID can never be
// resolved and answering statelessly would silently drop the conversation
// context the caller expects.
func rejectPreviousResponseID(req *core.ResponsesRequest) error {
if req == nil || req.PreviousResponseID == "" {
return nil
}
return core.NewInvalidRequestError(
"kimicode does not retain responses: previous_response_id is not supported", nil)
}

// adaptResponsesRequest pins store to false: the service retains no
// responses, so store=true fails upstream with a 400 (Postel's law — adapt
// instead of failing). previous_response_id is not adapted here;
// rejectPreviousResponseID rejects it instead.
func adaptResponsesRequest(req *core.ResponsesRequest) *core.ResponsesRequest {
if req == nil {
return nil
}
if req.Store == nil || !*req.Store {
return req
}
cp := *req
disabled := false
cp.Store = &disabled
return &cp
}
Loading
Loading