Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion docs/lld/07-ai-features.md
Original file line number Diff line number Diff line change
Expand Up @@ -49,7 +49,7 @@ flowchart TD
- The body is cleaned before it is forwarded (`sanitizeChatBody` in `server/chatmodels.go`; the body is read with a 1 MB limit). Only `model`, `messages`, `stream`, `temperature`, `max_tokens`, `max_completion_tokens`, `reasoning_effort` and `response_format` get through, and `response_format` only as `{"type":"json_object"}`.
- An admin can change the numbers of Genie (LLD 13 section 3, "Genie limits"): what it is told about the page (the page reads them from `settings.js`, `genieContext`), how long its conversation is, the answer size of each model (`answerCap` replaces the built-in cap in `sanitizeChatBody`), and OpenRouter's time limit and hosts (`openRouterRouting`).
- An admin can switch models off and choose the default (LLD 13 section 3, `genie.disabledModels` and `genie.defaultModel`). A request for a model that is off is answered `503 model_unavailable`, code `model_disabled`, before the balance is touched; the pages read the same lists from `settings.js` and do not offer the model (`js/model-choice.js`), starting visitors on the default. A visitor whose page is still open sees a card with no "Try again" and a button for another model.
- Only three models are accepted: `gpt-6-luna`, `gpt-4o-mini` and `google/gemma-4-31b-it`. Any other name (or none) is answered by `gpt-4o-mini`, so older callers such as the blog editor keep working.
- Only these models are accepted: the three built-in ones, `gpt-6-luna`, `gpt-4o-mini` and `google/gemma-4-31b-it`, and the OpenRouter models an admin added in the dashboard and switched on (`server/custom_models.go`, `GenieSettings.CustomModels`; `modelByID` finds either). Any other name (or none) is answered by the default model, so older callers such as the blog editor keep working. An added model has its own answer size, room for thinking tokens (`ThinkingRoom`, added on top of the answer size, because the thinking of a model that thinks first counts against `max_tokens`), whether its host takes JSON mode (`response_format` is dropped when it does not) and the providers that may answer (`provider.only`; none named means OpenRouter chooses and the body carries no provider rule). A browser cannot change any of these.
- Each model keeps only its own settings. Luna: `reasoning_effort` (`none`, `low`, `medium` or `high`, else `low`) and `max_completion_tokens` (default by effort 1000 / 2000 / 3000 / 4000, at most 7000); no `temperature` or `max_tokens`. 4o mini: `temperature` (0 to 2) and `max_tokens` (default 800, at most 3000); no effort. Gemma: the same as 4o mini, with `max_tokens` at most 4000. A request that is not a chat request (no `messages` list) gets a 400 with no balance charged.
- The model, effort and token lists exist in three places that have to match: `server/chatmodels.go`, `js/model-choice.js` and, for the allowed efforts, the tests in `server/chatmodels_test.go`. A model change is checked first with a `curl` of the exact request body against `/v1/chat/completions`, because models differ in which parameters they accept (Luna refuses `max_tokens`).
- OpenAI errors are returned with their original status and body (the proxy used to answer 200 for them; the status is now passed on, which is what lets the question generator's retry on 400 work). Proxy errors use OpenAI's error shape: `{"error":{"message","type","param","code"}}`.
Expand Down
8 changes: 8 additions & 0 deletions docs/lld/13-admin-dashboard.md
Original file line number Diff line number Diff line change
Expand Up @@ -128,6 +128,14 @@ The dashboard's own reads (GET under `/admin/` that succeed) are left out of the

`statsStore` counts, where `wrapControls` lets a terminal through, one terminal per language per day, and the visitors of the day. A visitor is a keyed hash of the account or, for a guest, the address, different on each day, so one visitor who opens several terminals counts once. Only today keeps these hashes (to count across a restart); older days keep numbers only, 60 days in all. The file is `admin-stats.json`, saved at most every 30 seconds. A gateway counts the terminals of its workers; a worker's own port is not counted.

### OpenRouter models

The three built-in models (GPT-6 Luna, GPT-4o mini, Gemma 4 31B) are fixed: they can be switched off, not removed. Under Site settings, *OpenRouter models* lists the models an admin added, each with a switch, a name, an answer size, room for thinking, "its host takes JSON mode", the providers allowed to answer and a "free tier" mark; *Add model* takes an OpenRouter id (`author/model-name`, lower case, optionally `:free`) and adds it off, *Remove* takes it out after a question. Nothing else in the code has to change to use a new model.

The list is `GenieSettings.CustomModels`, part of the site settings document, so it is stored wherever the settings are: `settings.json`, or MongoDB or Firestore when one is configured (LLD 05), and every instance reads the same list. `nil` (never saved) shows the two free models that were tried (Nemotron 3 Super and North mini code), off; an admin who removes all of them keeps an empty list. `normalize` checks the ids (the OpenRouter pattern, not a built-in id, no duplicates), the name, the sizes and the providers, at most 20 models, and that at least one model stays on; the default model may be any model that is on. A page from before the list existed does not send it, and then the stored list is kept. Changes go to the audit log ("OpenRouter model qwen/qwen3-coder:free added (off)", "... switched on", "... removed").

Visitors: `settings.js` carries the added models that are on (`customModels`), and `model-choice.js` adds them to the picker under OpenRouter, with a "Free" tag. A model that is off is not listed and `handleChatProxy` refuses it like any switched-off model. A free model's provider may keep what is sent, so the dashboard says so and the privacy page does too; a model with no providers named is routed by OpenRouter.

### Languages per node

The site-wide switch turns a language off for everybody. The Languages card in a node's drawer (*Workers*, then the node) decides per node, for every language that has a terminal (JavaScript runs in the browser and is not listed): **Default** follows what the node declares (`--worker-languages`; a node that did not list its languages runs everything), **Off** refuses new terminals of it although the node can run it, **On** takes it although the node did not declare it, for a language installed after the worker started. Open terminals keep running. The choices are saved with the site settings (`SiteSettings.NodeLanguages`, node id to `{off, on}`), by `/admin/workers/<id>/languages` only: the Settings page does not send them and keeps what is stored. They are audited ("worker pi-1: rappel switched off").
Expand Down
96 changes: 96 additions & 0 deletions docs/model-costs.txt
Original file line number Diff line number Diff line change
@@ -0,0 +1,96 @@
OPENREPL: MODELS AND TOKEN COSTS
Generated 2026-10-08 from OpenRouter's public catalog (https://openrouter.ai/api/v1/models). Prices change; check the catalog before relying on them.
Prices are US dollars per 1,000,000 tokens (in = prompt, out = completion). ctx = context window, maxout = most tokens of one answer.
json = the host takes response_format json_object (Agent mode); tools = takes tool calls.

====================================================================================================
1. WHAT OPENREPL USES TODAY
====================================================================================================
openai/gpt-6-luna ctx 1050000 maxout 128000 in 0.1 out 0.5 json=Y tools=Y
openai/gpt-4o-mini ctx 128000 maxout 16384 in 0.15 out 0.6 json=Y tools=Y
google/gemma-4-31b-it ctx 262144 maxout 16384 in 0.09 out 0.34 json=Y tools=Y
(GPT-6 Luna and GPT-4o mini are billed by OpenAI directly with the OpenAI key; Gemma is billed through OpenRouter credit.)
Example: an answer of 2,000 tokens on a model that costs 0.60 per 1M output tokens costs 2000 / 1,000,000 * 0.60 = $0.0012.

====================================================================================================
2. OPENROUTER FREE MODELS (20)
====================================================================================================
Cost: 0 per token. Limits on the account, shared by every visitor: 20 requests a minute, 50 a day (1,000 a day once $10 of credit
has been bought at any time). A negative credit balance can still give 402. Free providers may keep and use what is sent.
Models that think first spend tokens before the answer: leave 3,000 or more tokens of room.

apodex/apodex-1.1-mini:free ctx 262144 maxout 235929 in 0 (free) out 0 (free) json=Y tools=Y
cohere/north-mini-code:free ctx 256000 maxout 64000 in 0 (free) out 0 (free) json=- tools=Y
dots-studio/dots-3-note-preview:free ctx 512000 maxout 460800 in 0 (free) out 0 (free) json=Y tools=Y expires 2026-12-31
google/gemma-4-26b-a4b-it:free ctx 262144 maxout 32768 in 0 (free) out 0 (free) json=Y tools=Y
google/gemma-4-31b-it:free ctx 262144 maxout 32768 in 0 (free) out 0 (free) json=Y tools=Y
google/lyria-3-clip-preview ctx 1048576 maxout 65536 in 0 (free) out 0 (free) json=Y tools=-
google/lyria-3-pro-preview ctx 1048576 maxout 65536 in 0 (free) out 0 (free) json=Y tools=-
inclusionai/ling-3.0-flash-sante:free ctx 262144 maxout 32768 in 0 (free) out 0 (free) json=- tools=Y
inclusionai/ling-3.1-flash ctx 262144 maxout 32768 in 0 (free) out 0 (free) json=- tools=Y
liquid/lfm-2.5-2.6b:free ctx 65536 maxout 8192 in 0 (free) out 0 (free) json=Y tools=Y
nvidia/nemotron-3-nano-omni-30b-a3b-reasoning:free ctx 256000 maxout 65536 in 0 (free) out 0 (free) json=- tools=Y
nvidia/nemotron-3-super-120b-a12b:free ctx 262144 maxout 235929 in 0 (free) out 0 (free) json=Y tools=Y
nvidia/nemotron-3-ultra-550b-a55b:free ctx 1000000 maxout 65536 in 0 (free) out 0 (free) json=- tools=Y
nvidia/nemotron-3.5-content-safety:free ctx 128000 maxout 8192 in 0 (free) out 0 (free) json=- tools=-
nvidia/nemotron-3.5-lightning:free ctx 1000000 maxout 65536 in 0 (free) out 0 (free) json=- tools=Y
openrouter/free ctx 200000 maxout - in 0 (free) out 0 (free) json=Y tools=Y
poolside/laguna-s-2.1:free ctx 262144 maxout 32768 in 0 (free) out 0 (free) json=- tools=Y expires 2026-10-31
poolside/laguna-xs-2.1:free ctx 262144 maxout 32768 in 0 (free) out 0 (free) json=- tools=Y expires 2026-10-31
thinkingmachines/inkling-small:free ctx 1048576 maxout 262144 in 0 (free) out 0 (free) json=- tools=Y
thinkingmachines/inkling:free ctx 1048576 maxout 262144 in 0 (free) out 0 (free) json=- tools=Y

Tried through OpenRouter with the owner's key on 2026-10-08:
works nvidia/nemotron-3-super-120b-a12b:free (thinks first, takes JSON mode)
works cohere/north-mini-code:free (thinks first, no JSON mode)
weak apodex/apodex-1.1-mini:free (used all 300 tokens thinking, empty answer)
429 google/gemma-4-31b-it:free and google/gemma-4-26b-a4b-it:free (Google AI Studio shared pool exhausted)
not tried: everything else in this section

====================================================================================================
3. OPENAI MODELS (NOT FREE)
====================================================================================================
OpenAI has no free API models. These are OpenAI's own models as OpenRouter lists them, cheapest input price first;
OpenAI's direct prices are normally the same. Free credit, when OpenAI gives it, is a promotion, not a model.

openai/gpt-oss-20b ctx 131072 maxout 32768 in 0.018 out 0.09 json=Y tools=Y
openai/gpt-oss-20b:batch ctx 131072 maxout 117964 in 0.024 out 0.112 json=Y tools=Y
openai/gpt-5-nano:batch ctx 400000 maxout 128000 in 0.025 out 0.2 json=Y tools=Y
openai/gpt-oss-120b:batch ctx 131072 maxout 117964 in 0.0296 out 0.136 json=Y tools=Y
openai/gpt-oss-120b ctx 131072 maxout 117964 in 0.037 out 0.17 json=Y tools=Y
openai/gpt-6-luna-pro:batch ctx 1050000 maxout 128000 in 0.05 out 0.25 json=Y tools=Y
openai/gpt-6-luna:batch ctx 1050000 maxout 128000 in 0.05 out 0.25 json=Y tools=Y
openai/gpt-5-nano ctx 400000 maxout 128000 in 0.05 out 0.4 json=Y tools=Y
openai/gpt-4.1-nano:batch ctx 1047576 maxout 32768 in 0.05 out 0.2 json=Y tools=Y
openai/gpt-oss-safeguard-20b ctx 131072 maxout 65536 in 0.075 out 0.3 json=Y tools=Y
openai/gpt-4o-mini:batch ctx 128000 maxout 16384 in 0.075 out 0.3 json=Y tools=Y
openai/gpt-6-luna-pro ctx 1050000 maxout 128000 in 0.1 out 0.5 json=Y tools=Y
openai/gpt-6-luna ctx 1050000 maxout 128000 in 0.1 out 0.5 json=Y tools=Y
openai/gpt-5.6-luna-pro:batch ctx 1050000 maxout 128000 in 0.1 out 0.6 json=Y tools=Y
openai/gpt-5.6-luna:batch ctx 1050000 maxout 128000 in 0.1 out 0.6 json=Y tools=Y
openai/gpt-5.4-nano:batch ctx 400000 maxout 128000 in 0.1 out 0.625 json=Y tools=Y
openai/gpt-4.1-nano ctx 1047576 maxout 32768 in 0.1 out 0.4 json=Y tools=Y
openai/gpt-5-mini:batch ctx 400000 maxout 128000 in 0.125 out 1 json=Y tools=Y
openai/gpt-4o-mini ctx 128000 maxout 16384 in 0.15 out 0.6 json=Y tools=Y
openai/gpt-4o-mini-2024-07-18 ctx 128000 maxout 16384 in 0.15 out 0.6 json=Y tools=Y
openai/gpt-5.6-luna-pro ctx 1050000 maxout 128000 in 0.2 out 1.2 json=Y tools=Y
openai/gpt-5.6-luna ctx 1050000 maxout 128000 in 0.2 out 1.2 json=Y tools=Y
openai/gpt-5.4-nano ctx 400000 maxout 128000 in 0.2 out 1.25 json=Y tools=Y
openai/gpt-4.1-mini:batch ctx 1047576 maxout 32768 in 0.2 out 0.8 json=Y tools=Y
openai/gpt-5.1-codex-mini ctx 400000 maxout 128000 in 0.25 out 2 json=Y tools=Y
openai/gpt-5-mini ctx 400000 maxout 128000 in 0.25 out 2 json=Y tools=Y
openai/gpt-3.5-turbo:batch ctx 16385 maxout 4096 in 0.25 out 0.75 json=Y tools=Y
openai/gpt-5.4-mini:batch ctx 400000 maxout 128000 in 0.375 out 2.25 json=Y tools=Y
openai/gpt-4.1-mini ctx 1047576 maxout 32768 in 0.4 out 1.6 json=Y tools=Y
openai/gpt-3.5-turbo ctx 16385 maxout 4096 in 0.5 out 1.5 json=Y tools=Y
openai/o4-mini:batch ctx 200000 maxout 100000 in 0.55 out 2.2 json=Y tools=Y
openai/o3-mini:batch ctx 200000 maxout 100000 in 0.55 out 2.2 json=Y tools=Y
openai/gpt-audio-mini ctx 128000 maxout 16384 in 0.6 out 2.4 json=Y tools=Y
openai/gpt-5.1:batch ctx 400000 maxout 128000 in 0.625 out 5 json=Y tools=Y
openai/gpt-5:batch ctx 400000 maxout 128000 in 0.625 out 5 json=Y tools=Y
openai/gpt-5.4-mini ctx 400000 maxout 128000 in 0.75 out 4.5 json=Y tools=Y
openai/gpt-5.2:batch ctx 400000 maxout 128000 in 0.875 out 7 json=Y tools=Y
openai/gpt-6-sol-pro:batch ctx 1050000 maxout 128000 in 1 out 5 json=Y tools=Y
openai/gpt-6-sol:batch ctx 1050000 maxout 128000 in 1 out 5 json=Y tools=Y
openai/gpt-5.6-terra-pro:batch ctx 1050000 maxout 128000 in 1 out 6 json=Y tools=Y
... 62 more openai/* models in the catalog
Loading
Loading