diff --git a/docs/lld/07-ai-features.md b/docs/lld/07-ai-features.md index c29ce356..94f9e799 100644 --- a/docs/lld/07-ai-features.md +++ b/docs/lld/07-ai-features.md @@ -49,7 +49,7 @@ flowchart TD - The body is cleaned before it is forwarded (`sanitizeChatBody` in `server/chatmodels.go`; the body is read with a 1 MB limit). Only `model`, `messages`, `stream`, `temperature`, `max_tokens`, `max_completion_tokens`, `reasoning_effort` and `response_format` get through, and `response_format` only as `{"type":"json_object"}`. - An admin can change the numbers of Genie (LLD 13 section 3, "Genie limits"): what it is told about the page (the page reads them from `settings.js`, `genieContext`), how long its conversation is, the answer size of each model (`answerCap` replaces the built-in cap in `sanitizeChatBody`), and OpenRouter's time limit and hosts (`openRouterRouting`). - An admin can switch models off and choose the default (LLD 13 section 3, `genie.disabledModels` and `genie.defaultModel`). A request for a model that is off is answered `503 model_unavailable`, code `model_disabled`, before the balance is touched; the pages read the same lists from `settings.js` and do not offer the model (`js/model-choice.js`), starting visitors on the default. A visitor whose page is still open sees a card with no "Try again" and a button for another model. -- Only three models are accepted: `gpt-6-luna`, `gpt-4o-mini` and `google/gemma-4-31b-it`. Any other name (or none) is answered by `gpt-4o-mini`, so older callers such as the blog editor keep working. +- Only these models are accepted: the three built-in ones, `gpt-6-luna`, `gpt-4o-mini` and `google/gemma-4-31b-it`, and the OpenRouter models an admin added in the dashboard and switched on (`server/custom_models.go`, `GenieSettings.CustomModels`; `modelByID` finds either). Any other name (or none) is answered by the default model, so older callers such as the blog editor keep working. An added model has its own answer size, room for thinking tokens (`ThinkingRoom`, added on top of the answer size, because the thinking of a model that thinks first counts against `max_tokens`), whether its host takes JSON mode (`response_format` is dropped when it does not) and the providers that may answer (`provider.only`; none named means OpenRouter chooses and the body carries no provider rule). A browser cannot change any of these. - Each model keeps only its own settings. Luna: `reasoning_effort` (`none`, `low`, `medium` or `high`, else `low`) and `max_completion_tokens` (default by effort 1000 / 2000 / 3000 / 4000, at most 7000); no `temperature` or `max_tokens`. 4o mini: `temperature` (0 to 2) and `max_tokens` (default 800, at most 3000); no effort. Gemma: the same as 4o mini, with `max_tokens` at most 4000. A request that is not a chat request (no `messages` list) gets a 400 with no balance charged. - The model, effort and token lists exist in three places that have to match: `server/chatmodels.go`, `js/model-choice.js` and, for the allowed efforts, the tests in `server/chatmodels_test.go`. A model change is checked first with a `curl` of the exact request body against `/v1/chat/completions`, because models differ in which parameters they accept (Luna refuses `max_tokens`). - OpenAI errors are returned with their original status and body (the proxy used to answer 200 for them; the status is now passed on, which is what lets the question generator's retry on 400 work). Proxy errors use OpenAI's error shape: `{"error":{"message","type","param","code"}}`. diff --git a/docs/lld/13-admin-dashboard.md b/docs/lld/13-admin-dashboard.md index c480743c..ede3b116 100644 --- a/docs/lld/13-admin-dashboard.md +++ b/docs/lld/13-admin-dashboard.md @@ -128,6 +128,14 @@ The dashboard's own reads (GET under `/admin/` that succeed) are left out of the `statsStore` counts, where `wrapControls` lets a terminal through, one terminal per language per day, and the visitors of the day. A visitor is a keyed hash of the account or, for a guest, the address, different on each day, so one visitor who opens several terminals counts once. Only today keeps these hashes (to count across a restart); older days keep numbers only, 60 days in all. The file is `admin-stats.json`, saved at most every 30 seconds. A gateway counts the terminals of its workers; a worker's own port is not counted. +### OpenRouter models + +The three built-in models (GPT-6 Luna, GPT-4o mini, Gemma 4 31B) are fixed: they can be switched off, not removed. Under Site settings, *OpenRouter models* lists the models an admin added, each with a switch, a name, an answer size, room for thinking, "its host takes JSON mode", the providers allowed to answer and a "free tier" mark; *Add model* takes an OpenRouter id (`author/model-name`, lower case, optionally `:free`) and adds it off, *Remove* takes it out after a question. Nothing else in the code has to change to use a new model. + +The list is `GenieSettings.CustomModels`, part of the site settings document, so it is stored wherever the settings are: `settings.json`, or MongoDB or Firestore when one is configured (LLD 05), and every instance reads the same list. `nil` (never saved) shows the two free models that were tried (Nemotron 3 Super and North mini code), off; an admin who removes all of them keeps an empty list. `normalize` checks the ids (the OpenRouter pattern, not a built-in id, no duplicates), the name, the sizes and the providers, at most 20 models, and that at least one model stays on; the default model may be any model that is on. A page from before the list existed does not send it, and then the stored list is kept. Changes go to the audit log ("OpenRouter model qwen/qwen3-coder:free added (off)", "... switched on", "... removed"). + +Visitors: `settings.js` carries the added models that are on (`customModels`), and `model-choice.js` adds them to the picker under OpenRouter, with a "Free" tag. A model that is off is not listed and `handleChatProxy` refuses it like any switched-off model. A free model's provider may keep what is sent, so the dashboard says so and the privacy page does too; a model with no providers named is routed by OpenRouter. + ### Languages per node The site-wide switch turns a language off for everybody. The Languages card in a node's drawer (*Workers*, then the node) decides per node, for every language that has a terminal (JavaScript runs in the browser and is not listed): **Default** follows what the node declares (`--worker-languages`; a node that did not list its languages runs everything), **Off** refuses new terminals of it although the node can run it, **On** takes it although the node did not declare it, for a language installed after the worker started. Open terminals keep running. The choices are saved with the site settings (`SiteSettings.NodeLanguages`, node id to `{off, on}`), by `/admin/workers//languages` only: the Settings page does not send them and keeps what is stored. They are audited ("worker pi-1: rappel switched off"). diff --git a/docs/model-costs.txt b/docs/model-costs.txt new file mode 100644 index 00000000..87341ad0 --- /dev/null +++ b/docs/model-costs.txt @@ -0,0 +1,96 @@ +OPENREPL: MODELS AND TOKEN COSTS +Generated 2026-10-08 from OpenRouter's public catalog (https://openrouter.ai/api/v1/models). Prices change; check the catalog before relying on them. +Prices are US dollars per 1,000,000 tokens (in = prompt, out = completion). ctx = context window, maxout = most tokens of one answer. +json = the host takes response_format json_object (Agent mode); tools = takes tool calls. + +==================================================================================================== +1. WHAT OPENREPL USES TODAY +==================================================================================================== +openai/gpt-6-luna ctx 1050000 maxout 128000 in 0.1 out 0.5 json=Y tools=Y +openai/gpt-4o-mini ctx 128000 maxout 16384 in 0.15 out 0.6 json=Y tools=Y +google/gemma-4-31b-it ctx 262144 maxout 16384 in 0.09 out 0.34 json=Y tools=Y +(GPT-6 Luna and GPT-4o mini are billed by OpenAI directly with the OpenAI key; Gemma is billed through OpenRouter credit.) +Example: an answer of 2,000 tokens on a model that costs 0.60 per 1M output tokens costs 2000 / 1,000,000 * 0.60 = $0.0012. + +==================================================================================================== +2. OPENROUTER FREE MODELS (20) +==================================================================================================== +Cost: 0 per token. Limits on the account, shared by every visitor: 20 requests a minute, 50 a day (1,000 a day once $10 of credit +has been bought at any time). A negative credit balance can still give 402. Free providers may keep and use what is sent. +Models that think first spend tokens before the answer: leave 3,000 or more tokens of room. + +apodex/apodex-1.1-mini:free ctx 262144 maxout 235929 in 0 (free) out 0 (free) json=Y tools=Y +cohere/north-mini-code:free ctx 256000 maxout 64000 in 0 (free) out 0 (free) json=- tools=Y +dots-studio/dots-3-note-preview:free ctx 512000 maxout 460800 in 0 (free) out 0 (free) json=Y tools=Y expires 2026-12-31 +google/gemma-4-26b-a4b-it:free ctx 262144 maxout 32768 in 0 (free) out 0 (free) json=Y tools=Y +google/gemma-4-31b-it:free ctx 262144 maxout 32768 in 0 (free) out 0 (free) json=Y tools=Y +google/lyria-3-clip-preview ctx 1048576 maxout 65536 in 0 (free) out 0 (free) json=Y tools=- +google/lyria-3-pro-preview ctx 1048576 maxout 65536 in 0 (free) out 0 (free) json=Y tools=- +inclusionai/ling-3.0-flash-sante:free ctx 262144 maxout 32768 in 0 (free) out 0 (free) json=- tools=Y +inclusionai/ling-3.1-flash ctx 262144 maxout 32768 in 0 (free) out 0 (free) json=- tools=Y +liquid/lfm-2.5-2.6b:free ctx 65536 maxout 8192 in 0 (free) out 0 (free) json=Y tools=Y +nvidia/nemotron-3-nano-omni-30b-a3b-reasoning:free ctx 256000 maxout 65536 in 0 (free) out 0 (free) json=- tools=Y +nvidia/nemotron-3-super-120b-a12b:free ctx 262144 maxout 235929 in 0 (free) out 0 (free) json=Y tools=Y +nvidia/nemotron-3-ultra-550b-a55b:free ctx 1000000 maxout 65536 in 0 (free) out 0 (free) json=- tools=Y +nvidia/nemotron-3.5-content-safety:free ctx 128000 maxout 8192 in 0 (free) out 0 (free) json=- tools=- +nvidia/nemotron-3.5-lightning:free ctx 1000000 maxout 65536 in 0 (free) out 0 (free) json=- tools=Y +openrouter/free ctx 200000 maxout - in 0 (free) out 0 (free) json=Y tools=Y +poolside/laguna-s-2.1:free ctx 262144 maxout 32768 in 0 (free) out 0 (free) json=- tools=Y expires 2026-10-31 +poolside/laguna-xs-2.1:free ctx 262144 maxout 32768 in 0 (free) out 0 (free) json=- tools=Y expires 2026-10-31 +thinkingmachines/inkling-small:free ctx 1048576 maxout 262144 in 0 (free) out 0 (free) json=- tools=Y +thinkingmachines/inkling:free ctx 1048576 maxout 262144 in 0 (free) out 0 (free) json=- tools=Y + +Tried through OpenRouter with the owner's key on 2026-10-08: + works nvidia/nemotron-3-super-120b-a12b:free (thinks first, takes JSON mode) + works cohere/north-mini-code:free (thinks first, no JSON mode) + weak apodex/apodex-1.1-mini:free (used all 300 tokens thinking, empty answer) + 429 google/gemma-4-31b-it:free and google/gemma-4-26b-a4b-it:free (Google AI Studio shared pool exhausted) + not tried: everything else in this section + +==================================================================================================== +3. OPENAI MODELS (NOT FREE) +==================================================================================================== +OpenAI has no free API models. These are OpenAI's own models as OpenRouter lists them, cheapest input price first; +OpenAI's direct prices are normally the same. Free credit, when OpenAI gives it, is a promotion, not a model. + +openai/gpt-oss-20b ctx 131072 maxout 32768 in 0.018 out 0.09 json=Y tools=Y +openai/gpt-oss-20b:batch ctx 131072 maxout 117964 in 0.024 out 0.112 json=Y tools=Y +openai/gpt-5-nano:batch ctx 400000 maxout 128000 in 0.025 out 0.2 json=Y tools=Y +openai/gpt-oss-120b:batch ctx 131072 maxout 117964 in 0.0296 out 0.136 json=Y tools=Y +openai/gpt-oss-120b ctx 131072 maxout 117964 in 0.037 out 0.17 json=Y tools=Y +openai/gpt-6-luna-pro:batch ctx 1050000 maxout 128000 in 0.05 out 0.25 json=Y tools=Y +openai/gpt-6-luna:batch ctx 1050000 maxout 128000 in 0.05 out 0.25 json=Y tools=Y +openai/gpt-5-nano ctx 400000 maxout 128000 in 0.05 out 0.4 json=Y tools=Y +openai/gpt-4.1-nano:batch ctx 1047576 maxout 32768 in 0.05 out 0.2 json=Y tools=Y +openai/gpt-oss-safeguard-20b ctx 131072 maxout 65536 in 0.075 out 0.3 json=Y tools=Y +openai/gpt-4o-mini:batch ctx 128000 maxout 16384 in 0.075 out 0.3 json=Y tools=Y +openai/gpt-6-luna-pro ctx 1050000 maxout 128000 in 0.1 out 0.5 json=Y tools=Y +openai/gpt-6-luna ctx 1050000 maxout 128000 in 0.1 out 0.5 json=Y tools=Y +openai/gpt-5.6-luna-pro:batch ctx 1050000 maxout 128000 in 0.1 out 0.6 json=Y tools=Y +openai/gpt-5.6-luna:batch ctx 1050000 maxout 128000 in 0.1 out 0.6 json=Y tools=Y +openai/gpt-5.4-nano:batch ctx 400000 maxout 128000 in 0.1 out 0.625 json=Y tools=Y +openai/gpt-4.1-nano ctx 1047576 maxout 32768 in 0.1 out 0.4 json=Y tools=Y +openai/gpt-5-mini:batch ctx 400000 maxout 128000 in 0.125 out 1 json=Y tools=Y +openai/gpt-4o-mini ctx 128000 maxout 16384 in 0.15 out 0.6 json=Y tools=Y +openai/gpt-4o-mini-2024-07-18 ctx 128000 maxout 16384 in 0.15 out 0.6 json=Y tools=Y +openai/gpt-5.6-luna-pro ctx 1050000 maxout 128000 in 0.2 out 1.2 json=Y tools=Y +openai/gpt-5.6-luna ctx 1050000 maxout 128000 in 0.2 out 1.2 json=Y tools=Y +openai/gpt-5.4-nano ctx 400000 maxout 128000 in 0.2 out 1.25 json=Y tools=Y +openai/gpt-4.1-mini:batch ctx 1047576 maxout 32768 in 0.2 out 0.8 json=Y tools=Y +openai/gpt-5.1-codex-mini ctx 400000 maxout 128000 in 0.25 out 2 json=Y tools=Y +openai/gpt-5-mini ctx 400000 maxout 128000 in 0.25 out 2 json=Y tools=Y +openai/gpt-3.5-turbo:batch ctx 16385 maxout 4096 in 0.25 out 0.75 json=Y tools=Y +openai/gpt-5.4-mini:batch ctx 400000 maxout 128000 in 0.375 out 2.25 json=Y tools=Y +openai/gpt-4.1-mini ctx 1047576 maxout 32768 in 0.4 out 1.6 json=Y tools=Y +openai/gpt-3.5-turbo ctx 16385 maxout 4096 in 0.5 out 1.5 json=Y tools=Y +openai/o4-mini:batch ctx 200000 maxout 100000 in 0.55 out 2.2 json=Y tools=Y +openai/o3-mini:batch ctx 200000 maxout 100000 in 0.55 out 2.2 json=Y tools=Y +openai/gpt-audio-mini ctx 128000 maxout 16384 in 0.6 out 2.4 json=Y tools=Y +openai/gpt-5.1:batch ctx 400000 maxout 128000 in 0.625 out 5 json=Y tools=Y +openai/gpt-5:batch ctx 400000 maxout 128000 in 0.625 out 5 json=Y tools=Y +openai/gpt-5.4-mini ctx 400000 maxout 128000 in 0.75 out 4.5 json=Y tools=Y +openai/gpt-5.2:batch ctx 400000 maxout 128000 in 0.875 out 7 json=Y tools=Y +openai/gpt-6-sol-pro:batch ctx 1050000 maxout 128000 in 1 out 5 json=Y tools=Y +openai/gpt-6-sol:batch ctx 1050000 maxout 128000 in 1 out 5 json=Y tools=Y +openai/gpt-5.6-terra-pro:batch ctx 1050000 maxout 128000 in 1 out 6 json=Y tools=Y +... 62 more openai/* models in the catalog diff --git a/scripts/test-openrouter-free-models.sh b/scripts/test-openrouter-free-models.sh new file mode 100755 index 00000000..92943ccc --- /dev/null +++ b/scripts/test-openrouter-free-models.sh @@ -0,0 +1,164 @@ +#!/usr/bin/env bash +# Which of OpenRouter's free models answer right now? +# +# OPENREPL_OPENROUTER_API_KEY=sk-or-... scripts/test-openrouter-free-models.sh [options] [model-id ...] +# +# Without model ids it takes every free chat model from OpenRouter's public +# catalog (zero price in and out, text out only) and sends each one a tiny +# request: reply with {"ok": true}, in JSON mode where the model takes it. +# A model passes when the answer holds {"ok": true}. The key is read from the +# environment and never printed. +# +# Options: +# --dry-run list the models that would be tried, send nothing (needs no key) +# --tokens N answer budget of each request (default 800: models that think first +# spend it before they answer, so a small number looks like a failure) +# --delay SEC wait between requests (default 4: the free tier allows 20 a minute) +# --yes do not ask before sending +# +# The free tier allows 50 requests a day (1,000 once $10 of credit has been bought +# at any time) for the whole account, and the same limits apply to your visitors, +# so a run uses as many of them as there are models. A 429 usually means that the +# provider's shared pool is busy, not that the model is gone: try it again later. + +set -u + +dry=0 +tokens=800 +delay=4 +assume_yes=0 +models=() +while [ $# -gt 0 ]; do + case $1 in + --dry-run) dry=1 ;; + --tokens) tokens=${2:?--tokens needs a number}; shift ;; + --delay) delay=${2:?--delay needs a number of seconds}; shift ;; + --yes | -y) assume_yes=1 ;; + -h | --help) sed -n '2,24p' "$0" | sed 's/^# \{0,1\}//'; exit 0 ;; + -*) echo "unknown option: $1" >&2; exit 2 ;; + *) models+=("$1") ;; + esac + shift +done + +for tool in curl python3; do + command -v "$tool" >/dev/null 2>&1 || { echo "$tool is needed" >&2; exit 2; } +done +case $tokens$delay in *[!0-9.]*) echo "--tokens and --delay take numbers" >&2; exit 2 ;; esac + +if [ "$dry" -eq 0 ] && [ -z "${OPENREPL_OPENROUTER_API_KEY:-}" ]; then + echo "Set OPENREPL_OPENROUTER_API_KEY first (the key is only read from the environment)." >&2 + exit 2 +fi + +# "id json" for each free chat model, from the public catalog +catalog() { + curl -s --max-time 40 https://openrouter.ai/api/v1/models | python3 -c ' +import json, sys +try: + data = json.load(sys.stdin)["data"] +except Exception as e: + sys.exit("could not read the model catalog: %s" % e) +for m in sorted(data, key=lambda m: m["id"]): + p = m.get("pricing", {}) + try: + free = float(p.get("prompt", "1")) == 0 and float(p.get("completion", "1")) == 0 + except ValueError: + free = False + out = (m.get("architecture") or {}).get("output_modalities") or ["text"] + if not free or out != ["text"] or "content-safety" in m["id"]: + continue + print(m["id"], "json" if "response_format" in (m.get("supported_parameters") or []) else "plain") +' +} + +list=$(catalog) || exit 1 +if [ ${#models[@]} -gt 0 ]; then + picked="" + for m in "${models[@]}"; do + line=$(printf '%s\n' "$list" | awk -v id="$m" '$1 == id') + picked="$picked${line:-$m plain}"$'\n' + done + list=${picked%$'\n'} +fi +count=$(printf '%s\n' "$list" | grep -c .) +[ "$count" -gt 0 ] || { echo "no models to try" >&2; exit 1; } + +echo "$count model(s):" +printf '%s\n' "$list" | awk '{ printf " %-52s %s\n", $1, ($2 == "json" ? "(JSON mode)" : "") }' +if [ "$dry" -eq 1 ]; then + echo "(dry run: nothing was sent)" + exit 0 +fi + +echo +echo "This sends $count request(s) and uses as many of the account's free-tier requests today" +echo "(50 a day, 1,000 after \$10 of credit was ever bought), at $delay s apart." +if [ "$assume_yes" -eq 0 ] && [ -t 0 ]; then + printf 'Go on? [y/N] ' + read -r answer + case $answer in y | Y | yes) ;; *) echo "stopped"; exit 1 ;; esac +fi + +work=$(mktemp -d) +trap 'rm -rf "$work"' EXIT + +# one request: prints a status word and the facts after it +try() { + local model=$1 mode=$2 body http secs + body=$(python3 - "$model" "$mode" "$tokens" <<'EOF' +import json, sys +model, mode, tokens = sys.argv[1], sys.argv[2], int(sys.argv[3]) +req = {"model": model, "max_tokens": tokens, + "messages": [{"role": "user", "content": 'Reply with the JSON {"ok": true}'}]} +if mode == "json": + req["response_format"] = {"type": "json_object"} +print(json.dumps(req)) +EOF + ) + read -r http secs < <(curl -s --max-time 120 -o "$work/answer.json" -w '%{http_code} %{time_total}\n' \ + https://openrouter.ai/api/v1/chat/completions \ + -H "Authorization: Bearer $OPENREPL_OPENROUTER_API_KEY" -H "Content-Type: application/json" -d "$body") + python3 - "$work/answer.json" "${http:-000}" "${secs:-0}" <<'EOF' +import json, sys +path, http, secs = sys.argv[1], sys.argv[2], float(sys.argv[3]) +try: + d = json.load(open(path)) +except Exception: + print("FAIL no answer (HTTP %s, %.1fs)" % (http, secs)); sys.exit() +if "error" in d: + e = d["error"]; md = e.get("metadata") or {} + why = (md.get("raw") or e.get("message") or "")[:110].replace("\n", " ") + word = {"429": "BUSY ", "402": "CREDIT", "404": "GONE ", "401": "KEY "}.get(str(e.get("code")), "FAIL ") + print("%s %s %s%s" % (word, e.get("code"), why, (" [" + md["provider_name"] + "]") if md.get("provider_name") else "")) + sys.exit() +try: + c = d["choices"][0]; text = (c["message"].get("content") or "").strip() + u = d.get("usage") or {}; think = (u.get("completion_tokens_details") or {}).get("reasoning_tokens", 0) + facts = "%.1fs, %s tokens (%s thinking), %s, finish=%s" % (secs, u.get("completion_tokens", "?"), think, d.get("provider", "?"), c.get("finish_reason")) + ok = '"ok"' in text and "true" in text.lower() + if ok: print("PASS " + facts) + elif not text: print("EMPTY the budget went on thinking, no answer: " + facts) + else: print("ODD answered %r: %s" % (text[:60], facts)) +except Exception as e: + print("FAIL unexpected answer: %s" % str(d)[:120]) +EOF +} + +pass=0 +results="" +n=0 +while read -r model mode; do + [ -n "$model" ] || continue + n=$((n + 1)) + result=$(try "$model" "$mode") + printf '%-52s %s\n' "$model" "$result" + results="$results$result"$'\n' + [ "$n" -lt "$count" ] && sleep "$delay" +done <<<"$list" + +pass=$(printf '%s' "$results" | grep -c '^PASS') +echo +echo "$pass of $count passed." +echo "PASS = answered {\"ok\": true}. BUSY = rate-limited (try later). EMPTY = thought through the whole budget (raise --tokens)." +echo "CREDIT = 402, check the account balance. GONE = no such model or host. KEY = the key was refused." diff --git a/src/resources/css/admin.css b/src/resources/css/admin.css index e22071cd..2e2046ec 100644 --- a/src/resources/css/admin.css +++ b/src/resources/css/admin.css @@ -518,3 +518,19 @@ dialog .buttons { display: flex; justify-content: flex-end; gap: 8px; margin-top .node-lang .what .sub { margin: 0; overflow-wrap: anywhere; } .node-lang select.field { width: 100%; } .node-lang .status { grid-column: 1 / -1; } + +/* the OpenRouter models an admin adds (Site settings, Genie) */ +.custom-models { display: grid; gap: 10px; margin-bottom: 12px; } +.custom-model { padding: 10px 12px; border: 1px solid var(--line); border-radius: 8px; } +.cm-head { display: flex; align-items: center; gap: 10px; } +.cm-title { flex: 1; min-width: 0; display: grid; gap: 3px; } +.cm-title .mono { color: var(--muted); font-size: 12.5px; overflow-wrap: anywhere; } +.cm-grid { display: grid; grid-template-columns: repeat(auto-fit, minmax(170px, 1fr)); gap: 10px; margin: 10px 0 6px; } +.cm-field { display: grid; gap: 3px; } +.cm-field > span { color: var(--muted); font-size: 12.5px; } +.cm-field small { color: var(--faint); font-size: 11.5px; } +.cm-checks { display: flex; flex-wrap: wrap; gap: 6px 18px; } +.cm-check { display: inline-flex; align-items: center; gap: 7px; } +.cm-add { display: grid; grid-template-columns: minmax(0, 1.4fr) minmax(0, 1fr) auto; gap: 8px; align-items: center; } +@media (max-width: 640px) { .cm-add { grid-template-columns: 1fr; } } + diff --git a/src/resources/js/admin.js b/src/resources/js/admin.js index 1d1b5d78..947a7eb3 100644 --- a/src/resources/js/admin.js +++ b/src/resources/js/admin.js @@ -912,6 +912,24 @@ root.appendChild(pageHead('Site settings', 'Switches that change what every visitor sees. Changes apply to the next page a visitor loads.')); root.appendChild(body); + // ---- the OpenRouter models an admin adds (the three built-in ones stay as they are) + var CUSTOM_ID = /^[a-z0-9][a-z0-9._-]{0,63}\/[a-z0-9][a-z0-9._:+-]{0,95}$/; + function hostsFrom(text) { + var seen = {}, out = []; + String(text || '').toLowerCase().split(/[\s,]+/).forEach(function (x) { if (x && !seen[x]) { seen[x] = 1; out.push(x); } }); + return out; + } + // one shape for a model from the server and from a row of the form, so that they compare + function cleanCustom(c) { + var id = String(c.id || '').trim().toLowerCase(); + return { id: id, name: String(c.name || '').trim() || id, enabled: !!c.enabled, free: !!c.free || /:free$/.test(id), + answerTokens: int(c.answerTokens), thinkingRoom: int(c.thinkingRoom), jsonMode: !!c.jsonMode, hosts: (c.hosts || []).slice() }; + } + function readCustom(r) { + return cleanCustom({ id: r.id, name: r.name.value, enabled: r.on.checked, free: r.free.checked, answerTokens: r.answer.value, + thinkingRoom: r.room.value, jsonMode: r.json.checked, hosts: hostsFrom(r.hosts.value) }); + } + function snapshot() { return { colorOfTheDay: form.colour.checked, @@ -920,6 +938,7 @@ disabledLanguages: languages.filter(function (l) { return !form.langs[l.value].checked; }).map(function (l) { return l.value; }).sort(), genie: { disabled: !form.genie.checked, guestPerMinute: num(form.guest.value), userPerMinute: num(form.user.value), disabledModels: models.filter(function (m) { return !form.models[m.id].checked; }).map(function (m) { return m.id; }).sort(), + customModels: (form.custom || []).map(readCustom), defaultModel: form.defModel.value, contextEditorChars: int(form.ctxEditor.value), contextTerminalChars: int(form.ctxTermChars.value), contextTerminalLines: int(form.ctxTermLines.value), historyMessages: int(form.history.value), @@ -947,7 +966,8 @@ maintenance: { enabled: !!s.maintenance.enabled, message: s.maintenance.message || '' }, disabledLanguages: (s.disabledLanguages || []).slice().sort(), genie: { disabled: !!s.genie.disabled, guestPerMinute: s.genie.guestPerMinute || 0, userPerMinute: s.genie.userPerMinute || 0, - disabledModels: (s.genie.disabledModels || []).slice().sort(), defaultModel: s.genie.defaultModel || '', + disabledModels: (s.genie.disabledModels || []).slice().sort(), customModels: (s.genie.customModels || []).map(cleanCustom), + defaultModel: s.genie.defaultModel || '', contextEditorChars: s.genie.contextEditorChars || 0, contextTerminalChars: s.genie.contextTerminalChars || 0, contextTerminalLines: s.genie.contextTerminalLines || 0, historyMessages: s.genie.historyMessages || 0, openRouterTimeoutSec: s.genie.openRouterTimeoutSec || 0, @@ -1006,6 +1026,78 @@ models.forEach(function (m) { defModel.appendChild(h('option', { value: m.id, text: m.name })); }); defModel.value = settings.genie.defaultModel || ''; form.defModel = defModel; + var wantedDefault = settings.genie.defaultModel || ''; + + // the models an admin added: one block each, with its own switch + form.custom = []; + var builtInIds = models.map(function (m) { return m.id; }); + var openRouterKeySet = models.some(function (m) { return m.provider === 'openrouter' && m.keySet; }); + var customList = h('div', { class: 'custom-models' }); + function addCustomRow(c) { + var r = { id: c.id }; + var on = toggle('f-cm-' + form.custom.length + '-' + Date.now(), !!c.enabled, (c.name || c.id) + ' available'); + r.on = on.input; + r.name = h('input', { type: 'text', class: 'field', maxlength: '40', 'aria-label': 'Name of ' + c.id, oninput: update }); + r.name.value = c.name && c.name !== c.id ? c.name : (c.name || ''); + r.free = h('input', { type: 'checkbox', checked: !!c.free || /:free$/.test(c.id), onchange: update, 'aria-label': c.id + ' is on a free tier' }); + r.json = h('input', { type: 'checkbox', checked: !!c.jsonMode, onchange: update, 'aria-label': c.id + ' takes JSON mode' }); + r.answer = h('input', { type: 'number', class: 'field', min: '500', max: '16000', step: '100', placeholder: '4000', 'aria-label': 'Answer size of ' + c.id, oninput: update }); + r.answer.value = c.answerTokens || ''; + r.room = h('input', { type: 'number', class: 'field', min: '0', max: '16000', step: '100', placeholder: '0', 'aria-label': 'Room for thinking of ' + c.id, oninput: update }); + r.room.value = c.thinkingRoom || ''; + r.hosts = h('input', { type: 'text', class: 'field', placeholder: 'OpenRouter chooses', 'aria-label': 'Providers allowed for ' + c.id, oninput: update }); + r.hosts.value = (c.hosts || []).join(', '); + var remove = h('button', { class: 'btn small', type: 'button', text: 'Remove', 'aria-label': 'Remove ' + c.id, onclick: function () { + form.custom = form.custom.filter(function (x) { return x !== r; }); + if (r.node.parentNode) r.node.parentNode.removeChild(r.node); + update(); + } }); + var field = function (label, input, hint) { return h('label', { class: 'cm-field' }, h('span', { text: label }), input, hint ? h('small', { text: hint }) : null); }; + var check = function (input, label) { return h('label', { class: 'cm-check' }, input, h('span', { text: label })); }; + r.node = h('div', { class: 'custom-model' }, + h('div', { class: 'cm-head' }, on.node, h('div', { class: 'cm-title' }, r.name, h('span', { class: 'mono', text: c.id })), remove), + h('div', { class: 'cm-grid' }, + field('Answer tokens', r.answer, 'Most tokens of an answer. Empty: 4000.'), + field('Room for thinking', r.room, 'Extra tokens for a model that thinks first.'), + field('Providers', r.hosts, 'Only these may answer (OpenRouter names, comma separated).')), + h('div', { class: 'cm-checks' }, check(r.free, 'Free tier (visitors are told)'), check(r.json, 'Its host takes JSON mode (Agent mode)'))); + form.custom.push(r); + customList.appendChild(r.node); + return r; + } + (settings.genie.customModels || []).forEach(addCustomRow); + var addId = h('input', { type: 'text', class: 'field', id: 'f-cm-new-id', placeholder: 'author/model-name:free', 'aria-label': 'OpenRouter model id', autocomplete: 'off' }); + var addName = h('input', { type: 'text', class: 'field', id: 'f-cm-new-name', maxlength: '40', placeholder: 'Name visitors see (optional)', 'aria-label': 'Name of the new model', autocomplete: 'off' }); + var addBtn = h('button', { class: 'btn', type: 'button', text: 'Add model', onclick: function () { + var id = addId.value.trim().toLowerCase(); + if (!CUSTOM_ID.test(id)) { toast('An OpenRouter id looks like author/model-name, for example nvidia/nemotron-3-super-120b-a12b:free.', 'error'); return; } + if (builtInIds.indexOf(id) >= 0 || form.custom.some(function (r) { return r.id === id; })) { toast(id + ' is in the list already.', 'error'); return; } + if (form.custom.length >= 20) { toast('At most 20 models can be added. Remove one first.', 'error'); return; } + var free = /:free$/.test(id); + addCustomRow({ id: id, name: addName.value.trim(), enabled: false, free: free, jsonMode: false, thinkingRoom: free ? 3000 : 0, hosts: [] }); + addId.value = ''; addName.value = ''; + update(); + } }); + addId.addEventListener('keydown', function (e) { if (e.key === 'Enter') { e.preventDefault(); addBtn.click(); } }); + + // the default model can be any model that is on, built in or added + function isModelOn(id) { + if (form.models[id]) return form.models[id].checked; + return form.custom.some(function (r) { return r.id === id && r.on.checked; }); + } + function syncDefaultModels() { + var keep = defModel.value || wantedDefault; + wantedDefault = ''; + Array.prototype.slice.call(defModel.options).forEach(function (o) { if (o.getAttribute('data-custom')) defModel.removeChild(o); }); + form.custom.forEach(function (r) { + var o = h('option', { value: r.id, text: r.name.value.trim() || r.id }); + o.setAttribute('data-custom', '1'); + defModel.appendChild(o); + }); + Array.prototype.forEach.call(defModel.options, function (o) { o.disabled = !!o.value && !isModelOn(o.value); }); + defModel.value = keep; + if (defModel.value !== keep || (keep && !isModelOn(keep))) defModel.value = ''; + } // the numbers: empty means the built-in value, shown as the placeholder function numberField(id, key, label, builtIn, range) { @@ -1071,8 +1163,7 @@ Array.prototype.forEach.call(form.host2.options, function (o) { o.disabled = !!o.value && o.value === form.host1.value; }); if (form.host2.value === form.host1.value) form.host2.value = ''; // a model that is off cannot be the default - Array.prototype.forEach.call(defModel.options, function (o) { o.disabled = !!o.value && !form.models[o.value].checked; }); - if (defModel.value && !form.models[defModel.value].checked) defModel.value = ''; + syncDefaultModels(); Array.prototype.forEach.call(langGrid.children, function (lab) { lab.className = form.langs[lab.getAttribute('data-lang')].checked ? '' : 'off'; }); @@ -1088,16 +1179,24 @@ if (s.genie.guestPerMinute < 0 || s.genie.guestPerMinute > 60 || s.genie.userPerMinute < 0 || s.genie.userPerMinute > 60) { toast('The Genie rates must be between 0 and 60 requests per minute.', 'error'); return; } - if (models.length && s.genie.disabledModels.length >= models.length) { + var customOn = s.genie.customModels.some(function (c) { return c.enabled; }); + if (models.length && s.genie.disabledModels.length >= models.length && !customOn) { toast('At least one model has to stay on.', 'error'); return; } var go = Promise.resolve(true); + var removed = (saved.genie.customModels || []).filter(function (c) { + return !s.genie.customModels.some(function (n) { return n.id === c.id; }); + }); + if (removed.length) { + go = ask({ title: 'Remove ' + removed.map(function (c) { return c.name || c.id; }).join(' and ') + '?', + body: 'It is taken out of the list. Visitors who chose it are asked to pick another model. You can add it again by its id.', confirm: 'Remove', danger: true }); + } var newlyOff = models.filter(function (m) { return s.genie.disabledModels.indexOf(m.id) >= 0 && (saved.genie.disabledModels || []).indexOf(m.id) < 0; }); if (newlyOff.length) { - go = ask({ title: 'Switch off ' + newlyOff.map(function (m) { return m.name; }).join(' and ') + '?', - body: 'Visitors who use it are asked to pick another model. A visitor with a page already open sees that on their next request.', confirm: 'Switch off', danger: true }); + go = go.then(function (ok) { return ok ? ask({ title: 'Switch off ' + newlyOff.map(function (m) { return m.name; }).join(' and ') + '?', + body: 'Visitors who use it are asked to pick another model. A visitor with a page already open sees that on their next request.', confirm: 'Switch off', danger: true }) : false; }); } if (s.maintenance.enabled && !saved.maintenance.enabled) { go = go.then(function (ok) { return ok ? ask({ title: 'Turn on maintenance mode?', body: 'Visitors will not be able to start terminals until you turn it off. Admins are not affected.', confirm: 'Turn on', danger: true }) : false; }); @@ -1240,6 +1339,11 @@ row('Default model', 'The model visitors start with, and the one that answers a request that names none. Built in is Luna for visitors and GPT-4o mini for older callers.', defModel), row('Guests', 'Requests per minute (0 or empty uses the built-in rate).', guest), row('Signed-in users', 'Requests per minute (0 or empty uses the built-in rate).', user) ]), + card('OpenRouter models', 'The three models above are built in and stay. Add any other model of OpenRouter by its id (find them at openrouter.ai/models), switch it on when you want visitors to have it, and remove it when you do not. A model that is off is not offered, and the server refuses it. Save to apply.', [ + openRouterKeySet ? null : h('p', { class: 'callout', text: 'No OpenRouter key is set on the server, so these models cannot answer yet. Add one under API keys below.' }), + h('p', { class: 'callout info', text: 'A model whose id ends in :free costs nothing, but its provider may keep and use what visitors send (their code and terminal output) and its limits are shared by everyone, so it can be busy. Visitors see it labelled free. Any other model is paid for from your OpenRouter credit. Name the providers that may answer if you want to know where the code goes; empty lets OpenRouter choose. The privacy page names only the hosts of Gemma and says other models are chosen by the site owner.' }), + customList, + h('div', { class: 'cm-add' }, addId, addName, addBtn) ]), card('Genie limits', 'What Genie is told about the page and how long its answers may be. An empty box means the built-in value. Changes reach a page the next time it loads.', [ row('Editor code', 'Characters of the editor\'s code sent with each question.', form.ctxEditor), row('Terminal output, characters', 'The most recent characters of the active terminal.', form.ctxTermChars), diff --git a/src/resources/js/model-choice.js b/src/resources/js/model-choice.js index 18a221d0..39e2de83 100644 --- a/src/resources/js/model-choice.js +++ b/src/resources/js/model-choice.js @@ -1,9 +1,9 @@ // The model and effort choice, shared by Genie (the chip in its panel, chat // widget) and the New question dialog (js/common.js). Loaded before both. // -// Three models are offered, and only these: server/chatmodels.go holds the same -// lists and sends nothing else on (a request for any other model is answered by -// GPT-4o mini). Two come from OpenAI; Gemma 4 31B comes through OpenRouter and +// The built-in models and the ones an admin added and switched on are offered, and +// only these: server/chatmodels.go and custom_models.go hold the same lists and +// send nothing else on (a request for any other model is answered by the default). Two come from OpenAI; Gemma 4 31B comes through OpenRouter and // is listed only when the server has an OpenRouter key (config.js sets // openrouter_enabled). The choice is kept on this device, and the default is // Luna with a Low effort. A page that wants a choice of its own (the blog @@ -45,6 +45,25 @@ }); } + // Models an admin added in the dashboard and switched on come with the page + // (settings.js: site_settings.customModels, the ones that are on). They all run + // through OpenRouter, and a free one says so. + if (window.openrouter_enabled === true) { + ((window.site_settings || {}).customModels || []).forEach(function (c) { + ALL.push({ + id: c.id, + name: c.name, + short: c.name.length > 16 ? c.name.slice(0, 15) + "…" : c.name, + group: "OpenRouter", + tag: c.free ? "Free" : "", + desc: c.free + ? "A free model, run through OpenRouter. Its provider may keep what you send, and it can be busy." + : "Run through OpenRouter.", + reasoning: false, + }); + }); + } + // An admin can switch models off and choose the one visitors start with // (settings.js: site_settings.disabledModels, defaultModel). The server refuses // a model that is off, so the page does not offer it. If that would leave diff --git a/src/resources/knowledge/genie.md b/src/resources/knowledge/genie.md index df5a0c78..23a13f76 100644 --- a/src/resources/knowledge/genie.md +++ b/src/resources/knowledge/genie.md @@ -3,4 +3,4 @@ title: Genie, the AI helper keywords: genie, ai, assistant, chat, help, ask, explain, fix, suggest, how to use genie, use genie, what can genie do, getting started with genie, model, models, effort, thinking, luna, gpt, gemma, openai, openrouter, ctrl k, requests link: /privacy.html --- -Genie is OpenREPL's AI helper. Open it with the Ask Genie button or Ctrl+K and type a question. It reads your editor code and the latest terminal output, so you can ask why something fails without pasting anything. Choose the model with the chip beside Send: GPT-6 Luna, which thinks first and has an effort setting, GPT-4o mini, which is quick, or Gemma 4 31B through OpenRouter. Each message uses one request; signed-in users get more. Snippets in answers have Insert and Replace buttons. Right-click selected text to explain it, fix it, comment it or write tests. Signed-in users can switch from Chat to Agent mode. Genie is optional. Do not paste passwords or keys into it. +Genie is OpenREPL's AI helper. Open it with the Ask Genie button or Ctrl+K and type a question. It reads your editor code and the latest terminal output, so you can ask why something fails without pasting anything. Choose the model with the chip beside Send: GPT-6 Luna, which thinks first and has an effort setting, GPT-4o mini, which is quick, or Gemma 4 31B through OpenRouter. Other models, some of them free, can appear in the list when the site owner adds them. Each message uses one request; signed-in users get more. Snippets in answers have Insert and Replace buttons. Right-click selected text to explain it, fix it, comment it or write tests. Signed-in users can switch from Chat to Agent mode. Genie is optional. Do not paste passwords or keys into it. diff --git a/src/resources/privacy.html b/src/resources/privacy.html index e2582f8b..164ecc95 100644 --- a/src/resources/privacy.html +++ b/src/resources/privacy.html @@ -56,7 +56,7 @@

Code links

When you choose Share, then Create code link, a copy of the code in your editor (up to 64 KB) is stored on our server and you get a short link. Anyone who has the link can read that code. Code links are kept until we remove them, so don't create one for code that contains secrets. To prevent abuse we count how many links each IP address creates in a short time.

Genie, the AI helper

-

When you use Genie, your messages, the code in your editor and the most recent output of your terminal are sent through our server to OpenAI to generate replies. You can choose which model answers: GPT-6 Luna or GPT-4o mini (both run by OpenAI), or Gemma 4 31B, which runs through OpenRouter on servers operated by ModelRun and, if that is unavailable, CoreWeave. If you choose Gemma, your messages go to OpenRouter and to whichever of those two answers, not to OpenAI. Genie is optional; if you don't open it and don't ask for a practice question, nothing is sent to OpenAI or OpenRouter.

+

When you use Genie, your messages, the code in your editor and the most recent output of your terminal are sent through our server to OpenAI to generate replies. You can choose which model answers: GPT-6 Luna or GPT-4o mini (both run by OpenAI), or Gemma 4 31B, which runs through OpenRouter on servers operated by ModelRun and, if that is unavailable, CoreWeave. If you choose Gemma, your messages go to OpenRouter and to whichever of those two answers, not to OpenAI. The site owner can also add other models that run through OpenRouter. If you choose one, your messages go to OpenRouter and to the provider that runs it, which the owner names or OpenRouter picks, not to OpenAI. A model marked free runs on a provider's free tier, and such a provider may keep what you send and use it under its own terms, so do not send anything private to a free model. Genie is optional; if you don't open it and don't ask for a practice question, nothing is sent to OpenAI or OpenRouter.

To answer questions about OpenREPL itself, Genie may add short passages from our own notes about the service, and from blog posts published on this site, to your message before it is sent to the model. This text is public. It is picked on our server by matching words in your question, with no extra service involved, and nothing about you is added to it. If nothing matches, nothing is added. Genie shows under its reply which notes were used. The Ask Genie button of the blog editor works the same way for the people who write posts. Practice questions do not use it.

Agent mode, when you are signed in and choose it in the Genie panel, lets Genie work on a task in several steps. Each step sends the same things as a Genie message (your task, the code in your editor and the recent output of your terminal) to the model you chose. Genie changes your editor, switches the language, runs your code, restarts your terminal, opens, switches or closes terminal tabs, types a line in your terminal, or works in your Files panel only after you allowed that kind of action (the names of your files are sent to the model when it lists them; renaming, moving and deleting are asked about every time), and you review every change to the editor before it is applied. You see the exact line before it is typed, and a line that could delete or change things is asked about every time. It stops when you press Stop or close the panel.

The right-click actions on selected code (Fix problems, Add comments, Write tests) send the selected code, the file it is in, the language and the recent terminal output to the model you chose; Explain is sent as a Genie message. On the practice page the coach (hints, a solution review and complexity) sends the editor, which holds the problem and your code, and the recent terminal output in the same way. Hints used are counted in your browser only.

diff --git a/src/server/chatmodels.go b/src/server/chatmodels.go index 5d52b227..3ea0c21b 100644 --- a/src/server/chatmodels.go +++ b/src/server/chatmodels.go @@ -44,6 +44,17 @@ type chatModel struct { Provider string Reasoning bool MaxTokensCap int + // The rest is for models an admin added (custom_models.go). + Custom bool + // Free: on a provider's free tier. + Free bool + // Hosts are the only OpenRouter providers that may answer; empty means + // OpenRouter chooses (a custom model) or Gemma's hosts (Gemma). + Hosts []string + // ThinkingRoom is added to the answer budget of a model that thinks first. + ThinkingRoom int + // NoJSONMode: the model's host takes no response_format. + NoJSONMode bool } // the order the models are listed in, in the dashboard and as a last resort @@ -64,7 +75,7 @@ func defaultModelID() string { if g.DefaultModel != "" && !g.ModelDisabled(g.DefaultModel) { return g.DefaultModel } - for _, id := range append([]string{defaultChatModel}, chatModelOrder...) { + for _, id := range append([]string{defaultChatModel}, g.allModelIDs()...) { if !g.ModelDisabled(id) { return id } @@ -91,8 +102,17 @@ func activeOpenRouterHosts() []string { return openRouterHosts } -func openRouterRouting() map[string]interface{} { +// openRouterRouting is the host rule of a request: the hosts the model names, the +// admin's choice for Gemma, or none for a model an admin added without hosts +// (nil: OpenRouter chooses). +func openRouterRouting(model chatModel) map[string]interface{} { hosts := activeOpenRouterHosts() + switch { + case len(model.Hosts) > 0: + hosts = model.Hosts + case model.Custom: + return nil + } return map[string]interface{}{ "order": hosts, "only": hosts, @@ -174,28 +194,30 @@ func sanitizeChatBody(body []byte) ([]byte, chatModel, error) { } } - // The question generator asks for JSON so that a quote or a newline inside - // the text cannot make the reply unparseable. No other format is passed on. - if raw, ok := in["response_format"]; ok { - var format struct { - Type string `json:"type"` - } - if json.Unmarshal(raw, &format) == nil && format.Type == "json_object" { - out["response_format"] = map[string]string{"type": "json_object"} - } - } - - model := chatModels[defaultModelID()] + model, _ := modelByID(defaultModelID()) if raw, ok := in["model"]; ok { var asked string if json.Unmarshal(raw, &asked) == nil { - if known, ok := chatModels[asked]; ok { + if known, ok := modelByID(asked); ok { model = known } } } out["model"] = model.ID + // The question generator asks for JSON so that a quote or a newline inside + // the text cannot make the reply unparseable. No other format is passed on, + // and none to a model whose host does not take it (the agent reads the answer + // as JSON all the same, and asks again when it cannot). + if raw, ok := in["response_format"]; ok && !model.NoJSONMode { + var format struct { + Type string `json:"type"` + } + if json.Unmarshal(raw, &format) == nil && format.Type == "json_object" { + out["response_format"] = map[string]string{"type": "json_object"} + } + } + if model.Reasoning { effort := defaultEffort if raw, ok := in["reasoning_effort"]; ok { @@ -209,7 +231,7 @@ func sanitizeChatBody(body []byte) ([]byte, chatModel, error) { out["reasoning_effort"] = effort out["max_completion_tokens"] = clampTokens(in["max_completion_tokens"], effortTokens[effort], minCompletionTokens, answerCap(model)) } else { - out["max_tokens"] = clampTokens(in["max_tokens"], defaultMaxTokens, minCompletionTokens, answerCap(model)) + out["max_tokens"] = clampTokens(in["max_tokens"], defaultMaxTokens, minCompletionTokens, answerCap(model)) + model.ThinkingRoom if raw, ok := in["temperature"]; ok { var t float64 if json.Unmarshal(raw, &t) == nil && !math.IsNaN(t) { @@ -218,7 +240,9 @@ func sanitizeChatBody(body []byte) ([]byte, chatModel, error) { } } if model.Provider == providerOpenRouter { - out["provider"] = openRouterRouting() + if rule := openRouterRouting(model); rule != nil { + out["provider"] = rule + } } sanitized, err := json.Marshal(out) return sanitized, model, err diff --git a/src/server/custom_models.go b/src/server/custom_models.go new file mode 100644 index 00000000..6a265472 --- /dev/null +++ b/src/server/custom_models.go @@ -0,0 +1,241 @@ +package server + +import ( + "fmt" + "regexp" + "strings" + "unicode" + "unicode/utf8" +) + +// Models an admin adds in the dashboard (Settings, Genie, "OpenRouter models"). +// +// The three built-in models (GPT-6 Luna, GPT-4o mini, Gemma 4 31B) are fixed: +// an admin can switch them off, not remove them. Any other model of OpenRouter +// is added to this list by its OpenRouter id, with a switch, and removed again, +// without a change of code. Visitors can only pick what is in the list and on: +// the server sends on nothing else, whatever the browser asks for. + +// the two models of defaultCustomModels +const ( + modelNemotronFree = "nvidia/nemotron-3-super-120b-a12b:free" + modelNorthFree = "cohere/north-mini-code:free" +) + +const ( + maxCustomModels = 20 + maxCustomNameChars = 40 + maxCustomHosts = 4 + defaultCustomAnswer = 4000 // the answer cap of a model that has none of its own + maxCustomThinkingRoom = 16000 // extra tokens for a model that thinks first + customModelIDMaxLength = 160 +) + +// the id OpenRouter gives a model: author/model, optionally with a variant +// (":free"). Lower case, so that two spellings are not two models. +var customModelIDPattern = regexp.MustCompile(`^[a-z0-9][a-z0-9._-]{0,63}/[a-z0-9][a-z0-9._:+-]{0,95}$`) + +// an OpenRouter provider slug, as in "nvidia" or "modelrun/fp4" +var hostSlugPattern = regexp.MustCompile(`^[a-z0-9][a-z0-9._/-]{0,39}$`) + +// CustomModel is an OpenRouter model an admin added. +type CustomModel struct { + // ID is OpenRouter's id of the model, for example "nvidia/nemotron-3-super-120b-a12b:free". + ID string `json:"id"` + // Name is what visitors see. + Name string `json:"name"` + // Enabled is the switch: a model that is off is not offered and the proxy + // refuses it. + Enabled bool `json:"enabled"` + // Free marks a model on a provider's free tier. It costs nothing, but its + // provider may keep and use what is sent, and its limits are shared; visitors + // are told so. An id that ends in ":free" is always marked. + Free bool `json:"free"` + // AnswerTokens is the most tokens of an answer (0: 4000). + AnswerTokens int `json:"answerTokens"` + // ThinkingRoom is added to that for a model that thinks before it answers: the + // thinking counts against the limit, and the answer would be cut off without + // room for it (0: none). + ThinkingRoom int `json:"thinkingRoom"` + // JSONMode says the model's host takes response_format json_object, which + // Agent mode asks for. Without it the agent asks for JSON in words, and asks + // again when it cannot read the answer. + JSONMode bool `json:"jsonMode"` + // Hosts are the OpenRouter providers (slugs) that may answer, and no others. + // Empty: OpenRouter chooses, which the privacy page cannot name. + Hosts []string `json:"hosts"` +} + +// defaultCustomModels are in the list until an admin saves one of their own: the +// two free models that were tried and answered, off. +func defaultCustomModels() []CustomModel { + return []CustomModel{ + {ID: modelNemotronFree, Name: "Nemotron 3 Super (free)", Free: true, + AnswerTokens: 4000, ThinkingRoom: 3000, JSONMode: true, Hosts: []string{"nvidia"}}, + {ID: modelNorthFree, Name: "North mini code (free)", Free: true, + AnswerTokens: 4000, ThinkingRoom: 3000, JSONMode: false, Hosts: []string{"cohere"}}, + } +} + +// customList is the list in effect: the admin's, or the defaults while nothing +// was ever saved (nil; an admin who removed them all has an empty list). +func (g GenieSettings) customList() []CustomModel { + if g.CustomModels != nil { + return g.CustomModels + } + return defaultCustomModels() +} + +// model is the entry of the proxy's table for it. +func (c CustomModel) model() chatModel { + answer := c.AnswerTokens + if answer <= 0 { + answer = defaultCustomAnswer + } + return chatModel{ + ID: c.ID, Name: c.Name, Provider: providerOpenRouter, MaxTokensCap: answer, + Free: c.Free || strings.HasSuffix(c.ID, ":free"), Hosts: append([]string(nil), c.Hosts...), + ThinkingRoom: c.ThinkingRoom, NoJSONMode: !c.JSONMode, Custom: true, + } +} + +// modelByID finds a built-in model or one an admin added (on or off). +func (g GenieSettings) modelByID(id string) (chatModel, bool) { + if m, ok := chatModels[id]; ok { + return m, true + } + for _, c := range g.customList() { + if c.ID == id { + return c.model(), true + } + } + return chatModel{}, false +} + +func modelByID(id string) (chatModel, bool) { return GetSiteSettings().Genie.modelByID(id) } + +// allModelIDs are the built-in models, then the ones an admin added. +func (g GenieSettings) allModelIDs() []string { + ids := append([]string(nil), chatModelOrder...) + for _, c := range g.customList() { + ids = append(ids, c.ID) + } + return ids +} + +// normalizeCustomModels cleans the list of an untrusted source and says what is +// wrong with it. nil stays nil: "never configured". +func normalizeCustomModels(in []CustomModel) ([]CustomModel, error) { + if in == nil { + return nil, nil + } + if len(in) > maxCustomModels { + return nil, fmt.Errorf("at most %d OpenRouter models can be added", maxCustomModels) + } + out := []CustomModel{} + seen := map[string]bool{} + for _, c := range in { + c.ID = strings.ToLower(strings.TrimSpace(c.ID)) + if len(c.ID) > customModelIDMaxLength || !customModelIDPattern.MatchString(c.ID) { + return nil, fmt.Errorf("%q is not an OpenRouter model id (it looks like author/model-name)", c.ID) + } + if _, builtin := chatModels[c.ID]; builtin { + return nil, fmt.Errorf("%s is one of the built-in models; they cannot be added or removed", c.ID) + } + if seen[c.ID] { + return nil, fmt.Errorf("%s was added twice", c.ID) + } + seen[c.ID] = true + + c.Name = strings.TrimSpace(c.Name) + if c.Name == "" { + c.Name = c.ID + } + if utf8.RuneCountInString(c.Name) > maxCustomNameChars { + return nil, fmt.Errorf("the name of %s is longer than %d characters", c.ID, maxCustomNameChars) + } + for _, r := range c.Name { + if unicode.IsControl(r) { + return nil, fmt.Errorf("the name of %s has a control character", c.ID) + } + } + if c.AnswerTokens != 0 && (c.AnswerTokens < minAnswerCap || c.AnswerTokens > maxAnswerCap) { + return nil, fmt.Errorf("the answer size of %s must be between %d and %d tokens (or empty)", c.ID, minAnswerCap, maxAnswerCap) + } + if c.ThinkingRoom < 0 || c.ThinkingRoom > maxCustomThinkingRoom { + return nil, fmt.Errorf("the room for thinking of %s must be between 0 and %d tokens", c.ID, maxCustomThinkingRoom) + } + if strings.HasSuffix(c.ID, ":free") { + c.Free = true + } + hosts := []string{} + seenHost := map[string]bool{} + for _, h := range c.Hosts { + h = strings.ToLower(strings.TrimSpace(h)) + if h == "" || seenHost[h] { + continue + } + if !hostSlugPattern.MatchString(h) { + return nil, fmt.Errorf("%q is not a provider name of %s", h, c.ID) + } + seenHost[h] = true + hosts = append(hosts, h) + } + if len(hosts) > maxCustomHosts { + return nil, fmt.Errorf("at most %d providers can be named for %s", maxCustomHosts, c.ID) + } + c.Hosts = hosts + out = append(out, c) + } + return out, nil +} + +// publicCustomModel is what a page needs of a model that is on. +type publicCustomModel struct { + ID string `json:"id"` + Name string `json:"name"` + Free bool `json:"free"` +} + +// publicCustomModels lists the models that are on, for the picker. +func (g GenieSettings) publicCustomModels() []publicCustomModel { + out := []publicCustomModel{} + for _, c := range g.customList() { + if c.Enabled { + out = append(out, publicCustomModel{ID: c.ID, Name: c.Name, Free: c.model().Free}) + } + } + return out +} + +// customModelChanges describes what an admin did to the list, for the audit log. +func customModelChanges(a, b GenieSettings) []string { + before := map[string]CustomModel{} + for _, c := range a.customList() { + before[c.ID] = c + } + var out []string + after := map[string]bool{} + for _, c := range b.customList() { + after[c.ID] = true + old, existed := before[c.ID] + switch { + case !existed: + state := "off" + if c.Enabled { + state = "on" + } + out = append(out, "OpenRouter model "+c.ID+" added ("+state+")") + case old.Enabled != c.Enabled: + out = append(out, "OpenRouter model "+c.ID+" "+map[bool]string{true: "switched on", false: "switched off"}[c.Enabled]) + case fmt.Sprint(old) != fmt.Sprint(c): + out = append(out, "OpenRouter model "+c.ID+" changed (name, limits or providers)") + } + } + for _, c := range a.customList() { + if !after[c.ID] { + out = append(out, "OpenRouter model "+c.ID+" removed") + } + } + return out +} diff --git a/src/server/custom_models_test.go b/src/server/custom_models_test.go new file mode 100644 index 00000000..dbb64b89 --- /dev/null +++ b/src/server/custom_models_test.go @@ -0,0 +1,265 @@ +package server + +import ( + "encoding/json" + "reflect" + "strings" + "testing" +) + +// sanitizedAsIs is sanitized without resetting the settings the test saved. +func sanitizedAsIs(t *testing.T, body string) map[string]interface{} { + t.Helper() + out, _, err := sanitizeChatBody([]byte(body)) + if err != nil { + t.Fatalf("sanitizeChatBody(%s): %v", body, err) + } + var got map[string]interface{} + if err := json.Unmarshal(out, &got); err != nil { + t.Fatalf("result is not JSON: %v", err) + } + return got +} + +func customOn(t *testing.T, id string) GenieSettings { + t.Helper() + g := GetSiteSettings().Genie + g.CustomModels = append([]CustomModel{}, g.customList()...) + for i := range g.CustomModels { + if g.CustomModels[i].ID == id { + g.CustomModels[i].Enabled = true + } + } + return g +} + +func TestModelsAnAdminAddsAreOffUntilSwitchedOn(t *testing.T) { + isolateSettings(t) + g := GetSiteSettings().Genie + // nothing was ever saved: the two free models are listed, off, and hidden from the page + if len(g.customList()) != 2 || !g.ModelDisabled(modelNemotronFree) || !g.ModelDisabled(modelNorthFree) { + t.Fatalf("defaults: %+v", g.customList()) + } + if pub := GetSiteSettings().public(); len(pub.CustomModels) != 0 { + t.Fatalf("the page is offered models that are off: %+v", pub.CustomModels) + } + // the built-in three are untouched + for _, id := range chatModelOrder { + if g.ModelDisabled(id) { + t.Errorf("%s is off", id) + } + } + + s := GetSiteSettings() + s.Genie = customOn(t, modelNemotronFree) + if err := SaveSiteSettings(s); err != nil { + t.Fatal(err) + } + g = GetSiteSettings().Genie + pub := GetSiteSettings().public().CustomModels + if g.ModelDisabled(modelNemotronFree) || !g.ModelDisabled(modelNorthFree) || len(pub) != 1 || pub[0].ID != modelNemotronFree || !pub[0].Free { + t.Fatalf("after switching one on: %+v / %+v", g.customList(), pub) + } + // a request for it is now for it, and for the one that is still off the proxy refuses (chatproxy.go) + if m, ok := modelByID(modelNemotronFree); !ok || GetSiteSettings().Genie.ModelDisabled(m.ID) { + t.Fatalf("the model that was switched on is off: %+v %v", m, ok) + } + got2 := sanitizedAsIs(t, `{"model":"`+modelNemotronFree+`","messages":[{"role":"user","content":"hi"}]}`) + if got2["model"] != modelNemotronFree { + t.Fatalf("the request went to %v", got2["model"]) + } +} + +func TestModelsCanBeAddedAndRemovedAndTheBuiltInOnesCannot(t *testing.T) { + isolateSettings(t) + s := GetSiteSettings() + s.Genie.CustomModels = []CustomModel{{ID: " Qwen/Qwen3-Coder:free ", Name: " Qwen coder ", Enabled: true, ThinkingRoom: 2000, JSONMode: true, Hosts: []string{" Together ", "together"}}} + if err := SaveSiteSettings(s); err != nil { + t.Fatal(err) + } + c := GetSiteSettings().Genie.CustomModels + if len(c) != 1 || c[0].ID != "qwen/qwen3-coder:free" || c[0].Name != "Qwen coder" || !c[0].Free || !reflect.DeepEqual(c[0].Hosts, []string{"together"}) { + t.Fatalf("cleaned: %+v", c) + } + // the added model is a model of the proxy, the default ones are gone with it + if m, ok := modelByID("qwen/qwen3-coder:free"); !ok || !m.Custom || m.Provider != providerOpenRouter { + t.Fatalf("lookup: %+v %v", m, ok) + } + if _, ok := modelByID(modelNemotronFree); ok { + t.Fatal("a removed model is still known") + } + // removing all of them is a choice too: it does not bring the defaults back + s = GetSiteSettings() + s.Genie.CustomModels = []CustomModel{} + if err := SaveSiteSettings(s); err != nil { + t.Fatal(err) + } + if got := GetSiteSettings().Genie.customList(); len(got) != 0 { + t.Fatalf("after removing all: %+v", got) + } + + for name, list := range map[string][]CustomModel{ + "a built-in model": {{ID: modelGemma}}, + "a bad id": {{ID: "not an id"}}, + "no author": {{ID: "justaname"}}, + "the same id twice": {{ID: "a/b"}, {ID: "A/B"}}, + "a bad provider": {{ID: "a/b", Hosts: []string{"x y"}}}, + "too long a name": {{ID: "a/b", Name: strings.Repeat("n", 41)}}, + "a control character": {{ID: "a/b", Name: "bad\x07name"}}, + "a silly answer size": {{ID: "a/b", AnswerTokens: 5}}, + "a silly thinking room": {{ID: "a/b", ThinkingRoom: 999999}}, + "too many providers": {{ID: "a/b", Hosts: []string{"a", "b", "c", "d", "e"}}}, + } { + s := GetSiteSettings() + s.Genie.CustomModels = list + if err := SaveSiteSettings(s); err == nil { + t.Errorf("%s was accepted", name) + } + } + many := []CustomModel{} + for i := 0; i <= maxCustomModels; i++ { + many = append(many, CustomModel{ID: "a/m" + strings.Repeat("x", i)}) + } + s = GetSiteSettings() + s.Genie.CustomModels = many + if err := SaveSiteSettings(s); err == nil { + t.Error("more models than the limit were accepted") + } +} + +func TestAnAddedModelGoesWhereItsSettingsSay(t *testing.T) { + isolateSettings(t) + s := GetSiteSettings() + s.Genie.CustomModels = []CustomModel{ + {ID: "vendor/coder:free", Name: "Coder", Enabled: true, AnswerTokens: 2000, ThinkingRoom: 1500, JSONMode: false, Hosts: []string{"vendor"}}, + {ID: "other/open-model", Name: "Open", Enabled: true, JSONMode: true}, + } + if err := SaveSiteSettings(s); err != nil { + t.Fatal(err) + } + got := sanitizedAsIs(t, `{"model":"vendor/coder:free","max_tokens":99999,"temperature":0.5,"response_format":{"type":"json_object"}, + "reasoning_effort":"high","provider":{"only":["evil"]},"messages":[{"role":"user","content":"hi"}]}`) + if got["model"] != "vendor/coder:free" || got["max_tokens"] != float64(3500) || got["temperature"] != 0.5 { + t.Errorf("body: %v", got) + } + if _, has := got["response_format"]; has { + t.Errorf("JSON mode went to a host that does not take it: %v", got) + } + if _, has := got["reasoning_effort"]; has { + t.Errorf("reasoning_effort reached a custom model: %v", got) + } + want := map[string]interface{}{"order": []interface{}{"vendor"}, "only": []interface{}{"vendor"}, "allow_fallbacks": false} + if !reflect.DeepEqual(got["provider"], want) { + t.Errorf("host rule from the browser got through, or is wrong: %v", got["provider"]) + } + + // no providers named: OpenRouter chooses, and the browser cannot say otherwise + got = sanitizedAsIs(t, `{"model":"other/open-model","provider":{"only":["evil"]},"response_format":{"type":"json_object"},"messages":[{"role":"user","content":"hi"}]}`) + if _, has := got["provider"]; has { + t.Errorf("a provider rule without hosts: %v", got["provider"]) + } + if got["max_tokens"] != float64(800) || !reflect.DeepEqual(got["response_format"], map[string]interface{}{"type": "json_object"}) { + t.Errorf("open model: %v", got) + } + // the host rule of Gemma is its own and an admin's choice of hosts stays Gemma's + got = sanitizedAsIs(t, `{"model":"google/gemma-4-31b-it","messages":[{"role":"user","content":"hi"}]}`) + if !reflect.DeepEqual(got["provider"].(map[string]interface{})["only"], []interface{}{"modelrun/fp4", "coreweave/fp4"}) { + t.Errorf("Gemma hosts: %v", got["provider"]) + } +} + +func TestTheSettingsPageKeepsTheModelListItDoesNotSend(t *testing.T) { + s, mux := adminTestServer(t) + _ = s + st := GetSiteSettings() + st.Genie.CustomModels = []CustomModel{{ID: "a/kept", Name: "Kept", Enabled: true}} + if err := SaveSiteSettings(st); err != nil { + t.Fatal(err) + } + // a dashboard page from before the list existed sends none + if w := post(mux, "/admin/settings", `{"colorOfTheDay":true,"genie":{"guestPerMinute":1}}`); w.Code != 200 { + t.Fatalf("%d %s", w.Code, w.Body.String()) + } + if got := GetSiteSettings().Genie.CustomModels; len(got) != 1 || got[0].ID != "a/kept" { + t.Fatalf("the list was lost: %+v", got) + } + // the dashboard sends the list it shows, and what it sends is what is kept + var r settingsReply + decode(t, get(mux, "/admin/settings"), &r) + if len(r.Genie.CustomModels) != 1 || r.Genie.CustomModels[0].ID != "a/kept" { + t.Fatalf("the dashboard is shown %+v", r.Genie.CustomModels) + } + w := post(mux, "/admin/settings", `{"genie":{"customModels":[{"id":"b/new","name":"New","enabled":true,"jsonMode":true}]}}`) + if w.Code != 200 { + t.Fatalf("%d %s", w.Code, w.Body.String()) + } + if got := GetSiteSettings().Genie.CustomModels; len(got) != 1 || got[0].ID != "b/new" || !got[0].Enabled { + t.Fatalf("saved: %+v", got) + } + // the audit log says what was added and removed + audit := get(mux, "/admin/audit").Body.String() + for _, want := range []string{"OpenRouter model b/new added (on)", "OpenRouter model a/kept removed"} { + if !strings.Contains(audit, want) { + t.Errorf("audit log lacks %q: %s", want, audit) + } + } + // a mistake is refused with a message, and changes nothing + w = post(mux, "/admin/settings", `{"genie":{"customModels":[{"id":"google/gemma-4-31b-it"}]}}`) + if w.Code != 400 || !strings.Contains(w.Body.String(), "built-in") { + t.Fatalf("a built-in id: %d %s", w.Code, w.Body.String()) + } + if got := GetSiteSettings().Genie.CustomModels; len(got) != 1 || got[0].ID != "b/new" { + t.Fatalf("a refused save changed the list: %+v", got) + } +} + +// ---- persistence -------------------------------------------------------------------- + +func TestTheModelListSurvivesARestartOfTheFileStore(t *testing.T) { + isolateSettings(t) + s := GetSiteSettings() + s.Genie.CustomModels = []CustomModel{{ID: "a/b", Name: "AB", Enabled: true, Free: true, AnswerTokens: 3000, ThinkingRoom: 500, JSONMode: true, Hosts: []string{"h"}}} + if err := SaveSiteSettings(s); err != nil { + t.Fatal(err) + } + want := GetSiteSettings().Genie.CustomModels + // a restart: nothing in memory, the file is read again + settingsMu.Lock() + siteSettings, settingsLoaded = SiteSettings{}, false + settingsMu.Unlock() + if got := GetSiteSettings().Genie.CustomModels; !reflect.DeepEqual(got, want) { + t.Fatalf("after a restart: %+v, want %+v", got, want) + } +} + +func TestTheModelListIsKeptInTheDatabaseAndAnotherServerSeesIt(t *testing.T) { + f := &fakeStore{} + withStore(t, f) + syncSettingsOnce() + s := GetSiteSettings() + s.Genie.CustomModels = []CustomModel{{ID: "a/b", Name: "AB", Enabled: true, JSONMode: true}, {ID: "c/d:free", Name: "CD"}} + if err := SaveSiteSettings(s); err != nil { + t.Fatal(err) + } + var stored struct { + Genie struct { + CustomModels []CustomModel `json:"customModels"` + } `json:"genie"` + } + if err := json.Unmarshal(f.data, &stored); err != nil || len(stored.Genie.CustomModels) != 2 || stored.Genie.CustomModels[0].ID != "a/b" || !stored.Genie.CustomModels[1].Free { + t.Fatalf("the database holds %s (%v)", f.data, err) + } + + // another server (or a restart) starts with nothing and reads the database + settingsMu.Lock() + siteSettings, settingsLoaded, settingsSynced = SiteSettings{}, false, false + settingsMu.Unlock() + syncSettingsOnce() + got := GetSiteSettings().Genie + if len(got.CustomModels) != 2 || got.ModelDisabled("a/b") || !got.ModelDisabled("c/d:free") { + t.Fatalf("the other server: %+v", got.CustomModels) + } + if pub := GetSiteSettings().public().CustomModels; len(pub) != 1 || pub[0].ID != "a/b" { + t.Fatalf("what its page is told: %+v", pub) + } +} diff --git a/src/server/settings.go b/src/server/settings.go index ae195442..0096e0d4 100644 --- a/src/server/settings.go +++ b/src/server/settings.go @@ -85,6 +85,10 @@ type GenieSettings struct { // picker hides them and the proxy answers 503 model_unavailable. At least // one model stays on. DisabledModels []string `json:"disabledModels"` + // CustomModels are the OpenRouter models an admin added, each with its own + // switch (custom_models.go). nil means none was ever saved: the defaults are + // listed, off. The three built-in models are not in it and cannot be removed. + CustomModels []CustomModel `json:"customModels"` // What Genie is told about the page, and how much of the conversation it // keeps: the editor's code (characters), the terminal's recent output // (characters and lines) and the number of messages. 0 is the built-in value. @@ -117,8 +121,17 @@ type GenieSettings struct { AgentMaxSteps int `json:"agentMaxSteps"` } -// ModelDisabled reports whether an admin switched the model off. +// ModelDisabled reports whether an admin switched the model off (a model an admin +// added is off until its switch is on). func (g GenieSettings) ModelDisabled(id string) bool { + if _, builtin := chatModels[id]; !builtin { + for _, c := range g.customList() { + if c.ID == id { + return !c.Enabled + } + } + return false + } for _, d := range g.DisabledModels { if d == id { return true @@ -216,17 +229,29 @@ func (s SiteSettings) normalize() (SiteSettings, error) { off = append(off, id) } sort.Strings(off) - if len(off) >= len(chatModels) { + s.Genie.DisabledModels = off + custom, err := normalizeCustomModels(s.Genie.CustomModels) + if err != nil { + return s, err + } + s.Genie.CustomModels = custom + anyOn := false + for _, id := range s.Genie.allModelIDs() { + if !s.Genie.ModelDisabled(id) { + anyOn = true + } + } + if !anyOn { return s, fmt.Errorf("at least one model has to stay switched on") } - s.Genie.DisabledModels = off s.Genie.DefaultModel = strings.TrimSpace(s.Genie.DefaultModel) if d := s.Genie.DefaultModel; d != "" { - if _, known := chatModels[d]; !known { + m, known := s.Genie.modelByID(d) + if !known { return s, fmt.Errorf("%q is not a model", d) } - if seenModel[d] { - return s, fmt.Errorf("the default model %s is switched off", chatModels[d].Name) + if s.Genie.ModelDisabled(d) { + return s, fmt.Errorf("the default model %s is switched off", m.Name) } } return s, nil @@ -335,7 +360,9 @@ type publicSettings struct { // the models that are switched off, and the one visitors start with (empty: // the built-in choice); model-choice.js reads them DisabledModels []string `json:"disabledModels"` - DefaultModel string `json:"defaultModel"` + // the models an admin added that are on, for the picker + CustomModels []publicCustomModel `json:"customModels"` + DefaultModel string `json:"defaultModel"` // what Genie reads from the page and keeps (the effective numbers) GenieContext genieContext `json:"genieContext"` // agent mode: whether the panel offers it (to signed-in users; the server @@ -452,6 +479,7 @@ func (s SiteSettings) public() publicSettings { DisabledLanguages: langs, GenieDisabled: s.Genie.Disabled, DisabledModels: models, + CustomModels: s.Genie.publicCustomModels(), DefaultModel: s.Genie.DefaultModel, GenieContext: s.Genie.context(), AgentEnabled: !s.Genie.AgentDisabled && !s.Genie.Disabled, @@ -544,6 +572,7 @@ func reply(s SiteSettings) settingsReply { if s.Genie.DisabledModels == nil { s.Genie.DisabledModels = []string{} } + s.Genie.CustomModels = append([]CustomModel{}, s.Genie.customList()...) // the list in effect, never null s.Secrets = nil return settingsReply{SiteSettings: s, Languages: languageChoices(), Models: modelInfos(), Defaults: currentGenieDefaults(), Store: store, Version: version} } @@ -570,8 +599,11 @@ func (server *Server) handleAdminSettings(rw http.ResponseWriter, req *http.Requ base = *posted.Version } before := GetSiteSettings() - s.Secrets = before.Secrets // the form cannot set keys; /admin/keys does - s.Admins = before.Admins // nor admins; /admin/admins does, for owners + s.Secrets = before.Secrets // the form cannot set keys; /admin/keys does + s.Admins = before.Admins // nor admins; /admin/admins does, for owners + if s.Genie.CustomModels == nil { + s.Genie.CustomModels = before.Genie.CustomModels // a page from before the list existed does not send it + } s.NodeLanguages = before.NodeLanguages // nor the languages of single nodes; /admin/workers does if err := SaveSiteSettings(s, base); err != nil { if _, bad := s.normalize(); bad != nil { @@ -638,6 +670,7 @@ func settingsChanges(a, b SiteSettings) []string { out = append(out, "model "+chatModels[id].Name+" "+map[bool]string{true: "switched off", false: "switched on"}[b.Genie.ModelDisabled(id)]) } } + out = append(out, customModelChanges(a.Genie, b.Genie)...) if a.Genie.DefaultModel != b.Genie.DefaultModel { name := "built-in" if m, ok := chatModels[b.Genie.DefaultModel]; ok {