Skip to content

Let admins manage a deployment's AI Gateway models via admin panel - #616

Merged
Maximo-Guk merged 25 commits into
mainfrom
maximo/model-overrides
Oct 2, 2026
Merged

Maximo-Guk merged 25 commits into
mainfrom
maximo/model-overrides

Conversation

@Maximo-Guk

@Maximo-Guk Maximo-Guk commented Sep 30, 2026 •

Copy link
Copy Markdown
Member

Why

A deployment could only change which models its AI Gateway offers by patching SUGGESTED_MODELS. That patch goes stale whenever the catalog changes, and some models can't go in the public catalog at all.

What

In AI Gateway mode, /admin gets a new Models tab:

Enable / test providers and set default reasoning for the deployment and more
image

Change reasoning levels, alter auto compaction levels, enable/hide/disable and test models
image

  • Modes. Each gateway model is Enabled, Hidden (out of the pickers, still works where it is in use) or Disabled (stops working, in existing chats and gadgets too).

Add new models dynamically, with an optional models.dev auto suggestion, also adds a "behaves" like setting which allows for models to behave like another model if PI is unaware what this model is. ( helps when a new model gets released that isn't in PI's catalog ) so it doesn't lose reasoning etc
image

Two more fixes are present: an approval no longer fails when the suspended turn's model stopped resolving (the chat gets the error instead), and the toast for a failed send, new chat or retry now includes the server's reason, such as the model being disabled.

This is so we can show them the following popup:
image

@github-actions github-actions Bot added workshop/frontend Changes to the Workshop frontend kernel Changes to the Workshop kernel workshop/shared Changes to shared Workshop APIs labels Sep 30, 2026
@ask-bonk

ask-bonk Bot commented Sep 30, 2026

Copy link
Copy Markdown

LGTM!

github run

@github-actions

Copy link
Copy Markdown

Preview: pr616-maximo-model-overrides

https://pr616-maximo-model-overrides-router.cloudflare-os-previews.workers.dev

Dashboard · deleted when this PR closes

@Maximo-Guk
Maximo-Guk force-pushed the maximo/model-overrides branch from f46f2e4 to a69b2c1 Compare September 30, 2026 20:00
ask-bonk[bot]

This comment was marked as resolved.

@ask-bonk

ask-bonk Bot commented Sep 30, 2026

Copy link
Copy Markdown

Review: 1 findings.

Posted 1 actionable inline comment.

github run

@github-actions github-actions Bot deleted a comment from ask-bonk Bot Sep 30, 2026
Base automatically changed from maximo/hidden-models to main September 30, 2026 21:08
@Maximo-Guk
Maximo-Guk force-pushed the maximo/model-overrides branch from a69b2c1 to 04499c6 Compare September 30, 2026 22:27
@ask-bonk

ask-bonk Bot commented Sep 30, 2026

Copy link
Copy Markdown

LGTM!

github run

@github-actions github-actions Bot deleted a comment from ask-bonk Bot Sep 30, 2026
@Maximo-Guk
Maximo-Guk force-pushed the maximo/model-overrides branch from 04499c6 to 6531981 Compare October 2, 2026 02:01
@ask-bonk

ask-bonk Bot commented Oct 2, 2026

Copy link
Copy Markdown

LGTM!

github run

@github-actions github-actions Bot deleted a comment from ask-bonk Bot Oct 2, 2026
@github-actions github-actions Bot deleted a comment from ask-bonk Bot Oct 2, 2026
@ask-bonk

ask-bonk Bot commented Oct 2, 2026

Copy link
Copy Markdown

LGTM!

github run

@Maximo-Guk
Maximo-Guk force-pushed the maximo/model-overrides branch from 789ecd7 to 12c6583 Compare October 2, 2026 16:06
@ask-bonk

ask-bonk Bot commented Oct 2, 2026

Copy link
Copy Markdown

LGTM!

github run

@github-actions github-actions Bot deleted a comment from ask-bonk Bot Oct 2, 2026
@ask-bonk

ask-bonk Bot commented Oct 2, 2026

Copy link
Copy Markdown

LGTM!

github run

@github-actions github-actions Bot deleted a comment from ask-bonk Bot Oct 2, 2026
Maximo-Guk and others added 7 commits October 2, 2026 14:46
#resumeSuspendedAgent restarts a turn once its awaited actions are
approved or its connection requests accepted. getChatContext throws
when the thread's model stopped resolving, and nothing caught it, so
an approval that had already been recorded surfaced as "Failed to
approve action" and the chat stalled without saying why.

The resolve is now caught: the failure is logged and posted to the
chat as the agent that was running, carrying the reason, and the
approve or accept RPC succeeds. With no agent message to attribute
the error to, it still throws. Nothing is waiting on a suspended chat
(the turn that suspended already ran #finishAgentTurn), so no waiter
needs releasing.

Co-Authored-By: Claude Code <noreply@anthropic.com>
In AI Gateway mode every gateway model now has a mode: enabled (in
pickers, resolves), hidden (out of pickers, still resolves) or
disabled (out of pickers, does not resolve, ID stays reserved). The
catalog supplies the default, hidden for a superseded model and
enabled otherwise, and AdminConfig.modelModes holds only the models
an admin changed, so an untouched model follows the catalog across
upgrades. AdminConfig.addedModels holds models the deployment adds
beside the catalog's.

GatewayModels merges the catalog, the added models and the modes into
one table, and listing and resolving a gateway model both moved onto
it from AiGatewayConfig, which is env-only again: no caller can
resolve a gateway model without the admin's modes. The catalog wins an
ID an added model also claims, and only added models carry their
limits in the resolved config.

A disabled model is refused with "The "<name>" model is disabled on
this deployment by an administrator." by getChatContext and by
LanguageModelGatekeeper.startSession, so a gadget's model binding
stops at its next call. Its ID shadows a stored model in any mode, so
a stale hand-added model under it is not listed, addable, editable or
deletable. An external message to a chat on a disabled model falls
back to another model, as it does for a deleted one.

getAiConfig() reports builtInModelIds, every gateway model's ID in
any mode. Nothing writes the two AdminConfig fields yet.

Co-Authored-By: Claude Code <noreply@anthropic.com>
AdminApi gains setGatewayModelMode, addGatewayModel and
removeGatewayModel, and getSettings() reports gatewayModels: every
gateway model with its mode, its default mode and whether the
deployment added it, plus the providers a model may be added under.
All of it applies in AI Gateway mode only; the three setters throw
outside it, and a gateway the environment misconfigures is logged and
left out of the settings rather than failing the admin panel.

The AdminSettings methods validate against the config the mutation is
handed, so two concurrent adds of one ID can't both pass and a refused
change leaves storage and the KV mirror alone. Setting a model to its
default mode deletes the override. A model can be added under a
provider the gateway both enables and serves, with an ID that neither
the catalog (under any provider) nor a stored added model has; it is
stored from its sanitized fields and starts in its default mode.
Removing one drops its mode too, and frees the ID: chats and
preferences naming it stop resolving, while a gadget binding minted
for it keeps running, so disabling is how to shut a model off.

GatewayModels.assertAddable is the one add-time check, and
addableProviders the one list of providers the gateway serves and
enables.

Co-Authored-By: Claude Code <noreply@anthropic.com>
The toasts for a failed send, new conversation, retry and first
message on the home page named the action and nothing else, so a chat
on a model an administrator disabled failed without saying why.

They now carry the server's message as their description, through
rpcFailureDescription() in rpcErrors.ts. It leaves the description
out for a non-Error throw, an empty message, and a connection or
Durable Object reset failure, whose message is a transport string.
The titles, the transient-failure handling and the rethrow that keeps
the typed message in the composer are unchanged.

Co-Authored-By: Claude Code <noreply@anthropic.com>
In AI Gateway mode the tab lists the models the deployment provides:
the catalog's, grouped by provider, and the ones the deployment
added. Each row shows the model's name, ID and limits and an Enabled /
Hidden / Disabled radio group that marks the default option and flags
a model that is off its default. Added models can be removed, behind
the shared confirmation dialog, and a form adds one under a provider
the gateway enables. Outside AI Gateway mode the tab holds a notice.

The panel shows what the server holds: every write is followed by a
re-read of the settings, the newest re-read wins, and every control is
locked while a write is in flight. A refused change shows the server's
message, in the form for an add and in a toast otherwise.

PROVIDER_LABELS and parseTokenLimit move from AddModelModal into
features/ai-models/modelForm.ts, which the add form shares. The
providers page marks a model as built in from getAiConfig()'s
builtInModelIds, so models the deployment added, and hidden or
disabled ones, are no longer offered Edit and Delete.

Co-Authored-By: Claude Code <noreply@anthropic.com>
docs/public-server.md describes the three modes and what each does to
pickers, chats and gadget model bindings, how defaults follow the
catalog, adding and removing a model, and the limits: Disabled alone
is not a spend control, a turn in progress finishes on a disabled
model, removing a model does not shut it off, a change is not instant
everywhere, and the title model is outside the modes. AGENTS.md gains
a bullet on where the modes are stored, merged and enforced.

Co-Authored-By: Claude Code <noreply@anthropic.com>
On a gateway deployment a model a user adds by hand needs no
credentials: it runs through the deployment's gateway. So while users
may add their own, disabling a gateway model is not a spend control,
since a user can add the same provider model under another ID.

AdminConfig.userModelsEnabled (default true, AI Gateway mode only)
says whether they may, set through AdminApi.setUserModelsEnabled and
reported by getSettings().gatewayModels and getAiConfig(). While it
is false:

- addModel and updateModel refuse with "Adding your own models is
  disabled on this deployment by an administrator."
- the models a user stored are neither listed nor resolved, so a chat
  on one fails with 'The "<name>" model can't be used: adding your
  own models is disabled on this deployment by an administrator.',
  setPreferredModel refuses it, and an external message falls back off
  it as it does off a disabled model;
- LanguageModelGatekeeper.startSession refuses a binding unless a
  gateway model has its provider and model, so bindings for models
  users added, and for added models since removed, stop at their next
  call.

Nothing stored is deleted or rewritten, deleteModel and getModelConfig
keep working, and turning the setting back on restores every stored
model as it was. A stored value that is not a boolean reads as the
default, and nothing reads the setting outside gateway mode.

GatewayModels exposes the setting as userModels and holds both
sentences in refuseUserModel(), beside refuseDisabled().

Co-Authored-By: Claude Code <noreply@anthropic.com>
Maximo-Guk and others added 10 commits October 2, 2026 14:46
GatewayModelRow.tsx takes the row component and the mode labels, which
the panel's legend also reads. Nothing it renders or does changes.

Co-Authored-By: Claude Code <noreply@anthropic.com>
CF_AI_GATEWAY_PROVIDERS becomes a floor. AdminConfig.addedProviders
holds the providers an admin turned on beside the ones it lists, and
GatewayModels.providers is both: the environment's set as written,
plus the admin's that the gateway serves. A provider an admin turns on
is as one the variable lists: its suggested models appear in their
default modes, models can be added under it, and users can add their
own under it while that is allowed.

Every reader of the allowlist reads GatewayModels.providers: the
model table, assertAddable(), #putModel in user.ts, and
getAiConfig().enabledProviders. AiGatewayConfig stays the
environment's alone.

AdminApi.setGatewayProviderEnabled() validates inside the mutation.
It refuses a provider the gateway does not serve, and refuses to turn
off one the variable lists. Off drops only the provider from
addedProviders: the modes and settings of its models and the models
added under it are kept and return with it. Like removing a provider
from the variable, off stops neither a model a user already added
under it nor a gadget binding already minted.

getSettings().gatewayModels.providerSettings lists every provider the
gateway serves, on or off, with who enabled it and whether its
requests need a CF_AI_GATEWAY_API_TOKEN the deployment lacks. That is
Google on a deployment with none. The environment listing it there
still fails the gateway config, as before. An admin turning it on
never does: the provider is reported as needing the token and its
requests are refused one by one by the existing guard.

The "not supported through AI Gateway" error no longer lists the
environment's providers, which are not all that is enabled.

Co-Authored-By: Claude Code <noreply@anthropic.com>
In AI Gateway mode the Worker holds no provider keys: they, or
credits, are stored in the gateway, and nothing in the Worker can
list them. AdminApi.testGatewayProvider() answers the question an
admin is left with by asking: it sends one small request to the first
of the provider's suggested models through the gateway, as the admin,
and reports what happened within 15 seconds.

A request that fails is a result, not a throw: the provider's or the
gateway's own words on one line, cut to 300 characters, with the
deployment's gateway token cut out should they repeat it, and the
HTTP status when the model runtime reports one. A provider that needs
a token the deployment lacks, and a model that does not answer in
time, are results too. It throws only outside AI Gateway mode and for
a provider the gateway does not serve.

It works on a provider that is off, since the environment alone
decides how a request is routed, so a provider can be tested before
it is turned on. It stores nothing and stays out of the config
mutation queue, so other admin calls run while the request is out.
It logs one gateway.provider.test event with the model, the outcome,
the status and the duration, and never the prompt, the response or
the failure's message.

The request is the same every time, so it carries cf-aig-skip-cache
for a gateway that caches responses. completeText() takes per-request
headers for that.

httpStatusFromError() also reads the status behind the prefix pi's
OpenAI adapter puts on its errors ("OpenAI API error (401): ...").
Without that a failed OpenAI test had no status. The same holds for a
failed OpenAI chat turn, which the overseer's triage now treats like
any other provider's: a 4xx is expected rather than reported as
unknown, and a 5xx is reported with its status. A failed Google
request still has none, because pi's Google adapter reports only the
response body.

Co-Authored-By: Claude Code <noreply@anthropic.com>
Each model's row gets a Settings section with its reasoning level and
its compaction budget. The level select offers "Deployment default
(...)", which clears the model's own level and names what applies
instead, then the levels the model takes. Where the level in effect
is one the model lacks, the row says so and leaves the fit to the
request. A model that takes no levels is offered none. The budget
field names the built-in budget, the maximum, and where a chat
compacts; it refuses what the server would, warns under 100,000
tokens, and Reset returns to the built-in budget. A write always
sends the model's whole settings, so neither setting clobbers the
other, and a stored budget over a maximum that has since shrunk is
clipped to it rather than refused. "Changed" also marks a model with
a setting of its own.

"Default reasoning level", beside the tab's other deployment-wide
settings, is the level of every listed model that has none of its
own.

The add form gets an optional "Behaves like" field, with Kumo's label
tooltip for what it does. It lists the chosen provider's catalog
models that the runtime knows. The added model's row names the model
it behaves like, and says when the choice is unused because this
version knows the model itself, or when it no longer knows the other
one.

The panel returns focus to the control a write was made from once the
write settles. The panel-wide busy flag disables every control while
a write is out, and a browser drops focus from a disabled control
without giving it back. A dialog's own focus is left to the dialog.
A budget refused on Enter in its own field is said aloud the way the
add form says a refused submit; the two share useFieldErrorAlert.

Both reasoning selects act only on an item press, since the select
also reports changes of its own when its options change under a
value it lacks.

docs/public-server.md describes the Settings section, the default
level and "Behaves like"; AGENTS.md says where the settings are
stored and applied.

Co-Authored-By: Claude Code <noreply@anthropic.com>
The tab gets a Providers section: one row for each provider the
gateway serves, with a switch that turns it on for the deployment. A
provider that CF_AI_GATEWAY_PROVIDERS lists is on and locked, with a
note saying so. A row warns when the provider's requests need a
CF_AI_GATEWAY_API_TOKEN the deployment lacks. A switch goes through
the panel's write path like every other setting.

Each row has a Test button, which stays outside the panel's busy
lock: a test runs whatever else the tab is doing, and one provider's
does not wait for another's. Its result stays in the row, in a status
region that is there before the text is, until the provider is tested
again: the model that answered, or "Failed (<status>): <message>",
with a hint on a 401 or 403 that the gateway may hold no key or
credits for the provider, or that the token may not be allowed to run
models. A call that could not be made is said in the same place. A
test in flight leaves its button enabled and ignores a second press,
since a browser takes focus from a button that becomes disabled.

The tab's copy no longer names the variable as the only way to have a
provider.

docs/public-server.md describes the section, what on and off do and
do not stop, the token warning and the Test, says that the variable
is a floor, and adds two limits: turning a provider off is not a
spend control while users may add their own models, and an admin
session can turn on any provider the gateway serves and spend on the
keys it holds. AGENTS.md says where the providers are stored, merged
and tested.

Co-Authored-By: Claude Code <noreply@anthropic.com>
A gateway model with no reasoning level, its own or the deployment's,
is asked for whatever makeHandle() gives its API, and the admin view
had no way to say what that is. AdminModelView.builtInReasoning says
it for each model: 'adaptive' where the model itself decides whether
and how much to reason, a level where the Workshop asks for that
effort, and null where no level is sent.

One function in ai-models.ts, builtInReasoning(), decides it from the
model's descriptor, and makeHandle() builds the request's options
from its answer, so the description and the request cannot drift.
Adaptive is the Anthropic models pi marks adaptive-capable, "medium"
an OpenAI Responses model that reasons, and null everything else:
Haiku 4.5, Gemini, the Workers AI models. gatewayBuiltInReasoning()
answers for the descriptor getModel() builds for a gateway model with
no level, "Behaves like" included, which is the same over either
transport.

Requests are unchanged. An OpenAI Responses model that does no
reasoning reads null rather than "medium" because pi sends such a
model no effort whatever it is handed.

The tests give the answer for every model of the catalog, fail when
the catalog gains one the table lacks, and read the answer back off
the request body a handle with no level sends.

Co-Authored-By: Claude Code <noreply@anthropic.com>
The Providers section owned everything about a test: the state of
each provider's test, the button, and the region its result is said
in. A model's row is about to offer a test of its own, so those move
out of GatewayProviders.tsx as they are.

useGatewayTests() keeps the tests by key, starts one unless a test of
the same key is in flight, and drops a failure that arrives once the
list is gone. The key is a provider here; the tests are held in a Map
so that any string can be one. GatewayTest.tsx has the button, which
is never natively disabled, and the status region with its result
forms.

No behaviour or markup changes, and the section's tests pass
unedited.

Co-Authored-By: Claude Code <noreply@anthropic.com>
The provider test asks a provider's first suggested model for a few
tokens, quickly and with none of a model's settings. It says nothing
of a model an admin added by ID, of what a model "behaves like", or
of what a reasoning level turns into on the wire.

AdminApi.testGatewayModel() sends one model the request a chat turn
would: the config the model runs with, so the reasoning level in
effect for it (its own, else the deployment's default, else its
built-in request) and the flags of the model it behaves like, through
the gateway, as the admin. The response is capped at 2,048 tokens, or
at the model's own cap where that is lower, and the model has 30
seconds to answer. A level that a model takes as a token budget comes
out of that cap, so such a model is tested with a smaller budget than
a chat gives it.

A model in any mode can be tested, so that an admin can try a hidden
or disabled one before enabling it. GatewayModels.runConfig() is the
config of a model whatever its mode, and resolve() is written in
terms of it; whether a model may run is still for resolve() and
refuseDisabled() to say. The config read is the authoritative one in
the durable object, not the KV mirror.

Both tests run through one helper in AdminSettings, which is the
provider test's body: a failed request, a missing token and a timeout
are results, the message is on one line with the gateway token cut
out, nothing is stored and the config mutation queue is not joined.
The provider test's request, timeout and log event are unchanged. A
model test logs gateway.model.test with the model, the outcome, the
status and the duration, and never the prompt, the response or the
failure's message. It throws outside AI Gateway mode and for an ID
that names no gateway model.

completeText() takes an optional `thinking`, off by default as it
was, for the one caller that wants a turn's request rather than a
quick one.

GatewayProviderTest is renamed GatewayModelTest, since both tests
return it.

Co-Authored-By: Claude Code <noreply@anthropic.com>
The tab offered "Built-in" as the default reasoning level and
"Deployment default (built-in)" in each model's Settings without
saying what a model is then asked for, which is not the same for
every model.

While the deployment sets no default, the first option of a model's
"Reasoning level" list names it from the server's builtInReasoning:
"Deployment default (built-in: Adaptive)", "(built-in: Medium)" or
"(built-in: no level sent)". With a default it reads "Deployment
default (<Level>)" as before. The help of "Default reasoning level"
says that Built-in sets no level and that each model's Settings name
what it is then asked for. The wording comes from one function beside
the level labels in modelForm.ts.

docs/public-server.md gives the three forms and which models read
which in this version's catalog: adaptive thinking for every Claude
model but Haiku 4.5, effort medium for the OpenAI models, and no
level for Haiku 4.5, Gemini and the Workers AI models. It says how an
added model the runtime does not know is asked, and that the option
names what the Workshop asks for and not what the runtime or the
provider then does: DeepSeek V4 Pro has its thinking turned off, and
a Claude model whose effort the runtime manages is sent effort high.
AGENTS.md says where the answer is decided.

Co-Authored-By: Claude Code <noreply@anthropic.com>
Each model's row gets a Test button beside its modes, outside
Settings, for a model in any mode. It calls
AdminApi.testGatewayModel(), which sends the model one request the
way a chat turn would, and the result is said in the row, in a status
region that is there before the text is: "Answered through the
gateway.", "Failed (<status>): <message>" with the hint on a 401 or
403, or that the test could not be run. Like a provider's Test it
stays outside the panel's busy lock, leaves its button enabled while
it runs and ignores a second press. A note above the lists says what
a test sends and that it can use up to 2,048 output tokens.

The panel owns the models' tests. A write to a model that goes
through (its mode, its settings, its removal) forgets that model's
result, and a write to the default reasoning level forgets every
model's, so a pass is not left standing over a model that has
changed since. A failed write, a write to something else and a
re-read keep them. useGatewayTests() gains clearTest() and
clearTests() for that: the answer of a test forgotten while in flight
is dropped when it arrives, and the key can be tested again at once.
It tracks the request each key waits on in a ref, so two presses in
one tick also send one request.

GatewayTestStatus takes the subject of the test, since a provider's
pass names the model that answered and a model's does not.

docs/public-server.md describes the model Test: what the request
carries, that hidden and disabled models can be tested and a model
whose provider is off cannot, the result forms, when a result is
cleared, how it differs from a provider's Test and that it costs
more, what a pass proves, and that a Claude model which takes its
level as a token budget is tested with a smaller budget than a chat
gives it. It adds to the limits that tests have no rate limit beyond
the admin check. AGENTS.md names testGatewayModel() and runConfig()
where it describes the tests and where the level is applied.

Co-Authored-By: Claude Code <noreply@anthropic.com>
@Maximo-Guk
Maximo-Guk force-pushed the maximo/model-overrides branch from 5921fb4 to 0bc1589 Compare October 2, 2026 19:53
@ask-bonk

ask-bonk Bot commented Oct 2, 2026

Copy link
Copy Markdown

LGTM!

github run

@github-actions github-actions Bot deleted a comment from ask-bonk Bot Oct 2, 2026
@Maximo-Guk
Maximo-Guk marked this pull request as ready for review October 2, 2026 20:51
@Maximo-Guk Maximo-Guk changed the title Let a deployment hide and add gateway models with CF_AI_GATEWAY_MODELS Let admins manage a deployment's AI Gateway models Oct 2, 2026
@Maximo-Guk Maximo-Guk changed the title Let admins manage a deployment's AI Gateway models Let admins manage a deployment's AI Gateway models via admin panel Oct 2, 2026
devin-ai-integration[bot]

This comment was marked as resolved.

An added model's limits were checked one by one: each a positive
integer. A response is reserved out of the context window, though, so
a model whose output limit is not under its window, or a Workers AI
model with a window of 32,768 tokens or fewer and no output limit,
was added with a prompt budget of zero or less. Every turn of a chat
on it then finds its prompt over budget, and nothing told the admin
why.

addGatewayModel() now refuses such a model, with the size of the
reservation and what to do about it. It asks getModelTokenLimits(),
which is what a turn asks, so the two can't disagree about the room.
The add form shows the refusal beside the fields as it shows any
other.

A model stored with such limits is still refused a compaction budget;
its test stores one directly now, since it can't be added.

Co-Authored-By: Claude Code <noreply@anthropic.com>
@ask-bonk

ask-bonk Bot commented Oct 2, 2026

Copy link
Copy Markdown

LGTM!

github run

@github-actions github-actions Bot deleted a comment from ask-bonk Bot Oct 2, 2026
Comment thread docs/public-server.md Outdated
Maximo-Guk and others added 2 commits October 2, 2026 17:38
The tab's strings, error messages, timeouts and reasoning tables are visible
in the UI or in the code's doc comments.

Co-Authored-By: Claude Code <noreply@anthropic.com>
A disabled model stops what runs with nobody present too, and a scheduled
task that keeps failing ends up dead, which enabling the model again does
not undo. Hidden and Enabled still apply at once.

Co-Authored-By: Claude Code <noreply@anthropic.com>
@ask-bonk

ask-bonk Bot commented Oct 2, 2026

Copy link
Copy Markdown

LGTM!

github run

@Maximo-Guk

Copy link
Copy Markdown
Member Author

Added confirmation when disabling, thanks for the suggestion @ndisidore!
image

@Maximo-Guk
Maximo-Guk merged commit 631ef44 into main Oct 2, 2026
21 of 22 checks passed
@github-actions

github-actions Bot commented Oct 2, 2026

Copy link
Copy Markdown

Eval results

Verdict: ⚪ Unchanged. No task moved beyond what 10 runs can tell apart from noise.

Task Score Δ score Fisher test Cache hits Avg min Avg steps
change-calendar 100% 0 pp p = 1.00 88% → 86%
−2 pp
3.2 → 3.6 26.0 → 24.8
chess 100% → 100% (1 run error) candidate run errors — 97% → 96% 7.9 → 9.1 65.2 → 54.2
incident-desk 90% → 100% +10 pp p = 1.00 94% → 95%
+1 pp
3.6 → 4.4 28.9 → 34.8
worker-logs 100% 0 pp p = 1.00 93% → 92%
−1 pp
3.6 → 3.8 26.0 → 25.1
Failed checks
Task Check Failed
incident-desk t1 opens-acknowledges-and-resolves-in-order 1/10 → 0/10
incident-desk t1 simultaneous-acknowledges-yield-exactly-one-owner 1/10 → 0/10

Run · trajectories and raw results

@Maximo-Guk
Maximo-Guk deleted the maximo/model-overrides branch October 2, 2026 23:06
@github-actions github-actions Bot deleted a comment from ask-bonk Bot Oct 2, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

kernel Changes to the Workshop kernel workshop/frontend Changes to the Workshop frontend workshop/shared Changes to shared Workshop APIs

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants