fix(assistant): titles and learning use an admin-set utility model (chrono-llm-public / gpt-6-luna) with diagnostics and safe fallback (0.61.3) - #1773
Merged
Conversation
added 3 commits
October 5, 2026 15:34
…y model with diagnostics and safe fallback Production titles silently kept the first message: one_shot_text tried the first discoverable inference service, could pick reasoning or non-chat models with a 128-token budget, and stopped without fallback or logs after any failure. Background generation (titles, learning analysis, voice classifiers) now uses an admin-settable utility route seeded to chrono-llm-public / gpt-6-luna (GET/PUT /api/v1/admin/settings/utility-inference; overrides and explicit clears survive restarts), resolved through live platform-key ACLs and billed to the acting person. Fallback stays on the utility service's best live small chat model and reaches the person's own services only when the utility service is unavailable; at most three attempts, never after an ambiguous or paid call. Reasoning models get minimal effort and an output budget; every attempt logs a metadata-only stage.
📊 Code coverage
Gate: line coverage must stay at or above the threshold. Ratchet plan (W21): Backend → 55%, CLI → 50%, Frontend → 30% by quarter end. |
…iability # Conflicts: # Cargo.lock # backend/Cargo.toml # cli/Cargo.toml # cli/src/wizard/bundle-meta/index.hash # frontend/package-lock.json # frontend/package.json
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Fixes automatic conversation titles in production: threads kept the first user message as their title. Per the owner, background generation now uses NyxID's auto-connected
chrono-llm-publicwith a weak model (gpt-6-luna), as admin-editable configuration rather than a hardcoded policy.Cause (ranked; production had no diagnostics)
one_shot_texttried the first discoverable inference service (platform-available first, then by slug, so Chrono gateways precedellm-openai) and picked any "nano/mini"-style model. Then:Every failure returned
Nonesilently.Changes
{service_slug, model}is seeded at startup tochrono-llm-public/gpt-6-luna.GET/PUT /api/v1/admin/settings/utility-inference(admin only, audited). Settingnullclears it durably; admin overrides and clears survive restarts viautility_inference_admin_modified.$set, so old replicas cannot clobber the fields.shutdown_date) and non-chat models (search, deep-research, codex, computer-use, instruct, -pro and others) are excluded.Backward compatibility
The new fields are serde-defaulted. Absent configuration keeps legacy selection. Learning stays dormant with the flag off.
Validation
--all-targets -D warnings; fmt. No frontend changes; a settings UI can follow.$setwrites (rolling-safe), nodeny_unknown_fieldsonPlatformSettings, and that the voice callers only gained a diagnostic label.