Skip to content

fix(assistant): titles and learning use an admin-set utility model (chrono-llm-public / gpt-6-luna) with diagnostics and safe fallback (0.61.3) - #1773

Merged
chronoai-kai merged 4 commits into
mainfrom
fix/oneshot-title-reliability
Oct 5, 2026
Merged

chronoai-kai merged 4 commits into
mainfrom
fix/oneshot-title-reliability

Conversation

@chronoai-kai

Copy link
Copy Markdown
Contributor

Summary

Fixes automatic conversation titles in production: threads kept the first user message as their title. Per the owner, background generation now uses NyxID's auto-connected chrono-llm-public with a weak model (gpt-6-luna), as admin-editable configuration rather than a hardcoded policy.

Cause (ranked; production had no diagnostics)

one_shot_text tried the first discoverable inference service (platform-available first, then by slug, so Chrono gateways precede llm-openai) and picked any "nano/mini"-style model. Then:

  1. a reasoning model could spend the 128-token budget without visible text;
  2. a 4xx from an incompatible model or request stopped generation with no fallback;
  3. discovery, credential or timeout failures.

Every failure returned None silently.

Changes

  • Utility route:
    • An atomic platform setting {service_slug, model} is seeded at startup to chrono-llm-public / gpt-6-luna.
    • GET/PUT /api/v1/admin/settings/utility-inference (admin only, audited). Setting null clears it durably; admin overrides and clears survive restarts via utility_inference_admin_modified.
    • Writes use targeted $set, so old replicas cannot clobber the fields.
  • Routing and billing: titles, agent-learning analysis and the voice classifiers resolve the utility service through live platform-key ACLs and are billed to the acting person.
    • Fallback: the best live small chat model on the same service.
    • The person's own services are used only if the utility service is unavailable to them.
  • Safety:
    • At most three attempts, never after an ambiguous or paid call; definite refusals settle at zero usage.
    • Retired (shutdown_date) and non-chat models (search, deep-research, codex, computer-use, instruct, -pro and others) are excluded.
    • Reasoning models get minimal effort and an output budget. Unsupported parameters are cached per service/model after a 400.
  • Diagnostics: one metadata-only warn log per attempt (caller, service, model, stage), never prompts, inputs, outputs or keys.

Backward compatibility

The new fields are serde-defaulted. Absent configuration keeps legacy selection. Learning stays dormant with the flag off.

Validation

  • Implementer, at both default and 1.5 MiB stacks: 91 passed (one-shot 27, titles 7, utility configuration/API 4, learning 15, platform settings 6, voice 32). Rust 1.98.1 clippy --all-targets -D warnings; fmt. No frontend changes; a settings UI can follow.
  • Reviewer: checked admin-only access, durable clears, $set writes (rolling-safe), no deny_unknown_fields on PlatformSettings, and that the voice callers only gained a diagnostic label.

chrono-kw added 3 commits October 5, 2026 15:34
…y model with diagnostics and safe fallback

Production titles silently kept the first message: one_shot_text tried the
first discoverable inference service, could pick reasoning or non-chat models
with a 128-token budget, and stopped without fallback or logs after any
failure. Background generation (titles, learning analysis, voice classifiers)
now uses an admin-settable utility route seeded to chrono-llm-public /
gpt-6-luna (GET/PUT /api/v1/admin/settings/utility-inference; overrides and
explicit clears survive restarts), resolved through live platform-key ACLs and
billed to the acting person. Fallback stays on the utility service's best live
small chat model and reaches the person's own services only when the utility
service is unavailable; at most three attempts, never after an ambiguous or
paid call. Reasoning models get minimal effort and an output budget; every
attempt logs a metadata-only stage.
@github-actions

github-actions Bot commented Oct 5, 2026 •

Copy link
Copy Markdown

📊 Code coverage

Component Lines Threshold Status Δ vs base
Backend (nyxid) 87.06% 73% ✅ 🔺 +0.12
CLI (nyxid-cli) 69.23% 64% ✅ — 0.00
Frontend (vitest) 72.27% 15% ✅ — 0.00

Gate: line coverage must stay at or above the threshold. Ratchet plan (W21): Backend → 55%, CLI → 50%, Frontend → 30% by quarter end.

…iability

# Conflicts:
#	Cargo.lock
#	backend/Cargo.toml
#	cli/Cargo.toml
#	cli/src/wizard/bundle-meta/index.hash
#	frontend/package-lock.json
#	frontend/package.json
@chronoai-kai
chronoai-kai merged commit b3cbfe8 into main Oct 5, 2026
37 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant