Skip to content
Merged
4 changes: 2 additions & 2 deletions docs/architecture/companion.md
Original file line number Diff line number Diff line change
Expand Up @@ -39,9 +39,9 @@ Model routing runs in-process in the gateway: `servers/gateway/routes/llm-router
| Role | Provider / model | Engine | Notes |
|------|------------------|--------|-------|
| Fast voice (default) | `crow-voice/qwen3.5-4b` (`:8011`) | vLLM-ROCm | Text-only. Qwen3.5-4B is natively vision-language, but its ViT encoder OOMs (256 GiB) under vLLM-ROCm multimodal profiling on gfx1151, so image/video input is disabled (`--limit-mm-per-prompt`). Registered `alwaysResident` with **no mutex group** so it co-resides with the 35B and can never evict it. |
| Escalation (agentic) | `crow-chat/qwen3.6-35b-a3b` (`:8003`) | llama.cpp Vulkan | The daily-driver MoE; **multimodal** (mmproj). Vision-bearing turns escalate here (or to `grackle-vision`). |
| Escalation (agentic) | `crow-chat/qwen3.6-35b-a3b` (`:8003`) | llama.cpp Vulkan | The daily-driver MoE; **multimodal** (mmproj). Vision-bearing turns escalate here. |

Vision on this node is served by the multimodal 35B (stable on Vulkan) and the on-demand `grackle-vision` model — **not** by the fast 4B — so a text-only fast model loses no capability; image turns simply escalate. See [GPU orchestration](/architecture/gateway) for the `mutexGroup` eviction model.
Vision routes through the smart-router's capability pick — the first enabled provider with an image-capable model, or (if none is tagged) the profile fallback, then crow-chat (on crow, the multimodal 35B) — **not** the fast 4B, so a text-only fast model loses no capability; image turns simply escalate. See [GPU orchestration](/architecture/gateway) for the `mutexGroup` eviction model.

### Three model registries

Expand Down
4 changes: 2 additions & 2 deletions docs/es/architecture/companion.md
Original file line number Diff line number Diff line change
Expand Up @@ -39,9 +39,9 @@ El enrutamiento de modelos corre en proceso dentro del gateway: `servers/gateway
| Rol | Proveedor / modelo | Motor | Notas |
|------|------------------|--------|-------|
| Voz rápida (predeterminado) | `crow-voice/qwen3.5-4b` (`:8011`) | vLLM-ROCm | Solo texto. Qwen3.5-4B es nativamente visión-lenguaje, pero su encoder ViT se queda sin memoria (OOM, 256 GiB) bajo el perfilado multimodal de vLLM-ROCm en gfx1151, así que la entrada de imagen/video está deshabilitada (`--limit-mm-per-prompt`). Registrado `alwaysResident` **sin grupo de mutex**, de modo que coexiste con el 35B y nunca puede desalojarlo. |
| Escalado (agéntico) | `crow-chat/qwen3.6-35b-a3b` (`:8003`) | llama.cpp Vulkan | El MoE de uso diario; **multimodal** (mmproj). Los turnos con visión escalan aquí (o a `grackle-vision`). |
| Escalado (agéntico) | `crow-chat/qwen3.6-35b-a3b` (`:8003`) | llama.cpp Vulkan | El MoE de uso diario; **multimodal** (mmproj). Los turnos con visión escalan aquí. |

La visión en este nodo la sirven el 35B multimodal (estable en Vulkan) y el modelo bajo demanda `grackle-vision` — **no** el 4B rápido — así que un modelo rápido de solo texto no pierde ninguna capacidad; los turnos con imágenes simplemente escalan. Consulta la [orquestación de GPU](/es/architecture/gateway) para el modelo de desalojo por `mutexGroup`.
La visión se enruta mediante la selección por capacidad del smart-router — el primer proveedor habilitado con un modelo con capacidad de imagen o, si ninguno está etiquetado así, el fallback del perfil y luego crow-chat (en crow, el 35B multimodal) — **no** el 4B rápido — así que un modelo rápido de solo texto no pierde ninguna capacidad; los turnos con imágenes simplemente escalan. Consulta la [orquestación de GPU](/es/architecture/gateway) para el modelo de desalojo por `mutexGroup`.

### Tres registros de modelos

Expand Down
2 changes: 1 addition & 1 deletion docs/es/guide/ai-providers.md
Original file line number Diff line number Diff line change
Expand Up @@ -232,7 +232,7 @@ Cuando hay configurado un proveedor de embeddings, Crow mejora la búsqueda de m

### Requisitos

- Una entrada de proveedor de embeddings en `models.json` — funciona cualquier endpoint de embeddings compatible con OpenAI (un modelo de embeddings local en vLLM/llama.cpp, Ollama con `nomic-embed-text`, o un proveedor en la nube). Por defecto, Crow busca un proveedor llamado `grackle-embed`.
- Una entrada de proveedor de embeddings en `models.json` — funciona cualquier endpoint de embeddings compatible con OpenAI (un modelo de embeddings local en vLLM/llama.cpp, Ollama con `nomic-embed-text`, o un proveedor en la nube). Por defecto, Crow no depende de ningún host concreto: elige automáticamente el proveedor habilitado con el `id` más bajo que tenga un modelo etiquetado para embeddings.
- Eso es todo — los embeddings se almacenan como BLOBs simples en la tabla `memory_embeddings` y se comparan dentro del propio proceso, lo cual es más que suficiente a la escala de una base de conocimiento personal.

### Cómo funciona
Expand Down
4 changes: 2 additions & 2 deletions docs/guide/ai-providers.md
Original file line number Diff line number Diff line change
Expand Up @@ -237,15 +237,15 @@ When an embedding provider is configured, Crow enhances memory search with **sem

### Choosing the embedding provider

Crow uses `grackle-embed` by default, but the provider is configurable so you can point semantic search at whatever embedder you run. Resolution order (first match wins):
Crow has no hard-coded default provider — it picks one automatically, but the provider is configurable so you can point semantic search at whatever embedder you run. Resolution order (first match wins):

1. **`CROW_EMBED_PROVIDER`** environment variable — best for headless/scripted runs and the gateway (loaded from `.env`).
2. **`embed_provider`** key in `dashboard_settings` — stored in the shared `crow.db`, so it reaches **every** process (the gateway, the MCP servers Claude Code spawns, the sync/backfill scripts) with no re-registration. Set it once:
```sql
INSERT INTO dashboard_settings (key, value) VALUES ('embed_provider', '<provider-id>')
ON CONFLICT(key) DO UPDATE SET value = excluded.value;
```
3. **`grackle-embed`** fallback (preserves prior behavior).
3. **The lowest-id enabled provider with an embed-tagged model** — a capability pick, not a named host. If no provider is tagged for embedding, semantic search stays off.

The value is the provider `id` as registered (e.g. by an embedding bundle). After changing it, allow up to ~30s for the in-process cache to refresh (running processes re-probe the new provider automatically — no restart needed).

Expand Down
Loading
Loading