feat: enable Cohere Command A in chat and advanced mode - #135
Merged
Merged
Conversation
Ports the Cohere LLM client from the cohere-router branch and wires it into main. The model was already provisioned in Azure AI Foundry but never appeared in the app because the config loading, client, and model list endpoint were never merged. - Add CohereLLMClient (Azure AI Hub + Serverless endpoint auto-detection, 3-attempt exponential backoff, asyncio.to_thread for blocking HTTP) - Load [cohere] section from secrets.toml into env vars on startup - Route model_type="cohere" in both extract_entities and generate_paragraph - Expose Cohere in /api/models when COHERE_AZURE_ENDPOINT + KEY are set - Add "cohere" → 256K to chat memory context window estimator (Command A) - Map provider="Cohere" → model_type="cohere" in ChatPage getModelConfig - Add Cohere as Tier 6 in auto-selection priority (modelSelection.ts) - Add "cohere" to extraction provider cost tracking map
…v stack - EntityExtractionPage: add "cohere" branch to all three provider→model_type mappings so Cohere no longer silently routes as Azure during extraction - BatchStudySelectionPage: add missing Cohere accordion section to the AI Models Selection panel - docker-compose.yml: fix <AZURITE_DEV_KEY> placeholder with real well-known Azurite dev key; add --skipApiVersionCheck; wire AZURE_OPENAI_*, COHERE_*, and AZURE_DOC_INTELLIGENCE_* env vars to backend service via .env substitution
Jordan-Leis
force-pushed
the
feature/cohere-command-a
branch
from
June 19, 2026 18:13
4c70bd3 to
286322e
Compare
- Use urlparse().hostname + endswith() for hub/serverless detection instead of substring `in` check, preventing spoofed URLs from matching (CodeQL CWE-020 / incomplete URL sanitization) - Add explanatory comment to bare except json.JSONDecodeError pass so intent is clear (plain-text fallback is deliberate)
There was a problem hiding this comment.
Pull request overview
This PR ports and wires a Cohere (Command A) client into the app end-to-end, making Cohere available for chat, advanced workflows, and extraction/generation when Azure AI Foundry credentials are configured.
Changes:
- Adds a new backend
CohereLLMClientand routesmodel_type="cohere"through extraction and paragraph generation. - Exposes Cohere Command A in
/api/models, updates chat-memory context window estimation for Cohere, and adds Cohere to extraction cost tracking provider mapping. - Updates frontend model/provider mapping, auto-selection priority, and UI grouping to include Cohere models.
Reviewed changes
Copilot reviewed 11 out of 11 changed files in this pull request and generated 5 comments.
Show a summary per file
| File | Description |
|---|---|
| frontend/utils/modelSelection.ts | Adds Cohere into the auto-selection priority list. |
| frontend/components/EntityExtractionPage.tsx | Maps Cohere provider strings to backend model_type="cohere" for extraction flows. |
| frontend/components/ChatPage.tsx | Maps provider "Cohere" to backend model_type="cohere" for chat requests. |
| frontend/components/BatchStudySelectionPage.tsx | Adds a Cohere accordion section to the batch model selection UI. |
| docker-compose.yml | Adds Cohere/Azure OpenAI env vars for local container runs; adjusts Azurite flags and connection string placeholder. |
| backend/services/llm/llm_service.py | Instantiates and routes Cohere client for extraction + paragraph generation. |
| backend/services/llm/cohere.py | Introduces Cohere client implementation targeting Azure AI Foundry hub/serverless endpoints. |
| backend/services/chat_memory/chat_memory_service.py | Adds Cohere context-window estimate (256k). |
| backend/core/config.py | Loads [cohere] secrets.toml section into environment variables. |
| backend/api/server/router.py | Adds Cohere Command A to /models when configured. |
| backend/api/extractions/router.py | Adds Cohere to the extraction provider→cost-tracking map. |
💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.
Comment on lines
+82
to
+86
| // Tier 6 — Cohere | ||
| { | ||
| match: (m) => m.provider === "Cohere", | ||
| modelType: "cohere", | ||
| }, |
Comment on lines
740
to
+743
| if model_type in ("macbook", "vllm"): | ||
| return 32_000 | ||
| if model_type == "cohere": | ||
| return 256_000 # Command A context window |
- cohere.py: replace time.sleep() with await asyncio.sleep() in both retry paths so backoff no longer blocks the event loop - modelSelection.ts: add Cohere branch in modelConfigToSelection() so auto-selected Cohere models get modelType "cohere" not "azure" - router.py: use COHERE_MODEL_NAME env var for model id instead of hard-coded "cohere-command-a" - test_chat_memory_service.py: add test asserting Cohere context window returns 256k tokens
spencer-crook
approved these changes
Jun 19, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Ports the Cohere LLM client from the cohere-router branch and wires it into main. The model was already provisioned in Azure AI Foundry but never appeared in the app because the config loading, client, and model list endpoint were never merged.