Skip to content

feat: enable Cohere Command A in chat and advanced mode - #135

Merged
Jordan-Leis merged 5 commits into
mainfrom
feature/cohere-command-a
Jul 10, 2026
Merged

Jordan-Leis merged 5 commits into
mainfrom
feature/cohere-command-a

Conversation

@Jordan-Leis

Copy link
Copy Markdown
Collaborator

Ports the Cohere LLM client from the cohere-router branch and wires it into main. The model was already provisioned in Azure AI Foundry but never appeared in the app because the config loading, client, and model list endpoint were never merged.

  • Add CohereLLMClient (Azure AI Hub + Serverless endpoint auto-detection, 3-attempt exponential backoff, asyncio.to_thread for blocking HTTP)
  • Load [cohere] section from secrets.toml into env vars on startup
  • Route model_type="cohere" in both extract_entities and generate_paragraph
  • Expose Cohere in /api/models when COHERE_AZURE_ENDPOINT + KEY are set
  • Add "cohere" → 256K to chat memory context window estimator (Command A)
  • Map provider="Cohere" → model_type="cohere" in ChatPage getModelConfig
  • Add Cohere as Tier 6 in auto-selection priority (modelSelection.ts)
  • Add "cohere" to extraction provider cost tracking map

Ports the Cohere LLM client from the cohere-router branch and wires it
into main. The model was already provisioned in Azure AI Foundry but
never appeared in the app because the config loading, client, and model
list endpoint were never merged.

- Add CohereLLMClient (Azure AI Hub + Serverless endpoint auto-detection,
  3-attempt exponential backoff, asyncio.to_thread for blocking HTTP)
- Load [cohere] section from secrets.toml into env vars on startup
- Route model_type="cohere" in both extract_entities and generate_paragraph
- Expose Cohere in /api/models when COHERE_AZURE_ENDPOINT + KEY are set
- Add "cohere" → 256K to chat memory context window estimator (Command A)
- Map provider="Cohere" → model_type="cohere" in ChatPage getModelConfig
- Add Cohere as Tier 6 in auto-selection priority (modelSelection.ts)
- Add "cohere" to extraction provider cost tracking map
Comment thread backend/services/llm/cohere.py Fixed
Comment thread backend/services/llm/cohere.py Fixed
Comment thread backend/services/llm/cohere.py Fixed
Jordan-Leis and others added 2 commits June 19, 2026 18:11
…v stack

- EntityExtractionPage: add "cohere" branch to all three provider→model_type
  mappings so Cohere no longer silently routes as Azure during extraction
- BatchStudySelectionPage: add missing Cohere accordion section to the AI
  Models Selection panel
- docker-compose.yml: fix <AZURITE_DEV_KEY> placeholder with real well-known
  Azurite dev key; add --skipApiVersionCheck; wire AZURE_OPENAI_*, COHERE_*,
  and AZURE_DOC_INTELLIGENCE_* env vars to backend service via .env substitution
@Jordan-Leis
Jordan-Leis force-pushed the feature/cohere-command-a branch from 4c70bd3 to 286322e Compare June 19, 2026 18:13
- Use urlparse().hostname + endswith() for hub/serverless detection
  instead of substring `in` check, preventing spoofed URLs from
  matching (CodeQL CWE-020 / incomplete URL sanitization)
- Add explanatory comment to bare except json.JSONDecodeError pass
  so intent is clear (plain-text fallback is deliberate)

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This PR ports and wires a Cohere (Command A) client into the app end-to-end, making Cohere available for chat, advanced workflows, and extraction/generation when Azure AI Foundry credentials are configured.

Changes:

  • Adds a new backend CohereLLMClient and routes model_type="cohere" through extraction and paragraph generation.
  • Exposes Cohere Command A in /api/models, updates chat-memory context window estimation for Cohere, and adds Cohere to extraction cost tracking provider mapping.
  • Updates frontend model/provider mapping, auto-selection priority, and UI grouping to include Cohere models.

Reviewed changes

Copilot reviewed 11 out of 11 changed files in this pull request and generated 5 comments.

Show a summary per file
File Description
frontend/utils/modelSelection.ts Adds Cohere into the auto-selection priority list.
frontend/components/EntityExtractionPage.tsx Maps Cohere provider strings to backend model_type="cohere" for extraction flows.
frontend/components/ChatPage.tsx Maps provider "Cohere" to backend model_type="cohere" for chat requests.
frontend/components/BatchStudySelectionPage.tsx Adds a Cohere accordion section to the batch model selection UI.
docker-compose.yml Adds Cohere/Azure OpenAI env vars for local container runs; adjusts Azurite flags and connection string placeholder.
backend/services/llm/llm_service.py Instantiates and routes Cohere client for extraction + paragraph generation.
backend/services/llm/cohere.py Introduces Cohere client implementation targeting Azure AI Foundry hub/serverless endpoints.
backend/services/chat_memory/chat_memory_service.py Adds Cohere context-window estimate (256k).
backend/core/config.py Loads [cohere] secrets.toml section into environment variables.
backend/api/server/router.py Adds Cohere Command A to /models when configured.
backend/api/extractions/router.py Adds Cohere to the extraction provider→cost-tracking map.

💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.

Comment on lines +82 to +86
// Tier 6 — Cohere
{
match: (m) => m.provider === "Cohere",
modelType: "cohere",
},
Comment thread backend/services/llm/cohere.py Outdated
Comment thread backend/services/llm/cohere.py Outdated
Comment thread backend/api/server/router.py Outdated
Comment on lines 740 to +743
if model_type in ("macbook", "vllm"):
return 32_000
if model_type == "cohere":
return 256_000 # Command A context window
- cohere.py: replace time.sleep() with await asyncio.sleep() in both
  retry paths so backoff no longer blocks the event loop
- modelSelection.ts: add Cohere branch in modelConfigToSelection() so
  auto-selected Cohere models get modelType "cohere" not "azure"
- router.py: use COHERE_MODEL_NAME env var for model id instead of
  hard-coded "cohere-command-a"
- test_chat_memory_service.py: add test asserting Cohere context window
  returns 256k tokens
@Jordan-Leis
Jordan-Leis merged commit 55851c4 into main Jul 10, 2026
10 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants