feat(node): Add Mistral AI integration - #24243
Conversation
Add span-streaming (`traceLifecycle: 'stream'`) node integration tests for a planned `@mistralai/mistralai` gen_ai integration, mirroring the OpenAI suite. Covers chat, embeddings, agents (invoke_agent) and fim (text_completion), across PII-off, PII-on and explicit-integration-option variants. These tests are expected to fail until the `mistralAIIntegration` / `instrumentMistralClient` instrumentation is implemented (TDD step 1). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
size-limit report 📦
|
Instrument `@mistralai/mistralai` v2 with gen_ai spans, turning the previously-failing integration tests green. - Automatic `mistralAIIntegration()` via the orchestrion diagnostics channels (default-on in Node) - Manual `instrumentMistralClient()` proxy for edge runtimes - Covers chat, embeddings, agents (invoke_agent) and fim (text_completion), including streaming, with `recordInputs` / `recordOutputs` controls Mistral's typed responses/usage are camelCase, so the response/stream mapping reads `promptTokens`/`completionTokens`/`totalTokens` and `choices[].finishReason` directly. `@mistralai/mistralai` v2 is ESM-only, so the CJS test variants are marked `failsOnCjs`. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
- Remove `failsOnCjs` from the Mistral suite: CJS auto-instrumentation works on a full build (it only failed under a partial local rebuild), so the tests pass in both ESM and CJS. - Re-export `mistralAIIntegration` / `instrumentMistralClient` from the dependent SDK packages (aws-serverless, bun, elysia, deno, google-cloud-serverless, astro, cloudflare, vercel-edge) so the node-exports consistency check passes. - Add `Mistral` to the Deno default-integrations snapshot. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Exclude the fim endpoints for now — `text_completion` is not yet used by any other AI integration, so defer it to a follow-up. Remaining scope: chat, embeddings, and agents. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
`@mistralai/mistralai` v2 ships no CJS build, so CJS consumers load it via `require(esm)`, whose auto-instrumentation is inconsistent across Node versions (works on 24/26, fails on 22). The SDK's native mode is ESM, so use `createEsmTests` and cover it there only. Also give the embeddings mock a distinct id per call shape so the single-input span is targeted unambiguously. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
`@langchain/mistralai` drives the `@mistralai/mistralai` SDK under the hood, so with the LangChain integration active both it and `mistralAIIntegration` would instrument the same call, producing two spans. Add `Mistral` to LangChain's `SKIPPED_PROVIDERS`, matching the existing OpenAI/Anthropic/Google handling. This also puts the previously unused `MISTRAL_INTEGRATION_NAME` constant to use. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
- Record request tool definitions on `gen_ai.tool.definitions` (gated by recordInputs), restoring the `recordInputs` arg to `extractRequestAttributes`. - Accumulate streamed tool calls by index (concatenating fragmented `function.arguments`) and emit them on `gen_ai.response.tool_calls`; non-streaming tool calls were already captured. - Add `scenario-tools.mjs` + a test asserting tool definitions and tool calls for both streaming and non-streaming chat. Brings Mistral to parity with the OpenAI integration for function tools. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Read the span-name token from the already-extracted request attributes (model, or the agent name for invoke_agent spans) instead of re-reading the raw params — mirroring the OpenAI integration. Removes the `getModelForSpanName` helper. No change to emitted attributes. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
- Hoist the single `recordOutputs` check to wrap both content and tool calls (was checked twice), matching the OpenAI streaming path. - Drop the dead `?? ''` in tool-call argument accumulation (the first chunk already seeds it to ''). - Simplify the chunk unwrap to a single object guard. No behavior change. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Align `gen_ai.provider.name` (and the matching `sentry.origin`) with the package-scope slug `mistralai` — consistent with how `@google/genai` maps to `google_genai`, and with the existing LangChain path which already reports Mistral as `mistralai`. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…lAiClient Align the manual instrumentation API with the other providers, which all use the `instrument<Provider>AiClient` form (`instrumentOpenAiClient`, `instrumentAnthropicAiClient`, `instrumentWorkersAiClient`). Renamed the export and all per-package re-exports. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…tion # Conflicts: # packages/deno/test/__snapshots__/mod.test.ts.snap
- Set error status on the span when a Mistral stream throws mid-iteration, so failed streams no longer end as successful gen_ai spans. - Decide streaming from the SDK method alone; `stream: true` on `complete` still returns a completion in v2, so it must not be wrapped as a stream. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The default-on Mistral AI integration grows the @sentry/node bundle past the current budgets (130.52 kB > 130, 109.42 kB > 109). Raise both limits. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Cancelling a stream can reject a `read()` that is already in flight, because the cancel aborts the HTTP body the stream reads from. The cancel wrappers awaited the underlying cancel before ending the span, so that rejection reached `settle` first and recorded a deliberate abort as `internal_error`. The caller's intent is now recorded synchronously, before the underlying cancel is awaited, and an error arriving afterwards no longer sets error status. That makes the outcome independent of which of the two settles first. A stream that fails on its own still ends as `internal_error`; only a cancel the caller asked for suppresses it. Note that a spec `ReadableStream` resolves an in-flight read with `done: true` on cancel rather than rejecting it, so this is only reachable on a stream backed by a live connection. The regression test drives the reader directly for that reason. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…Through `tee`, `pipeTo` and `pipeThrough` acquire their reader through internal slots rather than by calling the public `getReader`, so patching `getReader` and the async iterator left them uninstrumented. A stream drained that way delivered every chunk correctly and produced no gen_ai span at all: the span was opened, never ended, and so never flushed. The surrounding trace still looked complete, which made the loss silent. These three now pull through the wrapped reader, so accumulation and span ending stay in one place and a teed stream is still recorded once rather than twice. The sub-stream follows `monitorStream` in `@sentry/deno` with two deliberate differences: chunks are pulled on demand rather than drained in a `start` loop, so backpressure still reaches the source, and `cancel` is forwarded, which is the hole #24054 fixed there. `@sentry/openai` and friends need none of this because their streams are plain async iterables. Mistral's `EventStream` extends `ReadableStream`, so it has drain paths they do not. Workers AI solves the same problem by returning a new stream, which is not available here: the orchestrion path publishes on `asyncEnd` after the caller already holds the stream, and Node's `tracingChannel` returns the original value regardless of what a subscriber assigns to `ctx.result`. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
RulaKhaled
left a comment
There was a problem hiding this comment.
overall looks good, direction of instrumentation makes sense. i pushed three commits that fold in some changes:
- manual path no longer swaps the EventStream for a bare async generator, so getReader/tee/pipeTo/cancel survive
- spans close on every drain path now, not just for await: getReader(), cancel(), tee(), pipeTo(), pipeThrough(). the last three take their reader from internal slots, so they were delivering every chunk fine and producing no gen_ai span at all
- a cancelled stream isn't recorded as internal_error anymore
- invoke_agent spans drop the agent id from the name under span streaming, ids are per-agent so it was high cardinality
- gen_ai.request.stream comes from the method, streaming methods get their own channels
- chat.parse / chat.parseStream produce spans, they call the request fns directly rather than this.complete so they were invisible
- array-shaped streamed content is recorded instead of dropped
two that are minor. gen_ai.response.text is now a stringified array per the conventions spec instead of one concatenated blob, and gen_ai.output.messages is set on both paths since it's the non-deprecated replacement. that meant moving setOutputMessagesAttribute out of ai/workers-ai/utils into ai/core/utils so both share it, workers-ai call sites are unchanged and its suites still pass.
- also bumped the two @sentry/node size limits by 1 KB.
- 32 unit tests plus an integration scenario for instrumentMistralAiClient, which had no coverage at all before.
- stacked the e2e app in #24378 — one express app run twice, dev through the runtime hook and prod through an esbuild build, 24 tests across both. needs retargeting to develop once this lands.
`chat.parse` and `chat.parseStream` had orchestrion entries but were only exercised against a stubbed client, so nothing proved the transform actually matched those methods on the real SDK. Drives both through the mock server with a zod `responseFormat` and asserts one span each. They call the underlying request functions directly rather than `this.complete` / `this.stream`, so the assertion also pins down that they neither go uninstrumented nor produce a second span from the sibling method. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
mydea
left a comment
There was a problem hiding this comment.
generally fine, a few small nits and seems worth it to do a follow up with a bit of cleanup!
|
|
||
| const finishReasons = choices | ||
| .map(choice => choice.finishReason) | ||
| .filter((reason): reason is string => typeof reason === 'string'); |
There was a problem hiding this comment.
l: Maybe we should also check against empty strings? It was done for responseTexts as well
| .filter((reason): reason is string => typeof reason === 'string'); | |
| .filter((reason): reason is string => typeof reason === 'string' && Boolean(reason)); |
There was a problem hiding this comment.
| return { | ||
| responseText: contentToString(message?.content), | ||
| toolCalls: Array.isArray(message?.toolCalls) ? message.toolCalls : undefined, | ||
| finishReason: typeof choice.finishReason === 'string' ? choice.finishReason : undefined, |
There was a problem hiding this comment.
| }; | ||
| }); | ||
|
|
||
| const responseTexts = outputMessages.map(message => message.responseText).filter(Boolean); |
There was a problem hiding this comment.
super-l: Theoretically we could turn around .filter with .map so we could reduce the second iteration a little (not sure if this performance improvement is worth it though):
const responseTexts = outputMessages
.filter(message => Boolean(message.responseText))
.map(message => message.responseText);| if (Array.isArray(response.choices)) { | ||
| const choices = response.choices as Array<Record<string, unknown>>; | ||
|
|
||
| const finishReasons = choices |
There was a problem hiding this comment.
| import { extractSystemInstructions, setTokenUsageAttributes } from '../core/utils'; | ||
| import { extractSystemInstructions, setOutputMessagesAttribute, setTokenUsageAttributes } from '../core/utils'; | ||
| // Re-exported so `workers-ai/streaming.ts` keeps importing it from this module. | ||
| export { setOutputMessagesAttribute }; |
There was a problem hiding this comment.
l: Moving the export to after the imports would make it easier to see that something is exported, even though it wouldn't make a difference for the implenentation
…trument-manual.mjs Co-authored-by: Francesco Gringl-Novy <francesco.novy@sentry.io>
…trument-with-options.mjs Co-authored-by: Francesco Gringl-Novy <francesco.novy@sentry.io>
…trument-with-pii.mjs Co-authored-by: Francesco Gringl-Novy <francesco.novy@sentry.io>
…trument.mjs Co-authored-by: Francesco Gringl-Novy <francesco.novy@sentry.io>
Co-authored-by: Francesco Gringl-Novy <francesco.novy@sentry.io>
| } | ||
| } | ||
|
|
||
| /** |
There was a problem hiding this comment.
Bug: For streaming channels, if a result is not a valid async-iterable, the span ends prematurely and silently fails to record any stream data or response attributes.
Severity: MEDIUM
Suggested Fix
The logic in the end event subscriber should be modified. When deferSpanEnd returns false for a channel configured for streaming, it should not proceed to call endBoundSpan immediately. Instead, it should handle this case as a failed stream instrumentation, potentially by logging a warning, to avoid silently creating an incomplete trace.
Prompt for AI Agent
Review the code at the location below. A potential bug has been identified by an AI
agent. Verify if this is a real issue. If it is, propose a fix; if not, explain why it's
not valid.
Location: packages/server-utils/src/integrations/mistral.ts#L64-L67
Potential issue: In the auto-instrumentation for Mistral streaming channels, the
`deferSpanEnd` callback's behavior depends on `instrumentEventStream`. If a streaming
method returns a result that is not a valid async-iterable (e.g., due to an SDK version
change), `instrumentEventStream` returns `false`. This causes the span to end
immediately instead of being deferred. The `beforeSpanEnd` callback,
`addResponseAttributes`, is then incorrectly called with the stream-like object. While
this does not crash due to defensive checks, it results in a silent failure where the
span closes prematurely without recording any stream data or response attributes,
leading to incomplete traces.
There was a problem hiding this comment.
Cursor Bugbot has reviewed your changes and found 1 potential issue.
❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, enable autofix in the Cursor dashboard.
Reviewed by Cursor Bugbot for commit 6d04ea5. Configure here.
| if (choice.finishReason) { | ||
| state.finishReasons.push(choice.finishReason); | ||
| } | ||
| } |
There was a problem hiding this comment.
Streaming merges multiple choices
Medium Severity
Streamed completions with n greater than 1 fold every choice into one output. Text from all choices is concatenated, and tool calls are keyed only by index, so two choices that both emit index 0 overwrite or join each other's function.arguments. Non-streaming already keeps one gen_ai.output.messages entry per choice.
Additional Locations (1)
Reviewed by Cursor Bugbot for commit 6d04ea5. Configure here.
…24378) Stacked on #24243. Adds an E2E test app for the Mistral integration. The suite runs twice against the same Express app, selected by TEST_ENV. development serves unbundled ESM through the runtime --import hook; production serves an esbuild bundle transformed by sentryEsbuildPlugin. Every assertion runs against both paths, which matters most for a provider that ships ESM-only. What it covers: - gen_ai spans for streaming and non-streaming calls, with request, response and usage attributes. - The tee() and pipeThrough drain paths, each asserting one span and only one. - A failed call reaching Sentry as an error, with the span marked errored and on the same trace. - A manual span under the request span, and the gen_ai span directly under the manual one. - dataloader spans in the same trace as gen_ai spans. Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Documents the Mistral AI integration added in getsentry/sentry-javascript#24243, including automatic and manual setup, privacy controls, and supported operations. Restricts the page and discovery links to supported platforms. ## IS YOUR CHANGE URGENT? - [x] No deadline: Not urgent, can wait up to 1 week+ --------- Co-authored-by: GPT-6 <codex@openai.com> Co-authored-by: Andrei <168741329+andreiborza@users.noreply.github.com> Co-authored-by: Jan Peer Stöcklmair <jan.peer@sentry.io> Co-authored-by: Andrei Borza <andrei.borza@sentry.io>
Export `instrumentMistralAiClient` from the SvelteKit and Remix server entry points, matching their existing AI manual wrappers. Completes the wrapper exports for #24243. --------- Co-authored-by: GPT-6 <codex@openai.com>


Adds a
gen_aiintegration for the@mistralai/mistralaiv2 SDK.Sentry.mistralAIIntegration()Sentry.instrumentMistralAiClient(client)Instruments
chat.complete/stream(gen_ai.chat),embeddings.create(gen_ai.embeddings), andagents.complete/stream(gen_ai.invoke_agent), including streaming, withrecordInputs/recordOutputs(PII) controls. Providermistral, originauto.ai.mistral.Trace from my local sample app:

Sorry for the large PR and the AI integrations more generally could also use some refactors. However, to get this out soon I suggest to follow up on this with a broader sweep.