You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
Commit 9d52581
Browse filesBrowse the repository at this point in the historyBrowse files
description: 'Text to synthesize, or a list of turns for multi-speaker input. Each turn has its own text, voice, and instructions. Multi-speaker input is currently supported by Gemini TTS models only.'
28617
+
example: 'Hello world'
28609
28618
SpeechInputReference:
28610
28619
description: 'Reference content part for stateless voice cloning or voice design'
28611
28620
discriminator:
@@ -28713,9 +28722,7 @@ components:
28713
28722
voice: 'en_paul_neutral'
28714
28723
properties:
28715
28724
input:
28716
-
description: 'Text to synthesize'
28717
-
example: 'Hello world'
28718
-
type: 'string'
28725
+
$ref: '#/components/schemas/SpeechInput'
28719
28726
input_references:
28720
28727
description: 'Reference content for stateless voice cloning or voice design. Audio mode: one to three `input_audio` parts, each optionally paired with a `text` part carrying its transcript (a single clip accepts its transcript before or after it; with multiple clips each transcript immediately follows its clip); only routed to endpoints that support voice cloning (and multiple references when more than one part is sent). Image mode: exactly one `image_url` part; only routed to endpoints that support image references. The two modes cannot be mixed. An empty array is treated as no reference.'
28721
28728
example:
@@ -28727,6 +28734,10 @@ components:
28727
28734
items:
28728
28735
$ref: '#/components/schemas/SpeechInputReference'
28729
28736
type: 'array'
28737
+
instructions:
28738
+
description: 'Delivery instructions for the whole request, such as tone, pacing, or emotion. Supported by OpenAI gpt-4o-mini-tts and Gemini TTS models. Ignored by other providers.'
28739
+
example: 'Speak in a warm and friendly tone.'
28740
+
type: 'string'
28730
28741
model:
28731
28742
description: 'TTS model identifier'
28732
28743
example: 'mistralai/voxtral-mini-tts-2603'
@@ -28788,6 +28799,24 @@ components:
28788
28799
- 'model'
28789
28800
- 'input'
28790
28801
type: 'object'
28802
+
SpeechTurn:
28803
+
properties:
28804
+
instructions:
28805
+
description: 'Delivery instructions for this turn, such as tone, pacing, or emotion. Overrides the top-level `instructions`.'
28806
+
example: 'whispering'
28807
+
type: 'string'
28808
+
text:
28809
+
description: 'Text to synthesize'
28810
+
example: 'Hi Jane.'
28811
+
type: 'string'
28812
+
voice:
28813
+
description: 'Voice for this turn. Defaults to the top-level `voice`.'
28814
+
example: 'Kore'
28815
+
minLength: 1
28816
+
type: 'string'
28817
+
required:
28818
+
- 'text'
28819
+
type: 'object'
28791
28820
StopServerToolsWhen:
28792
28821
description: 'Stop conditions for the server-tool agent loop. Any condition firing halts the loop (OR logic). When set, this overrides `max_tool_calls`. When a condition fires while the model is still emitting tool calls, the pending tool calls are executed and one final turn is made with tool calls disabled so the response ends with a natural-language answer instead of an unfinished tool call.'
0 commit comments