La voix dans Mastra
Le système Voice de Mastra fournit une interface unifiée pour les interactions vocales. Il permet d'intégrer à vos applications la synthèse vocale (TTS), la transcription vocale (STT) et les échanges vocaux en temps réel (STS).
Ajouter la voix aux agentsLien direct vers Ajouter la voix aux agents
Transmettez un fournisseur vocal à un agent à l'aide de la propriété voice. Selon le fournisseur configuré, cette même propriété prend en charge la synthèse vocale (TTS), la transcription vocale (STT) et les échanges vocaux en temps réel (STS).
import { Agent } from '@mastra/core/agent'
import { OpenAIVoice } from '@mastra/voice-openai'
// Initialize OpenAI voice for TTS
const voiceAgent = new Agent({
id: 'voice-agent',
name: 'Voice Agent',
instructions: 'You are a voice assistant that can help users with their tasks.',
model: 'openai/gpt-5.6-sol',
voice: new OpenAIVoice(),
})
Vous pouvez ensuite utiliser les fonctionnalités vocales suivantes :
Synthèse vocale (TTS)Lien direct vers Synthèse vocale (TTS)
Transformez les réponses de votre agent en parole naturelle grâce aux fonctionnalités TTS de Mastra. Choisissez parmi plusieurs fournisseurs, comme OpenAI, ElevenLabs et bien d'autres.
Pour découvrir les options de configuration détaillées et les fonctionnalités avancées, consultez notre guide sur la synthèse vocale.
- OpenAI
- Azure
- ElevenLabs
- PlayAI
- Cloudflare
- Deepgram
- Inworld
- Speechify
- Sarvam
- Murf
import { Agent } from '@mastra/core/agent'
import { OpenAIVoice } from '@mastra/voice-openai'
import { playAudio } from '@mastra/node-audio'
const voiceAgent = new Agent({
id: 'voice-agent',
name: 'Voice Agent',
instructions: 'You are a voice assistant that can help users with their tasks.',
model: 'openai/gpt-5.6-sol',
voice: new OpenAIVoice(),
})
const { text } = await voiceAgent.generate('What color is the sky?')
// Convert text to speech to an Audio Stream
const audioStream = await voiceAgent.voice.speak(text, {
speaker: 'default', // Optional: specify a speaker
responseFormat: 'wav', // Optional: specify a response format
})
playAudio(audioStream)
Consultez la référence OpenAI Voice pour en savoir plus sur le fournisseur vocal OpenAI.
import { Agent } from '@mastra/core/agent'
import { AzureVoice } from '@mastra/voice-azure'
import { playAudio } from '@mastra/node-audio'
const voiceAgent = new Agent({
id: 'voice-agent',
name: 'Voice Agent',
instructions: 'You are a voice assistant that can help users with their tasks.',
model: 'openai/gpt-5.6-sol',
voice: new AzureVoice(),
})
const { text } = await voiceAgent.generate('What color is the sky?')
// Convert text to speech to an Audio Stream
const audioStream = await voiceAgent.voice.speak(text, {
speaker: 'en-US-JennyNeural', // Optional: specify a speaker
})
playAudio(audioStream)
Consultez la référence Azure Voice pour en savoir plus sur le fournisseur vocal Azure.
import { Agent } from '@mastra/core/agent'
import { ElevenLabsVoice } from '@mastra/voice-elevenlabs'
import { playAudio } from '@mastra/node-audio'
const voiceAgent = new Agent({
id: 'voice-agent',
name: 'Voice Agent',
instructions: 'You are a voice assistant that can help users with their tasks.',
model: 'openai/gpt-5.6-sol',
voice: new ElevenLabsVoice(),
})
const { text } = await voiceAgent.generate('What color is the sky?')
// Convert text to speech to an Audio Stream
const audioStream = await voiceAgent.voice.speak(text, {
speaker: 'default', // Optional: specify a speaker
})
playAudio(audioStream)
Consultez la référence ElevenLabs Voice pour en savoir plus sur le fournisseur vocal ElevenLabs.
import { Agent } from '@mastra/core/agent'
import { PlayAIVoice } from '@mastra/voice-playai'
import { playAudio } from '@mastra/node-audio'
const voiceAgent = new Agent({
id: 'voice-agent',
name: 'Voice Agent',
instructions: 'You are a voice assistant that can help users with their tasks.',
model: 'openai/gpt-5.6-sol',
voice: new PlayAIVoice(),
})
const { text } = await voiceAgent.generate('What color is the sky?')
// Convert text to speech to an Audio Stream
const audioStream = await voiceAgent.voice.speak(text, {
speaker: 'default', // Optional: specify a speaker
})
playAudio(audioStream)
Consultez la référence PlayAI Voice pour en savoir plus sur le fournisseur vocal PlayAI.
import { Agent } from '@mastra/core/agent'
import { GoogleVoice } from '@mastra/voice-google'
import { playAudio } from '@mastra/node-audio'
const voiceAgent = new Agent({
id: 'voice-agent',
name: 'Voice Agent',
instructions: 'You are a voice assistant that can help users with their tasks.',
model: 'openai/gpt-5.6-sol',
voice: new GoogleVoice(),
})
const { text } = await voiceAgent.generate('What color is the sky?')
// Convert text to speech to an Audio Stream
const audioStream = await voiceAgent.voice.speak(text, {
speaker: 'en-US-Studio-O', // Optional: specify a speaker
})
playAudio(audioStream)
Consultez la référence Google Voice pour en savoir plus sur le fournisseur vocal Google.
import { Agent } from '@mastra/core/agent'
import { CloudflareVoice } from '@mastra/voice-cloudflare'
import { playAudio } from '@mastra/node-audio'
const voiceAgent = new Agent({
id: 'voice-agent',
name: 'Voice Agent',
instructions: 'You are a voice assistant that can help users with their tasks.',
model: 'openai/gpt-5.6-sol',
voice: new CloudflareVoice(),
})
const { text } = await voiceAgent.generate('What color is the sky?')
// Convert text to speech to an Audio Stream
const audioStream = await voiceAgent.voice.speak(text, {
speaker: 'default', // Optional: specify a speaker
})
playAudio(audioStream)
Consultez la référence Cloudflare Voice pour en savoir plus sur le fournisseur vocal Cloudflare.
import { Agent } from '@mastra/core/agent'
import { DeepgramVoice } from '@mastra/voice-deepgram'
import { playAudio } from '@mastra/node-audio'
const voiceAgent = new Agent({
id: 'voice-agent',
name: 'Voice Agent',
instructions: 'You are a voice assistant that can help users with their tasks.',
model: 'openai/gpt-5.6-sol',
voice: new DeepgramVoice(),
})
const { text } = await voiceAgent.generate('What color is the sky?')
// Convert text to speech to an Audio Stream
const audioStream = await voiceAgent.voice.speak(text, {
speaker: 'aura-english-us', // Optional: specify a speaker
})
playAudio(audioStream)
Consultez la référence Deepgram Voice pour en savoir plus sur le fournisseur vocal Deepgram.
import { Agent } from '@mastra/core/agent'
import { InworldVoice } from '@mastra/voice-inworld'
import { playAudio } from '@mastra/node-audio'
const voiceAgent = new Agent({
id: 'voice-agent',
name: 'Voice Agent',
instructions: 'You are a voice assistant that can help users with their tasks.',
model: 'openai/gpt-5.6-sol',
voice: new InworldVoice(),
})
const { text } = await voiceAgent.generate('What color is the sky?')
// Convert text to speech to an Audio Stream
const audioStream = await voiceAgent.voice.speak(text, {
speaker: 'Dennis', // Optional: specify a speaker
})
playAudio(audioStream)
Consultez la référence Inworld Voice pour en savoir plus sur le fournisseur vocal Inworld.
import { Agent } from '@mastra/core/agent'
import { SpeechifyVoice } from '@mastra/voice-speechify'
import { playAudio } from '@mastra/node-audio'
const voiceAgent = new Agent({
id: 'voice-agent',
name: 'Voice Agent',
instructions: 'You are a voice assistant that can help users with their tasks.',
model: 'openai/gpt-5.6-sol',
voice: new SpeechifyVoice(),
})
const { text } = await voiceAgent.generate('What color is the sky?')
// Convert text to speech to an Audio Stream
const audioStream = await voiceAgent.voice.speak(text, {
speaker: 'matthew', // Optional: specify a speaker
})
playAudio(audioStream)
Consultez la référence Speechify Voice pour en savoir plus sur le fournisseur vocal Speechify.
import { Agent } from '@mastra/core/agent'
import { SarvamVoice } from '@mastra/voice-sarvam'
import { playAudio } from '@mastra/node-audio'
const voiceAgent = new Agent({
id: 'voice-agent',
name: 'Voice Agent',
instructions: 'You are a voice assistant that can help users with their tasks.',
model: 'openai/gpt-5.6-sol',
voice: new SarvamVoice(),
})
const { text } = await voiceAgent.generate('What color is the sky?')
// Convert text to speech to an Audio Stream
const audioStream = await voiceAgent.voice.speak(text, {
speaker: 'shubh', // Optional: specify a bulbul:v3 speaker
})
playAudio(audioStream)
Consultez la référence Sarvam Voice pour en savoir plus sur le fournisseur vocal Sarvam.
import { Agent } from '@mastra/core/agent'
import { MurfVoice } from '@mastra/voice-murf'
import { playAudio } from '@mastra/node-audio'
const voiceAgent = new Agent({
id: 'voice-agent',
name: 'Voice Agent',
instructions: 'You are a voice assistant that can help users with their tasks.',
model: 'openai/gpt-5.6-sol',
voice: new MurfVoice(),
})
const { text } = await voiceAgent.generate('What color is the sky?')
// Convert text to speech to an Audio Stream
const audioStream = await voiceAgent.voice.speak(text, {
speaker: 'default', // Optional: specify a speaker
})
playAudio(audioStream)
Consultez la référence Murf Voice pour en savoir plus sur le fournisseur vocal Murf.
Transcription vocale (STT)Lien direct vers Transcription vocale (STT)
Transcrivez du contenu parlé à l'aide de fournisseurs comme OpenAI, ElevenLabs et bien d'autres. Pour découvrir les options de configuration détaillées, consultez le guide sur la transcription vocale.
Vous pouvez télécharger un fichier audio d'exemple ici.
- OpenAI
- Azure
- ElevenLabs
- Cloudflare
- Deepgram
- Inworld
- Sarvam
import { Agent } from '@mastra/core/agent'
import { OpenAIVoice } from '@mastra/voice-openai'
import { createReadStream } from 'fs'
const voiceAgent = new Agent({
id: 'voice-agent',
name: 'Voice Agent',
instructions: 'You are a voice assistant that can help users with their tasks.',
model: 'openai/gpt-5.6-sol',
voice: new OpenAIVoice(),
})
// Use an audio file from a URL
const audioStream = await createReadStream('./how_can_i_help_you.mp3')
// Convert audio to text
const transcript = await voiceAgent.voice.listen(audioStream)
console.log(`User said: ${transcript}`)
// Generate a response based on the transcript
const { text } = await voiceAgent.generate(transcript)
Consultez la référence OpenAI Voice pour en savoir plus sur le fournisseur vocal OpenAI.
import { createReadStream } from 'fs'
import { Agent } from '@mastra/core/agent'
import { AzureVoice } from '@mastra/voice-azure'
import { createReadStream } from 'fs'
const voiceAgent = new Agent({
id: 'voice-agent',
name: 'Voice Agent',
instructions: 'You are a voice assistant that can help users with their tasks.',
model: 'openai/gpt-5.6-sol',
voice: new AzureVoice(),
})
// Use an audio file from a URL
const audioStream = await createReadStream('./how_can_i_help_you.mp3')
// Convert audio to text
const transcript = await voiceAgent.voice.listen(audioStream)
console.log(`User said: ${transcript}`)
// Generate a response based on the transcript
const { text } = await voiceAgent.generate(transcript)
Consultez la référence Azure Voice pour en savoir plus sur le fournisseur vocal Azure.
import { Agent } from '@mastra/core/agent'
import { ElevenLabsVoice } from '@mastra/voice-elevenlabs'
import { createReadStream } from 'fs'
const voiceAgent = new Agent({
id: 'voice-agent',
name: 'Voice Agent',
instructions: 'You are a voice assistant that can help users with their tasks.',
model: 'openai/gpt-5.6-sol',
voice: new ElevenLabsVoice(),
})
// Use an audio file from a URL
const audioStream = await createReadStream('./how_can_i_help_you.mp3')
// Convert audio to text
const transcript = await voiceAgent.voice.listen(audioStream)
console.log(`User said: ${transcript}`)
// Generate a response based on the transcript
const { text } = await voiceAgent.generate(transcript)
Consultez la référence ElevenLabs Voice pour en savoir plus sur le fournisseur vocal ElevenLabs.
import { Agent } from '@mastra/core/agent'
import { GoogleVoice } from '@mastra/voice-google'
import { createReadStream } from 'fs'
const voiceAgent = new Agent({
id: 'voice-agent',
name: 'Voice Agent',
instructions: 'You are a voice assistant that can help users with their tasks.',
model: 'openai/gpt-5.6-sol',
voice: new GoogleVoice(),
})
// Use an audio file from a URL
const audioStream = await createReadStream('./how_can_i_help_you.mp3')
// Convert audio to text
const transcript = await voiceAgent.voice.listen(audioStream)
console.log(`User said: ${transcript}`)
// Generate a response based on the transcript
const { text } = await voiceAgent.generate(transcript)
Consultez la référence Google Voice pour en savoir plus sur le fournisseur vocal Google.
import { Agent } from '@mastra/core/agent'
import { CloudflareVoice } from '@mastra/voice-cloudflare'
import { createReadStream } from 'fs'
const voiceAgent = new Agent({
id: 'voice-agent',
name: 'Voice Agent',
instructions: 'You are a voice assistant that can help users with their tasks.',
model: 'openai/gpt-5.6-sol',
voice: new CloudflareVoice(),
})
// Use an audio file from a URL
const audioStream = await createReadStream('./how_can_i_help_you.mp3')
// Convert audio to text
const transcript = await voiceAgent.voice.listen(audioStream)
console.log(`User said: ${transcript}`)
// Generate a response based on the transcript
const { text } = await voiceAgent.generate(transcript)
Consultez la référence Cloudflare Voice pour en savoir plus sur le fournisseur vocal Cloudflare.
import { Agent } from '@mastra/core/agent'
import { DeepgramVoice } from '@mastra/voice-deepgram'
import { createReadStream } from 'fs'
const voiceAgent = new Agent({
id: 'voice-agent',
name: 'Voice Agent',
instructions: 'You are a voice assistant that can help users with their tasks.',
model: 'openai/gpt-5.6-sol',
voice: new DeepgramVoice(),
})
// Use an audio file from a URL
const audioStream = await createReadStream('./how_can_i_help_you.mp3')
// Convert audio to text
const transcript = await voiceAgent.voice.listen(audioStream)
console.log(`User said: ${transcript}`)
// Generate a response based on the transcript
const { text } = await voiceAgent.generate(transcript)
Consultez la référence Deepgram Voice pour en savoir plus sur le fournisseur vocal Deepgram.
import { Agent } from '@mastra/core/agent'
import { InworldVoice } from '@mastra/voice-inworld'
import { createReadStream } from 'fs'
const voiceAgent = new Agent({
id: 'voice-agent',
name: 'Voice Agent',
instructions: 'You are a voice assistant that can help users with their tasks.',
model: 'openai/gpt-5.6-sol',
voice: new InworldVoice(),
})
// Use an audio file from a URL
const audioStream = await createReadStream('./how_can_i_help_you.mp3')
// Convert audio to text
const transcript = await voiceAgent.voice.listen(audioStream)
console.log(`User said: ${transcript}`)
// Generate a response based on the transcript
const { text } = await voiceAgent.generate(transcript)
Consultez la référence Inworld Voice pour en savoir plus sur le fournisseur vocal Inworld.
import { Agent } from '@mastra/core/agent'
import { SarvamVoice } from '@mastra/voice-sarvam'
import { createReadStream } from 'fs'
const voiceAgent = new Agent({
id: 'voice-agent',
name: 'Voice Agent',
instructions: 'You are a voice assistant that can help users with their tasks.',
model: 'openai/gpt-5.6-sol',
voice: new SarvamVoice(),
})
// Use an audio file from a URL
const audioStream = await createReadStream('./how_can_i_help_you.mp3')
// Convert audio to text
const transcript = await voiceAgent.voice.listen(audioStream)
console.log(`User said: ${transcript}`)
// Generate a response based on the transcript
const { text } = await voiceAgent.generate(transcript)
Consultez la référence Sarvam Voice pour en savoir plus sur le fournisseur vocal Sarvam.
Échanges vocaux (STS)Lien direct vers Échanges vocaux (STS)
Créez des expériences conversationnelles grâce aux fonctionnalités d'échanges vocaux. L'API unifiée permet des interactions vocales en temps réel entre les utilisateurs et les agents d'IA. Pour découvrir les options de configuration détaillées et les fonctionnalités avancées, consultez le guide sur les échanges vocaux.
- OpenAI
- AWS Nova Sonic
- Inworld Realtime
- xAI
import { Agent } from '@mastra/core/agent'
import { playAudio, getMicrophoneStream } from '@mastra/node-audio'
import { OpenAIRealtimeVoice } from '@mastra/voice-openai-realtime'
const voiceAgent = new Agent({
id: 'voice-agent',
name: 'Voice Agent',
instructions: 'You are a voice assistant that can help users with their tasks.',
model: 'openai/gpt-5.6-sol',
voice: new OpenAIRealtimeVoice(),
})
// Listen for agent audio responses
voiceAgent.voice.on('speaker', ({ audio }) => {
playAudio(audio)
})
// Initiate the conversation
await voiceAgent.voice.speak('How can I help you today?')
// Send continuous audio from the microphone
const micStream = getMicrophoneStream()
await voiceAgent.voice.send(micStream)
Consultez la référence OpenAI Voice pour en savoir plus sur le fournisseur vocal OpenAI.
import { Agent } from '@mastra/core/agent'
import { playAudio, getMicrophoneStream } from '@mastra/node-audio'
import { GeminiLiveVoice } from '@mastra/voice-google-gemini-live'
const voiceAgent = new Agent({
id: 'voice-agent',
name: 'Voice Agent',
instructions: 'You are a voice assistant that can help users with their tasks.',
model: 'openai/gpt-5.6-sol',
voice: new GeminiLiveVoice({
// Live API mode
apiKey: process.env.GOOGLE_API_KEY,
model: 'gemini-2.0-flash-exp',
speaker: 'Puck',
debug: true,
// Vertex AI alternative:
// vertexAI: true,
// project: 'your-gcp-project',
// location: 'us-central1',
// serviceAccountKeyFile: '/path/to/service-account.json',
}),
})
// Connect before using speak/send
await voiceAgent.voice.connect()
// Listen for agent audio responses
voiceAgent.voice.on('speaker', ({ audio }) => {
playAudio(audio)
})
// Listen for text responses and transcriptions
voiceAgent.voice.on('writing', ({ text, role }) => {
console.log(`${role}: ${text}`)
})
// Initiate the conversation
await voiceAgent.voice.speak('How can I help you today?')
// Send continuous audio from the microphone
const micStream = getMicrophoneStream()
await voiceAgent.voice.send(micStream)
Consultez la référence Google Gemini Live pour en savoir plus sur le fournisseur vocal Google Gemini Live.
import { Agent } from '@mastra/core/agent'
import { playAudio, getMicrophoneStream } from '@mastra/node-audio'
import { NovaSonicVoice } from '@mastra/voice-aws-nova-sonic'
const voiceAgent = new Agent({
id: 'voice-agent',
name: 'Voice Agent',
instructions: 'You are a voice assistant that can help users with their tasks.',
model: 'openai/gpt-5.6-sol',
voice: new NovaSonicVoice({
region: 'us-east-1',
speaker: 'matthew',
// Static credentials are optional. The default AWS credential
// provider chain is used when none are passed.
}),
})
// Connect before using speak/send
await voiceAgent.voice.connect()
// Listen for assistant audio (Int16Array PCM)
voiceAgent.voice.on('speaking', ({ audioData }) => {
if (audioData) playAudio(audioData)
})
// Listen for transcribed text
voiceAgent.voice.on('writing', ({ text, role }) => {
console.log(`${role}: ${text}`)
})
// Initiate the conversation
await voiceAgent.voice.speak('How can I help you today?')
// Send continuous audio from the microphone
const micStream = getMicrophoneStream()
await voiceAgent.voice.send(micStream)
Consultez la référence AWS Nova Sonic pour en savoir plus sur le fournisseur vocal AWS Nova Sonic.
import { Agent } from '@mastra/core/agent'
import { playAudio, getMicrophoneStream } from '@mastra/node-audio'
import { InworldRealtimeVoice } from '@mastra/voice-inworld'
const voiceAgent = new Agent({
id: 'voice-agent',
name: 'Voice Agent',
instructions: 'You are a voice assistant that can help users with their tasks.',
model: 'openai/gpt-5.6-sol',
voice: new InworldRealtimeVoice({
apiKey: process.env.INWORLD_API_KEY,
model: 'inworld/models/gemma-4-26b-a4b-it',
speaker: 'Sarah',
}),
})
// Connect before using speak/send
await voiceAgent.voice.connect()
// Listen for agent audio (PCM stream)
voiceAgent.voice.on('speaker', stream => {
playAudio(stream)
})
// Listen for text responses and transcriptions
voiceAgent.voice.on('writing', ({ text, role }) => {
console.log(`${role}: ${text}`)
})
// Initiate the conversation
await voiceAgent.voice.speak('How can I help you today?')
// Send continuous audio from the microphone
const micStream = getMicrophoneStream()
await voiceAgent.voice.send(micStream)
Consultez la référence Inworld Realtime pour en savoir plus sur le fournisseur vocal Inworld Realtime.
import { Agent } from '@mastra/core/agent'
import { playAudio, getMicrophoneStream } from '@mastra/node-audio'
import { XAIRealtimeVoice } from '@mastra/voice-xai-realtime'
const voiceAgent = new Agent({
id: 'voice-agent',
name: 'Voice Agent',
instructions: 'You are a voice assistant that can help users with their tasks.',
model: 'xai/grok-4.3',
voice: new XAIRealtimeVoice({
apiKey: process.env.XAI_API_KEY,
model: 'grok-voice-think-fast-1.0',
speaker: 'eve',
turnDetection: { type: 'server_vad' },
}),
})
// Connect before using speak/send
await voiceAgent.voice.connect()
// Listen for agent audio responses
voiceAgent.voice.on('speaker', audioStream => {
playAudio(audioStream)
})
// Listen for text responses and transcriptions
voiceAgent.voice.on('writing', ({ text, role }) => {
console.log(`${role}: ${text}`)
})
// Initiate the conversation
await voiceAgent.voice.speak('How can I help you today?')
// Send continuous audio from the microphone
const micStream = getMicrophoneStream()
await voiceAgent.voice.send(micStream)
Consultez la référence xAI Realtime Voice pour en savoir plus sur le fournisseur vocal xAI.
Voix en temps réelLien direct vers Voix en temps réel
Menez des appels en direct dans lesquels un utilisateur peut intervenir, depuis un navigateur ou par téléphone. Mastra confie la boucle audio à LiveKit, qui prend en charge la détection de l'activité vocale, la détection sémantique des tours de parole et l'interruption, tandis que votre agent génère chaque réponse à l'aide de son propre modèle, de ses tools et de sa memory. Pour connaître les options d'installation et de configuration, consultez le guide sur la voix en temps réel.
Configuration de VoiceLien direct vers Configuration de Voice
Chaque fournisseur vocal peut être configuré avec différents modèles et différentes options. Vous trouverez ci-dessous les options de configuration détaillées de tous les fournisseurs pris en charge :
- OpenAI
- Azure
- ElevenLabs
- PlayAI
- Cloudflare
- Deepgram
- Inworld
- Speechify
- Sarvam
- Murf
- OpenAI Realtime
- xAI Realtime
- Google Gemini Live
- AWS Nova Sonic
- Inworld Realtime
- AI SDK
// OpenAI Voice Configuration
const voice = new OpenAIVoice({
speechModel: {
name: 'gpt-3.5-turbo', // Example model name
apiKey: process.env.OPENAI_API_KEY,
language: 'en-US', // Language code
voiceType: 'neural', // Type of voice model
},
listeningModel: {
name: 'whisper-1', // Example model name
apiKey: process.env.OPENAI_API_KEY,
language: 'en-US', // Language code
format: 'wav', // Audio format
},
speaker: 'alloy', // Example speaker name
})
Consultez la référence OpenAI Voice pour en savoir plus sur le fournisseur vocal OpenAI.
// Azure Voice Configuration
const voice = new AzureVoice({
speechModel: {
name: 'en-US-JennyNeural', // Example model name
apiKey: process.env.AZURE_SPEECH_KEY,
region: process.env.AZURE_SPEECH_REGION,
language: 'en-US', // Language code
style: 'cheerful', // Voice style
pitch: '+0Hz', // Pitch adjustment
rate: '1.0', // Speech rate
},
listeningModel: {
name: 'en-US', // Example model name
apiKey: process.env.AZURE_SPEECH_KEY,
region: process.env.AZURE_SPEECH_REGION,
format: 'simple', // Output format
},
})
Consultez la référence Azure Voice pour en savoir plus sur le fournisseur vocal Azure.
// ElevenLabs Voice Configuration
const voice = new ElevenLabsVoice({
speechModel: {
voiceId: 'your-voice-id', // Example voice ID
model: 'eleven_multilingual_v2', // Example model name
apiKey: process.env.ELEVENLABS_API_KEY,
language: 'en', // Language code
emotion: 'neutral', // Emotion setting
},
// ElevenLabs may not have a separate listening model
})
Consultez la référence ElevenLabs Voice pour en savoir plus sur le fournisseur vocal ElevenLabs.
// PlayAI Voice Configuration
const voice = new PlayAIVoice({
speechModel: {
name: 'playai-voice', // Example model name
speaker: 'emma', // Example speaker name
apiKey: process.env.PLAYAI_API_KEY,
language: 'en-US', // Language code
speed: 1.0, // Speech speed
},
// PlayAI may not have a separate listening model
})
Consultez la référence PlayAI Voice pour en savoir plus sur le fournisseur vocal PlayAI.
// Google Voice Configuration
const voice = new GoogleVoice({
speechModel: {
name: 'en-US-Studio-O', // Example model name
apiKey: process.env.GOOGLE_API_KEY,
languageCode: 'en-US', // Language code
gender: 'FEMALE', // Voice gender
speakingRate: 1.0, // Speaking rate
},
listeningModel: {
name: 'en-US', // Example model name
sampleRateHertz: 16000, // Sample rate
},
})
Consultez la référence Google Voice pour en savoir plus sur le fournisseur vocal Google.
// Cloudflare Voice Configuration
const voice = new CloudflareVoice({
speechModel: {
name: 'cloudflare-voice', // Example model name
accountId: process.env.CLOUDFLARE_ACCOUNT_ID,
apiToken: process.env.CLOUDFLARE_API_TOKEN,
language: 'en-US', // Language code
format: 'mp3', // Audio format
},
// Cloudflare may not have a separate listening model
})
Consultez la référence Cloudflare Voice pour en savoir plus sur le fournisseur vocal Cloudflare.
// Deepgram Voice Configuration
const voice = new DeepgramVoice({
speechModel: {
name: 'nova-2', // Example model name
speaker: 'aura-english-us', // Example speaker name
apiKey: process.env.DEEPGRAM_API_KEY,
language: 'en-US', // Language code
tone: 'formal', // Tone setting
},
listeningModel: {
name: 'nova-2', // Example model name
format: 'flac', // Audio format
},
})
Consultez la référence Deepgram Voice pour en savoir plus sur le fournisseur vocal Deepgram.
// Inworld Voice Configuration
const voice = new InworldVoice({
speechModel: {
name: 'inworld-tts-2',
apiKey: process.env.INWORLD_API_KEY,
},
listeningModel: {
name: 'groq/whisper-large-v3',
apiKey: process.env.INWORLD_API_KEY,
},
speaker: 'Dennis',
audioEncoding: 'MP3',
sampleRateHertz: 48000,
language: 'en-US',
})
// Per-call options: `deliveryMode` is honored only by `inworld-tts-2`.
const audioStream = await voice.speak('Hello!', {
deliveryMode: 'BALANCED', // 'STABLE' | 'BALANCED' | 'CREATIVE'
language: 'en-US', // BCP-47 per-call override
})
Consultez la référence Inworld Voice pour en savoir plus sur le fournisseur vocal Inworld.
// Speechify Voice Configuration
const voice = new SpeechifyVoice({
speechModel: {
name: 'speechify-voice', // Example model name
speaker: 'matthew', // Example speaker name
apiKey: process.env.SPEECHIFY_API_KEY,
language: 'en-US', // Language code
speed: 1.0, // Speech speed
},
// Speechify may not have a separate listening model
})
Consultez la référence Speechify Voice pour en savoir plus sur le fournisseur vocal Speechify.
// Sarvam Voice Configuration
const voice = new SarvamVoice({
speechModel: {
model: 'bulbul:v3', // TTS model (bulbul:v2 or bulbul:v3)
apiKey: process.env.SARVAM_API_KEY,
language: 'en-IN', // BCP-47 language code
},
listeningModel: {
model: 'saarika:v2.5', // STT model (saarika:v2.5 or saaras:v3)
apiKey: process.env.SARVAM_API_KEY,
},
speaker: 'shubh', // Default bulbul:v3 speaker
})
Consultez la référence Sarvam Voice pour en savoir plus sur le fournisseur vocal Sarvam.
// Murf Voice Configuration
const voice = new MurfVoice({
speechModel: {
name: 'murf-voice', // Example model name
apiKey: process.env.MURF_API_KEY,
language: 'en-US', // Language code
emotion: 'happy', // Emotion setting
},
// Murf may not have a separate listening model
})
Consultez la référence Murf Voice pour en savoir plus sur le fournisseur vocal Murf.
// OpenAI Realtime Voice Configuration
const voice = new OpenAIRealtimeVoice({
speechModel: {
name: 'gpt-3.5-turbo', // Example model name
apiKey: process.env.OPENAI_API_KEY,
language: 'en-US', // Language code
},
listeningModel: {
name: 'whisper-1', // Example model name
apiKey: process.env.OPENAI_API_KEY,
format: 'ogg', // Audio format
},
speaker: 'alloy', // Example speaker name
})
Pour en savoir plus sur le fournisseur vocal OpenAI Realtime, consultez la référence OpenAI Realtime Voice.
// xAI Realtime Voice Configuration
const voice = new XAIRealtimeVoice({
apiKey: process.env.XAI_API_KEY,
model: 'grok-voice-think-fast-1.0',
speaker: 'eve',
instructions: 'You are a concise voice assistant.',
turnDetection: {
type: 'server_vad',
threshold: 0.85,
silence_duration_ms: 1000,
prefix_padding_ms: 333,
},
audio: {
input: { format: { type: 'audio/pcm', rate: 24000 } },
output: { format: { type: 'audio/pcm', rate: 24000 } },
},
serverTools: [
{ type: 'web_search' },
{
type: 'mcp',
server_url: 'https://mcp.example.com/mcp',
server_label: 'business-tools',
},
],
})
Consultez la référence xAI Realtime Voice pour en savoir plus sur le fournisseur vocal xAI Realtime.
// Google Gemini Live Voice Configuration
const voice = new GeminiLiveVoice({
speechModel: {
name: 'gemini-2.0-flash-exp', // Example model name
apiKey: process.env.GOOGLE_API_KEY,
},
speaker: 'Puck', // Example speaker name
// Google Gemini Live is a realtime bidirectional API without separate speech and listening models
})
Consultez la référence Google Gemini Live pour en savoir plus sur le fournisseur vocal Google Gemini Live.
// AWS Nova Sonic Voice Configuration
const voice = new NovaSonicVoice({
region: 'us-east-1',
speaker: 'matthew',
sessionConfig: {
inferenceConfiguration: {
temperature: 0.7,
maxTokens: 1024,
},
turnDetectionConfiguration: {
endpointingSensitivity: 'MEDIUM',
},
},
// AWS Nova Sonic is a realtime bidirectional API without separate speech and listening models
})
Consultez la référence AWS Nova Sonic pour en savoir plus sur le fournisseur vocal AWS Nova Sonic.
// Inworld Realtime Voice Configuration
const voice = new InworldRealtimeVoice({
apiKey: process.env.INWORLD_API_KEY,
model: 'inworld/models/gemma-4-26b-a4b-it',
speaker: 'Sarah',
// Typed Inworld realtime knobs (semantic VAD, playback speed, MCP tool routing, ...)
session: {
audio: {
output: { speed: 1.1 },
input: { turn_detection: { type: 'semantic_vad', eagerness: 'high' } },
},
},
})
Consultez la référence Inworld Realtime pour en savoir plus sur le fournisseur vocal Inworld Realtime.
// AI SDK Voice Configuration
import { CompositeVoice } from '@mastra/core/voice'
import { openai } from '@ai-sdk/openai'
import { elevenlabs } from '@ai-sdk/elevenlabs'
// Use AI SDK models directly - no need to install separate packages
const voice = new CompositeVoice({
input: openai.transcription('whisper-1'), // AI SDK transcription
output: elevenlabs.speech('eleven_turbo_v2'), // AI SDK speech
})
// Works seamlessly with your agent
const voiceAgent = new Agent({
id: 'aisdk-voice-agent',
name: 'AI SDK Voice Agent',
instructions: 'You are a helpful assistant with voice capabilities.',
model: 'openai/gpt-5.6-sol',
voice,
})
Utiliser plusieurs fournisseurs vocauxLien direct vers Utiliser plusieurs fournisseurs vocaux
Cet exemple montre comment créer et utiliser deux fournisseurs vocaux différents dans Mastra : OpenAI pour la transcription vocale (STT) et PlayAI pour la synthèse vocale (TTS).
Commencez par créer des instances des fournisseurs vocaux avec la configuration nécessaire.
import { OpenAIVoice } from '@mastra/voice-openai'
import { PlayAIVoice } from '@mastra/voice-playai'
import { CompositeVoice } from '@mastra/core/voice'
import { playAudio, getMicrophoneStream } from '@mastra/node-audio'
// Initialize OpenAI voice for STT
const input = new OpenAIVoice({
listeningModel: {
name: 'whisper-1',
apiKey: process.env.OPENAI_API_KEY,
},
})
// Initialize PlayAI voice for TTS
const output = new PlayAIVoice({
speechModel: {
name: 'playai-voice',
apiKey: process.env.PLAYAI_API_KEY,
},
})
// Combine the providers using CompositeVoice
const voice = new CompositeVoice({
input,
output,
})
// Implement voice interactions using the combined voice provider
const audioStream = getMicrophoneStream() // Assume this function gets audio input
const transcript = await voice.listen(audioStream)
// Log the transcribed text
console.log('Transcribed text:', transcript)
// Convert text to speech
const responseAudio = await voice.speak(`You said: ${transcript}`, {
speaker: 'default', // Optional: specify a speaker,
responseFormat: 'wav', // Optional: specify a response format
})
// Play the audio response
playAudio(responseAudio)
Utiliser les Model Providers d'AI SDKLien direct vers Utiliser les Model Providers d'AI SDK
Vous pouvez également utiliser directement les modèles AI SDK avec CompositeVoice :
import { CompositeVoice } from '@mastra/core/voice'
import { openai } from '@ai-sdk/openai'
import { elevenlabs } from '@ai-sdk/elevenlabs'
import { playAudio, getMicrophoneStream } from '@mastra/node-audio'
// Use AI SDK models directly - no provider setup needed
const voice = new CompositeVoice({
input: openai.transcription('whisper-1'), // AI SDK transcription
output: elevenlabs.speech('eleven_turbo_v2'), // AI SDK speech
})
// Works the same way as Mastra providers
const audioStream = getMicrophoneStream()
const transcript = await voice.listen(audioStream)
console.log('Transcribed text:', transcript)
// Convert text to speech
const responseAudio = await voice.speak(`You said: ${transcript}`, {
speaker: 'Rachel', // ElevenLabs voice
})
playAudio(responseAudio)
Vous pouvez aussi associer des modèles AI SDK à des fournisseurs Mastra :
import { CompositeVoice } from '@mastra/core/voice'
import { PlayAIVoice } from '@mastra/voice-playai'
import { groq } from '@ai-sdk/groq'
const voice = new CompositeVoice({
input: groq.transcription('whisper-large-v3'), // AI SDK for STT
output: new PlayAIVoice(), // Mastra provider for TTS
})
Pour en savoir plus sur CompositeVoice, consultez la référence CompositeVoice.