Mastra 中的 Voice
Mastra 的 Voice 系统为语音交互提供统一接口,让应用能够使用文本转语音(TTS)、语音转文本(STT)和实时语音转语音(STS)能力。
为 Agent 添加语音能力为 Agent 添加语音能力的直接链接
通过 voice 属性将语音 Provider 传给 Agent。根据配置的 Provider,同一属性支持文本转语音(TTS)、语音转文本(STT)和实时语音转语音(STS)。
import { Agent } from '@mastra/core/agent'
import { OpenAIVoice } from '@mastra/voice-openai'
// Initialize OpenAI voice for TTS
const voiceAgent = new Agent({
id: 'voice-agent',
name: 'Voice Agent',
instructions: 'You are a voice assistant that can help users with their tasks.',
model: 'openai/gpt-5.6-sol',
voice: new OpenAIVoice(),
})
随后可以使用以下语音能力:
文本转语音(TTS)文本转语音(TTS)的直接链接
使用 Mastra 的 TTS 能力将 Agent 响应转换为自然流畅的语音。 可以选择 OpenAI、ElevenLabs 等多个 Provider。
有关详细配置选项和高级功能,请参阅文本转语音指南。
- OpenAI
- Azure
- ElevenLabs
- PlayAI
- Cloudflare
- Deepgram
- Inworld
- Speechify
- Sarvam
- Murf
import { Agent } from '@mastra/core/agent'
import { OpenAIVoice } from '@mastra/voice-openai'
import { playAudio } from '@mastra/node-audio'
const voiceAgent = new Agent({
id: 'voice-agent',
name: 'Voice Agent',
instructions: 'You are a voice assistant that can help users with their tasks.',
model: 'openai/gpt-5.6-sol',
voice: new OpenAIVoice(),
})
const { text } = await voiceAgent.generate('What color is the sky?')
// Convert text to speech to an Audio Stream
const audioStream = await voiceAgent.voice.speak(text, {
speaker: 'default', // Optional: specify a speaker
responseFormat: 'wav', // Optional: specify a response format
})
playAudio(audioStream)
有关 OpenAI 语音 Provider 的更多信息,请参阅 OpenAI Voice 参考。
import { Agent } from '@mastra/core/agent'
import { AzureVoice } from '@mastra/voice-azure'
import { playAudio } from '@mastra/node-audio'
const voiceAgent = new Agent({
id: 'voice-agent',
name: 'Voice Agent',
instructions: 'You are a voice assistant that can help users with their tasks.',
model: 'openai/gpt-5.6-sol',
voice: new AzureVoice(),
})
const { text } = await voiceAgent.generate('What color is the sky?')
// Convert text to speech to an Audio Stream
const audioStream = await voiceAgent.voice.speak(text, {
speaker: 'en-US-JennyNeural', // Optional: specify a speaker
})
playAudio(audioStream)
有关 Azure 语音 Provider 的更多信息,请参阅 Azure Voice 参考。
import { Agent } from '@mastra/core/agent'
import { ElevenLabsVoice } from '@mastra/voice-elevenlabs'
import { playAudio } from '@mastra/node-audio'
const voiceAgent = new Agent({
id: 'voice-agent',
name: 'Voice Agent',
instructions: 'You are a voice assistant that can help users with their tasks.',
model: 'openai/gpt-5.6-sol',
voice: new ElevenLabsVoice(),
})
const { text } = await voiceAgent.generate('What color is the sky?')
// Convert text to speech to an Audio Stream
const audioStream = await voiceAgent.voice.speak(text, {
speaker: 'default', // Optional: specify a speaker
})
playAudio(audioStream)
有关 ElevenLabs 语音 Provider 的更多信息,请参阅 ElevenLabs Voice 参考。
import { Agent } from '@mastra/core/agent'
import { PlayAIVoice } from '@mastra/voice-playai'
import { playAudio } from '@mastra/node-audio'
const voiceAgent = new Agent({
id: 'voice-agent',
name: 'Voice Agent',
instructions: 'You are a voice assistant that can help users with their tasks.',
model: 'openai/gpt-5.6-sol',
voice: new PlayAIVoice(),
})
const { text } = await voiceAgent.generate('What color is the sky?')
// Convert text to speech to an Audio Stream
const audioStream = await voiceAgent.voice.speak(text, {
speaker: 'default', // Optional: specify a speaker
})
playAudio(audioStream)
有关 PlayAI 语音 Provider 的更多信息,请参阅 PlayAI Voice 参考。
import { Agent } from '@mastra/core/agent'
import { GoogleVoice } from '@mastra/voice-google'
import { playAudio } from '@mastra/node-audio'
const voiceAgent = new Agent({
id: 'voice-agent',
name: 'Voice Agent',
instructions: 'You are a voice assistant that can help users with their tasks.',
model: 'openai/gpt-5.6-sol',
voice: new GoogleVoice(),
})
const { text } = await voiceAgent.generate('What color is the sky?')
// Convert text to speech to an Audio Stream
const audioStream = await voiceAgent.voice.speak(text, {
speaker: 'en-US-Studio-O', // Optional: specify a speaker
})
playAudio(audioStream)
有关 Google 语音 Provider 的更多信息,请参阅 Google Voice 参考。
import { Agent } from '@mastra/core/agent'
import { CloudflareVoice } from '@mastra/voice-cloudflare'
import { playAudio } from '@mastra/node-audio'
const voiceAgent = new Agent({
id: 'voice-agent',
name: 'Voice Agent',
instructions: 'You are a voice assistant that can help users with their tasks.',
model: 'openai/gpt-5.6-sol',
voice: new CloudflareVoice(),
})
const { text } = await voiceAgent.generate('What color is the sky?')
// Convert text to speech to an Audio Stream
const audioStream = await voiceAgent.voice.speak(text, {
speaker: 'default', // Optional: specify a speaker
})
playAudio(audioStream)
有关 Cloudflare 语音 Provider 的更多信息,请参阅 Cloudflare Voice 参考。
import { Agent } from '@mastra/core/agent'
import { DeepgramVoice } from '@mastra/voice-deepgram'
import { playAudio } from '@mastra/node-audio'
const voiceAgent = new Agent({
id: 'voice-agent',
name: 'Voice Agent',
instructions: 'You are a voice assistant that can help users with their tasks.',
model: 'openai/gpt-5.6-sol',
voice: new DeepgramVoice(),
})
const { text } = await voiceAgent.generate('What color is the sky?')
// Convert text to speech to an Audio Stream
const audioStream = await voiceAgent.voice.speak(text, {
speaker: 'aura-english-us', // Optional: specify a speaker
})
playAudio(audioStream)
有关 Deepgram 语音 Provider 的更多信息,请参阅 Deepgram Voice 参考。
import { Agent } from '@mastra/core/agent'
import { InworldVoice } from '@mastra/voice-inworld'
import { playAudio } from '@mastra/node-audio'
const voiceAgent = new Agent({
id: 'voice-agent',
name: 'Voice Agent',
instructions: 'You are a voice assistant that can help users with their tasks.',
model: 'openai/gpt-5.6-sol',
voice: new InworldVoice(),
})
const { text } = await voiceAgent.generate('What color is the sky?')
// Convert text to speech to an Audio Stream
const audioStream = await voiceAgent.voice.speak(text, {
speaker: 'Dennis', // Optional: specify a speaker
})
playAudio(audioStream)
有关 Inworld 语音 Provider 的更多信息,请参阅 Inworld Voice 参考。
import { Agent } from '@mastra/core/agent'
import { SpeechifyVoice } from '@mastra/voice-speechify'
import { playAudio } from '@mastra/node-audio'
const voiceAgent = new Agent({
id: 'voice-agent',
name: 'Voice Agent',
instructions: 'You are a voice assistant that can help users with their tasks.',
model: 'openai/gpt-5.6-sol',
voice: new SpeechifyVoice(),
})
const { text } = await voiceAgent.generate('What color is the sky?')
// Convert text to speech to an Audio Stream
const audioStream = await voiceAgent.voice.speak(text, {
speaker: 'matthew', // Optional: specify a speaker
})
playAudio(audioStream)
有关 Speechify 语音 Provider 的更多信息,请参阅 Speechify Voice 参考。
import { Agent } from '@mastra/core/agent'
import { SarvamVoice } from '@mastra/voice-sarvam'
import { playAudio } from '@mastra/node-audio'
const voiceAgent = new Agent({
id: 'voice-agent',
name: 'Voice Agent',
instructions: 'You are a voice assistant that can help users with their tasks.',
model: 'openai/gpt-5.6-sol',
voice: new SarvamVoice(),
})
const { text } = await voiceAgent.generate('What color is the sky?')
// Convert text to speech to an Audio Stream
const audioStream = await voiceAgent.voice.speak(text, {
speaker: 'shubh', // Optional: specify a bulbul:v3 speaker
})
playAudio(audioStream)
有关 Sarvam 语音 Provider 的更多信息,请参阅 Sarvam Voice 参考。
import { Agent } from '@mastra/core/agent'
import { MurfVoice } from '@mastra/voice-murf'
import { playAudio } from '@mastra/node-audio'
const voiceAgent = new Agent({
id: 'voice-agent',
name: 'Voice Agent',
instructions: 'You are a voice assistant that can help users with their tasks.',
model: 'openai/gpt-5.6-sol',
voice: new MurfVoice(),
})
const { text } = await voiceAgent.generate('What color is the sky?')
// Convert text to speech to an Audio Stream
const audioStream = await voiceAgent.voice.speak(text, {
speaker: 'default', // Optional: specify a speaker
})
playAudio(audioStream)
有关 Murf 语音 Provider 的更多信息,请参阅 Murf Voice 参考。
语音转文本(STT)语音转文本(STT)的直接链接
使用 OpenAI、ElevenLabs 等 Provider 转录语音内容。有关详细配置选项等内容,请参阅语音转文本。
你可以从这里下载示例音频文件。
- OpenAI
- Azure
- ElevenLabs
- Cloudflare
- Deepgram
- Inworld
- Sarvam
import { Agent } from '@mastra/core/agent'
import { OpenAIVoice } from '@mastra/voice-openai'
import { createReadStream } from 'fs'
const voiceAgent = new Agent({
id: 'voice-agent',
name: 'Voice Agent',
instructions: 'You are a voice assistant that can help users with their tasks.',
model: 'openai/gpt-5.6-sol',
voice: new OpenAIVoice(),
})
// Use an audio file from a URL
const audioStream = await createReadStream('./how_can_i_help_you.mp3')
// Convert audio to text
const transcript = await voiceAgent.voice.listen(audioStream)
console.log(`User said: ${transcript}`)
// Generate a response based on the transcript
const { text } = await voiceAgent.generate(transcript)
有关 OpenAI 语音 Provider 的更多信息,请参阅 OpenAI Voice 参考。
import { createReadStream } from 'fs'
import { Agent } from '@mastra/core/agent'
import { AzureVoice } from '@mastra/voice-azure'
import { createReadStream } from 'fs'
const voiceAgent = new Agent({
id: 'voice-agent',
name: 'Voice Agent',
instructions: 'You are a voice assistant that can help users with their tasks.',
model: 'openai/gpt-5.6-sol',
voice: new AzureVoice(),
})
// Use an audio file from a URL
const audioStream = await createReadStream('./how_can_i_help_you.mp3')
// Convert audio to text
const transcript = await voiceAgent.voice.listen(audioStream)
console.log(`User said: ${transcript}`)
// Generate a response based on the transcript
const { text } = await voiceAgent.generate(transcript)
有关 Azure 语音 Provider 的更多信息,请参阅 Azure Voice 参考。
import { Agent } from '@mastra/core/agent'
import { ElevenLabsVoice } from '@mastra/voice-elevenlabs'
import { createReadStream } from 'fs'
const voiceAgent = new Agent({
id: 'voice-agent',
name: 'Voice Agent',
instructions: 'You are a voice assistant that can help users with their tasks.',
model: 'openai/gpt-5.6-sol',
voice: new ElevenLabsVoice(),
})
// Use an audio file from a URL
const audioStream = await createReadStream('./how_can_i_help_you.mp3')
// Convert audio to text
const transcript = await voiceAgent.voice.listen(audioStream)
console.log(`User said: ${transcript}`)
// Generate a response based on the transcript
const { text } = await voiceAgent.generate(transcript)
有关 ElevenLabs 语音 Provider 的更多信息,请参阅 ElevenLabs Voice 参考。
import { Agent } from '@mastra/core/agent'
import { GoogleVoice } from '@mastra/voice-google'
import { createReadStream } from 'fs'
const voiceAgent = new Agent({
id: 'voice-agent',
name: 'Voice Agent',
instructions: 'You are a voice assistant that can help users with their tasks.',
model: 'openai/gpt-5.6-sol',
voice: new GoogleVoice(),
})
// Use an audio file from a URL
const audioStream = await createReadStream('./how_can_i_help_you.mp3')
// Convert audio to text
const transcript = await voiceAgent.voice.listen(audioStream)
console.log(`User said: ${transcript}`)
// Generate a response based on the transcript
const { text } = await voiceAgent.generate(transcript)
有关 Google 语音 Provider 的更多信息,请参阅 Google Voice 参考。
import { Agent } from '@mastra/core/agent'
import { CloudflareVoice } from '@mastra/voice-cloudflare'
import { createReadStream } from 'fs'
const voiceAgent = new Agent({
id: 'voice-agent',
name: 'Voice Agent',
instructions: 'You are a voice assistant that can help users with their tasks.',
model: 'openai/gpt-5.6-sol',
voice: new CloudflareVoice(),
})
// Use an audio file from a URL
const audioStream = await createReadStream('./how_can_i_help_you.mp3')
// Convert audio to text
const transcript = await voiceAgent.voice.listen(audioStream)
console.log(`User said: ${transcript}`)
// Generate a response based on the transcript
const { text } = await voiceAgent.generate(transcript)
有关 Cloudflare 语音 Provider 的更多信息,请参阅 Cloudflare Voice 参考。
import { Agent } from '@mastra/core/agent'
import { DeepgramVoice } from '@mastra/voice-deepgram'
import { createReadStream } from 'fs'
const voiceAgent = new Agent({
id: 'voice-agent',
name: 'Voice Agent',
instructions: 'You are a voice assistant that can help users with their tasks.',
model: 'openai/gpt-5.6-sol',
voice: new DeepgramVoice(),
})
// Use an audio file from a URL
const audioStream = await createReadStream('./how_can_i_help_you.mp3')
// Convert audio to text
const transcript = await voiceAgent.voice.listen(audioStream)
console.log(`User said: ${transcript}`)
// Generate a response based on the transcript
const { text } = await voiceAgent.generate(transcript)
有关 Deepgram 语音 Provider 的更多信息,请参阅 Deepgram Voice 参考。
import { Agent } from '@mastra/core/agent'
import { InworldVoice } from '@mastra/voice-inworld'
import { createReadStream } from 'fs'
const voiceAgent = new Agent({
id: 'voice-agent',
name: 'Voice Agent',
instructions: 'You are a voice assistant that can help users with their tasks.',
model: 'openai/gpt-5.6-sol',
voice: new InworldVoice(),
})
// Use an audio file from a URL
const audioStream = await createReadStream('./how_can_i_help_you.mp3')
// Convert audio to text
const transcript = await voiceAgent.voice.listen(audioStream)
console.log(`User said: ${transcript}`)
// Generate a response based on the transcript
const { text } = await voiceAgent.generate(transcript)
有关 Inworld 语音 Provider 的更多信息,请参阅 Inworld Voice 参考。
import { Agent } from '@mastra/core/agent'
import { SarvamVoice } from '@mastra/voice-sarvam'
import { createReadStream } from 'fs'
const voiceAgent = new Agent({
id: 'voice-agent',
name: 'Voice Agent',
instructions: 'You are a voice assistant that can help users with their tasks.',
model: 'openai/gpt-5.6-sol',
voice: new SarvamVoice(),
})
// Use an audio file from a URL
const audioStream = await createReadStream('./how_can_i_help_you.mp3')
// Convert audio to text
const transcript = await voiceAgent.voice.listen(audioStream)
console.log(`User said: ${transcript}`)
// Generate a response based on the transcript
const { text } = await voiceAgent.generate(transcript)
有关 Sarvam 语音 Provider 的更多信息,请参阅 Sarvam Voice 参考。
语音转语音(STS)语音转语音(STS)的直接链接
使用语音转语音能力创建对话体验。统一 API 可实现用户与 AI Agent 之间的实时语音交互。 有关详细配置选项和高级功能,请参阅语音转语音。
- OpenAI
- AWS Nova Sonic
- Inworld Realtime
- xAI
import { Agent } from '@mastra/core/agent'
import { playAudio, getMicrophoneStream } from '@mastra/node-audio'
import { OpenAIRealtimeVoice } from '@mastra/voice-openai-realtime'
const voiceAgent = new Agent({
id: 'voice-agent',
name: 'Voice Agent',
instructions: 'You are a voice assistant that can help users with their tasks.',
model: 'openai/gpt-5.6-sol',
voice: new OpenAIRealtimeVoice(),
})
// Listen for agent audio responses
voiceAgent.voice.on('speaker', ({ audio }) => {
playAudio(audio)
})
// Initiate the conversation
await voiceAgent.voice.speak('How can I help you today?')
// Send continuous audio from the microphone
const micStream = getMicrophoneStream()
await voiceAgent.voice.send(micStream)
有关 OpenAI 语音 Provider 的更多信息,请参阅 OpenAI Voice 参考。
import { Agent } from '@mastra/core/agent'
import { playAudio, getMicrophoneStream } from '@mastra/node-audio'
import { GeminiLiveVoice } from '@mastra/voice-google-gemini-live'
const voiceAgent = new Agent({
id: 'voice-agent',
name: 'Voice Agent',
instructions: 'You are a voice assistant that can help users with their tasks.',
model: 'openai/gpt-5.6-sol',
voice: new GeminiLiveVoice({
// Live API mode
apiKey: process.env.GOOGLE_API_KEY,
model: 'gemini-2.0-flash-exp',
speaker: 'Puck',
debug: true,
// Vertex AI alternative:
// vertexAI: true,
// project: 'your-gcp-project',
// location: 'us-central1',
// serviceAccountKeyFile: '/path/to/service-account.json',
}),
})
// Connect before using speak/send
await voiceAgent.voice.connect()
// Listen for agent audio responses
voiceAgent.voice.on('speaker', ({ audio }) => {
playAudio(audio)
})
// Listen for text responses and transcriptions
voiceAgent.voice.on('writing', ({ text, role }) => {
console.log(`${role}: ${text}`)
})
// Initiate the conversation
await voiceAgent.voice.speak('How can I help you today?')
// Send continuous audio from the microphone
const micStream = getMicrophoneStream()
await voiceAgent.voice.send(micStream)
有关 Google Gemini Live 语音 Provider 的更多信息,请参阅 Google Gemini Live 参考。
import { Agent } from '@mastra/core/agent'
import { playAudio, getMicrophoneStream } from '@mastra/node-audio'
import { NovaSonicVoice } from '@mastra/voice-aws-nova-sonic'
const voiceAgent = new Agent({
id: 'voice-agent',
name: 'Voice Agent',
instructions: 'You are a voice assistant that can help users with their tasks.',
model: 'openai/gpt-5.6-sol',
voice: new NovaSonicVoice({
region: 'us-east-1',
speaker: 'matthew',
// Static credentials are optional. The default AWS credential
// provider chain is used when none are passed.
}),
})
// Connect before using speak/send
await voiceAgent.voice.connect()
// Listen for assistant audio (Int16Array PCM)
voiceAgent.voice.on('speaking', ({ audioData }) => {
if (audioData) playAudio(audioData)
})
// Listen for transcribed text
voiceAgent.voice.on('writing', ({ text, role }) => {
console.log(`${role}: ${text}`)
})
// Initiate the conversation
await voiceAgent.voice.speak('How can I help you today?')
// Send continuous audio from the microphone
const micStream = getMicrophoneStream()
await voiceAgent.voice.send(micStream)
有关 AWS Nova Sonic 语音 Provider 的更多信息,请参阅 AWS Nova Sonic 参考。
import { Agent } from '@mastra/core/agent'
import { playAudio, getMicrophoneStream } from '@mastra/node-audio'
import { InworldRealtimeVoice } from '@mastra/voice-inworld'
const voiceAgent = new Agent({
id: 'voice-agent',
name: 'Voice Agent',
instructions: 'You are a voice assistant that can help users with their tasks.',
model: 'openai/gpt-5.6-sol',
voice: new InworldRealtimeVoice({
apiKey: process.env.INWORLD_API_KEY,
model: 'inworld/models/gemma-4-26b-a4b-it',
speaker: 'Sarah',
}),
})
// Connect before using speak/send
await voiceAgent.voice.connect()
// Listen for agent audio (PCM stream)
voiceAgent.voice.on('speaker', stream => {
playAudio(stream)
})
// Listen for text responses and transcriptions
voiceAgent.voice.on('writing', ({ text, role }) => {
console.log(`${role}: ${text}`)
})
// Initiate the conversation
await voiceAgent.voice.speak('How can I help you today?')
// Send continuous audio from the microphone
const micStream = getMicrophoneStream()
await voiceAgent.voice.send(micStream)
有关 Inworld Realtime 语音 Provider 的更多信息,请参阅 Inworld Realtime 参考。
import { Agent } from '@mastra/core/agent'
import { playAudio, getMicrophoneStream } from '@mastra/node-audio'
import { XAIRealtimeVoice } from '@mastra/voice-xai-realtime'
const voiceAgent = new Agent({
id: 'voice-agent',
name: 'Voice Agent',
instructions: 'You are a voice assistant that can help users with their tasks.',
model: 'xai/grok-4.3',
voice: new XAIRealtimeVoice({
apiKey: process.env.XAI_API_KEY,
model: 'grok-voice-think-fast-1.0',
speaker: 'eve',
turnDetection: { type: 'server_vad' },
}),
})
// Connect before using speak/send
await voiceAgent.voice.connect()
// Listen for agent audio responses
voiceAgent.voice.on('speaker', audioStream => {
playAudio(audioStream)
})
// Listen for text responses and transcriptions
voiceAgent.voice.on('writing', ({ text, role }) => {
console.log(`${role}: ${text}`)
})
// Initiate the conversation
await voiceAgent.voice.speak('How can I help you today?')
// Send continuous audio from the microphone
const micStream = getMicrophoneStream()
await voiceAgent.voice.send(micStream)
有关 xAI 语音 Provider 的更多信息,请参阅 xAI Realtime Voice 参考。
实时语音实时语音的直接链接
在浏览器或电话中运行可由用户打断的实时通话。Mastra 将音频循环交给 LiveKit,由它处理语音活动检测、语义轮次检测和插话;Agent 则使用自己的模型、Tool 和 memory 生成每条回复。有关设置和配置选项,请参阅实时语音。
Voice 配置Voice 配置的直接链接
每个语音 Provider 都可以配置不同的模型和选项。以下是所有受支持 Provider 的详细配置选项:
- OpenAI
- Azure
- ElevenLabs
- PlayAI
- Cloudflare
- Deepgram
- Inworld
- Speechify
- Sarvam
- Murf
- OpenAI Realtime
- xAI Realtime
- Google Gemini Live
- AWS Nova Sonic
- Inworld Realtime
- AI SDK
// OpenAI Voice Configuration
const voice = new OpenAIVoice({
speechModel: {
name: 'gpt-3.5-turbo', // Example model name
apiKey: process.env.OPENAI_API_KEY,
language: 'en-US', // Language code
voiceType: 'neural', // Type of voice model
},
listeningModel: {
name: 'whisper-1', // Example model name
apiKey: process.env.OPENAI_API_KEY,
language: 'en-US', // Language code
format: 'wav', // Audio format
},
speaker: 'alloy', // Example speaker name
})
有关 OpenAI 语音 Provider 的更多信息,请参阅 OpenAI Voice 参考。
// Azure Voice Configuration
const voice = new AzureVoice({
speechModel: {
name: 'en-US-JennyNeural', // Example model name
apiKey: process.env.AZURE_SPEECH_KEY,
region: process.env.AZURE_SPEECH_REGION,
language: 'en-US', // Language code
style: 'cheerful', // Voice style
pitch: '+0Hz', // Pitch adjustment
rate: '1.0', // Speech rate
},
listeningModel: {
name: 'en-US', // Example model name
apiKey: process.env.AZURE_SPEECH_KEY,
region: process.env.AZURE_SPEECH_REGION,
format: 'simple', // Output format
},
})
有关 Azure 语音 Provider 的更多信息,请参阅 Azure Voice 参考。
// ElevenLabs Voice Configuration
const voice = new ElevenLabsVoice({
speechModel: {
voiceId: 'your-voice-id', // Example voice ID
model: 'eleven_multilingual_v2', // Example model name
apiKey: process.env.ELEVENLABS_API_KEY,
language: 'en', // Language code
emotion: 'neutral', // Emotion setting
},
// ElevenLabs may not have a separate listening model
})
有关 ElevenLabs 语音 Provider 的更多信息,请参阅 ElevenLabs Voice 参考。
// PlayAI Voice Configuration
const voice = new PlayAIVoice({
speechModel: {
name: 'playai-voice', // Example model name
speaker: 'emma', // Example speaker name
apiKey: process.env.PLAYAI_API_KEY,
language: 'en-US', // Language code
speed: 1.0, // Speech speed
},
// PlayAI may not have a separate listening model
})
有关 PlayAI 语音 Provider 的更多信息,请参阅 PlayAI Voice 参考。
// Google Voice Configuration
const voice = new GoogleVoice({
speechModel: {
name: 'en-US-Studio-O', // Example model name
apiKey: process.env.GOOGLE_API_KEY,
languageCode: 'en-US', // Language code
gender: 'FEMALE', // Voice gender
speakingRate: 1.0, // Speaking rate
},
listeningModel: {
name: 'en-US', // Example model name
sampleRateHertz: 16000, // Sample rate
},
})
有关 Google 语音 Provider 的更多信息,请参阅 Google Voice 参考。
// Cloudflare Voice Configuration
const voice = new CloudflareVoice({
speechModel: {
name: 'cloudflare-voice', // Example model name
accountId: process.env.CLOUDFLARE_ACCOUNT_ID,
apiToken: process.env.CLOUDFLARE_API_TOKEN,
language: 'en-US', // Language code
format: 'mp3', // Audio format
},
// Cloudflare may not have a separate listening model
})
有关 Cloudflare 语音 Provider 的更多信息,请参阅 Cloudflare Voice 参考。
// Deepgram Voice Configuration
const voice = new DeepgramVoice({
speechModel: {
name: 'nova-2', // Example model name
speaker: 'aura-english-us', // Example speaker name
apiKey: process.env.DEEPGRAM_API_KEY,
language: 'en-US', // Language code
tone: 'formal', // Tone setting
},
listeningModel: {
name: 'nova-2', // Example model name
format: 'flac', // Audio format
},
})
有关 Deepgram 语音 Provider 的更多信息,请参阅 Deepgram Voice 参考。
// Inworld Voice Configuration
const voice = new InworldVoice({
speechModel: {
name: 'inworld-tts-2',
apiKey: process.env.INWORLD_API_KEY,
},
listeningModel: {
name: 'groq/whisper-large-v3',
apiKey: process.env.INWORLD_API_KEY,
},
speaker: 'Dennis',
audioEncoding: 'MP3',
sampleRateHertz: 48000,
language: 'en-US',
})
// Per-call options: `deliveryMode` is honored only by `inworld-tts-2`.
const audioStream = await voice.speak('Hello!', {
deliveryMode: 'BALANCED', // 'STABLE' | 'BALANCED' | 'CREATIVE'
language: 'en-US', // BCP-47 per-call override
})
有关 Inworld 语音 Provider 的更多信息,请参阅 Inworld Voice 参考。
// Speechify Voice Configuration
const voice = new SpeechifyVoice({
speechModel: {
name: 'speechify-voice', // Example model name
speaker: 'matthew', // Example speaker name
apiKey: process.env.SPEECHIFY_API_KEY,
language: 'en-US', // Language code
speed: 1.0, // Speech speed
},
// Speechify may not have a separate listening model
})
有关 Speechify 语音 Provider 的更多信息,请参阅 Speechify Voice 参考。
// Sarvam Voice Configuration
const voice = new SarvamVoice({
speechModel: {
model: 'bulbul:v3', // TTS model (bulbul:v2 or bulbul:v3)
apiKey: process.env.SARVAM_API_KEY,
language: 'en-IN', // BCP-47 language code
},
listeningModel: {
model: 'saarika:v2.5', // STT model (saarika:v2.5 or saaras:v3)
apiKey: process.env.SARVAM_API_KEY,
},
speaker: 'shubh', // Default bulbul:v3 speaker
})
有关 Sarvam 语音 Provider 的更多信息,请参阅 Sarvam Voice 参考。
// Murf Voice Configuration
const voice = new MurfVoice({
speechModel: {
name: 'murf-voice', // Example model name
apiKey: process.env.MURF_API_KEY,
language: 'en-US', // Language code
emotion: 'happy', // Emotion setting
},
// Murf may not have a separate listening model
})
有关 Murf 语音 Provider 的更多信息,请参阅 Murf Voice 参考。
// OpenAI Realtime Voice Configuration
const voice = new OpenAIRealtimeVoice({
speechModel: {
name: 'gpt-3.5-turbo', // Example model name
apiKey: process.env.OPENAI_API_KEY,
language: 'en-US', // Language code
},
listeningModel: {
name: 'whisper-1', // Example model name
apiKey: process.env.OPENAI_API_KEY,
format: 'ogg', // Audio format
},
speaker: 'alloy', // Example speaker name
})
有关 OpenAI Realtime 语音 Provider 的更多信息,请参阅 OpenAI Realtime Voice 参考。
// xAI Realtime Voice Configuration
const voice = new XAIRealtimeVoice({
apiKey: process.env.XAI_API_KEY,
model: 'grok-voice-think-fast-1.0',
speaker: 'eve',
instructions: 'You are a concise voice assistant.',
turnDetection: {
type: 'server_vad',
threshold: 0.85,
silence_duration_ms: 1000,
prefix_padding_ms: 333,
},
audio: {
input: { format: { type: 'audio/pcm', rate: 24000 } },
output: { format: { type: 'audio/pcm', rate: 24000 } },
},
serverTools: [
{ type: 'web_search' },
{
type: 'mcp',
server_url: 'https://mcp.example.com/mcp',
server_label: 'business-tools',
},
],
})
有关 xAI 实时语音 Provider 的更多信息,请参阅 xAI Realtime Voice 参考。
// Google Gemini Live Voice Configuration
const voice = new GeminiLiveVoice({
speechModel: {
name: 'gemini-2.0-flash-exp', // Example model name
apiKey: process.env.GOOGLE_API_KEY,
},
speaker: 'Puck', // Example speaker name
// Google Gemini Live is a realtime bidirectional API without separate speech and listening models
})
有关 Google Gemini Live 语音 Provider 的更多信息,请参阅 Google Gemini Live 参考。
// AWS Nova Sonic Voice Configuration
const voice = new NovaSonicVoice({
region: 'us-east-1',
speaker: 'matthew',
sessionConfig: {
inferenceConfiguration: {
temperature: 0.7,
maxTokens: 1024,
},
turnDetectionConfiguration: {
endpointingSensitivity: 'MEDIUM',
},
},
// AWS Nova Sonic is a realtime bidirectional API without separate speech and listening models
})
有关 AWS Nova Sonic 语音 Provider 的更多信息,请参阅 AWS Nova Sonic 参考。
// Inworld Realtime Voice Configuration
const voice = new InworldRealtimeVoice({
apiKey: process.env.INWORLD_API_KEY,
model: 'inworld/models/gemma-4-26b-a4b-it',
speaker: 'Sarah',
// Typed Inworld realtime knobs (semantic VAD, playback speed, MCP tool routing, ...)
session: {
audio: {
output: { speed: 1.1 },
input: { turn_detection: { type: 'semantic_vad', eagerness: 'high' } },
},
},
})
有关 Inworld Realtime 语音 Provider 的更多信息,请参阅 Inworld Realtime 参考。
// AI SDK Voice Configuration
import { CompositeVoice } from '@mastra/core/voice'
import { openai } from '@ai-sdk/openai'
import { elevenlabs } from '@ai-sdk/elevenlabs'
// Use AI SDK models directly - no need to install separate packages
const voice = new CompositeVoice({
input: openai.transcription('whisper-1'), // AI SDK transcription
output: elevenlabs.speech('eleven_turbo_v2'), // AI SDK speech
})
// Works seamlessly with your agent
const voiceAgent = new Agent({
id: 'aisdk-voice-agent',
name: 'AI SDK Voice Agent',
instructions: 'You are a helpful assistant with voice capabilities.',
model: 'openai/gpt-5.6-sol',
voice,
})
使用多个语音 Provider使用多个语音 Provider的直接链接
此示例演示如何在 Mastra 中创建并使用两个不同的语音 Provider:使用 OpenAI 进行语音转文本(STT),使用 PlayAI 进行文本转语音(TTS)。
首先使用所需配置创建语音 Provider 实例。
import { OpenAIVoice } from '@mastra/voice-openai'
import { PlayAIVoice } from '@mastra/voice-playai'
import { CompositeVoice } from '@mastra/core/voice'
import { playAudio, getMicrophoneStream } from '@mastra/node-audio'
// Initialize OpenAI voice for STT
const input = new OpenAIVoice({
listeningModel: {
name: 'whisper-1',
apiKey: process.env.OPENAI_API_KEY,
},
})
// Initialize PlayAI voice for TTS
const output = new PlayAIVoice({
speechModel: {
name: 'playai-voice',
apiKey: process.env.PLAYAI_API_KEY,
},
})
// Combine the providers using CompositeVoice
const voice = new CompositeVoice({
input,
output,
})
// Implement voice interactions using the combined voice provider
const audioStream = getMicrophoneStream() // Assume this function gets audio input
const transcript = await voice.listen(audioStream)
// Log the transcribed text
console.log('Transcribed text:', transcript)
// Convert text to speech
const responseAudio = await voice.speak(`You said: ${transcript}`, {
speaker: 'default', // Optional: specify a speaker,
responseFormat: 'wav', // Optional: specify a response format
})
// Play the audio response
playAudio(responseAudio)
使用 AI SDK 模型 Provider使用 AI SDK 模型 Provider的直接链接
也可以直接将 AI SDK 模型与 CompositeVoice 配合使用:
import { CompositeVoice } from '@mastra/core/voice'
import { openai } from '@ai-sdk/openai'
import { elevenlabs } from '@ai-sdk/elevenlabs'
import { playAudio, getMicrophoneStream } from '@mastra/node-audio'
// Use AI SDK models directly - no provider setup needed
const voice = new CompositeVoice({
input: openai.transcription('whisper-1'), // AI SDK transcription
output: elevenlabs.speech('eleven_turbo_v2'), // AI SDK speech
})
// Works the same way as Mastra providers
const audioStream = getMicrophoneStream()
const transcript = await voice.listen(audioStream)
console.log('Transcribed text:', transcript)
// Convert text to speech
const responseAudio = await voice.speak(`You said: ${transcript}`, {
speaker: 'Rachel', // ElevenLabs voice
})
playAudio(responseAudio)
还可以混合使用 AI SDK 模型和 Mastra Provider:
import { CompositeVoice } from '@mastra/core/voice'
import { PlayAIVoice } from '@mastra/voice-playai'
import { groq } from '@ai-sdk/groq'
const voice = new CompositeVoice({
input: groq.transcription('whisper-large-v3'), // AI SDK for STT
output: new PlayAIVoice(), // Mastra provider for TTS
})
有关 CompositeVoice 的更多信息,请参阅 CompositeVoice 参考。