Mastra 中的 Voice
Mastra 的 Voice 系統提供統一的語音互動介面,讓應用程式能使用文字轉語音(TTS)、語音轉文字(STT)及即時語音轉語音(STS)功能。
為 Agent 加入 Voice「為 Agent 加入 Voice」的直接連結
使用 voice 屬性將 Voice Provider 傳入 Agent。依所設定的 Provider 而定,相同屬性可支援文字轉語音(TTS)、語音轉文字(STT)及即時語音轉語音(STS)。
import { Agent } from '@mastra/core/agent'
import { OpenAIVoice } from '@mastra/voice-openai'
// Initialize OpenAI voice for TTS
const voiceAgent = new Agent({
id: 'voice-agent',
name: 'Voice Agent',
instructions: 'You are a voice assistant that can help users with their tasks.',
model: 'openai/gpt-5.6-sol',
voice: new OpenAIVoice(),
})
接著即可使用下列 Voice 功能:
文字轉語音(TTS)「文字轉語音(TTS)」的直接連結
使用 Mastra 的 TTS 功能,將 Agent 回應轉換成自然的語音。 你可以從 OpenAI、ElevenLabs 等多個 Provider 中選擇。
如需詳細設定選項與進階功能,請參閱文字轉語音指南。
- OpenAI
- Azure
- ElevenLabs
- PlayAI
- Cloudflare
- Deepgram
- Inworld
- Speechify
- Sarvam
- Murf
import { Agent } from '@mastra/core/agent'
import { OpenAIVoice } from '@mastra/voice-openai'
import { playAudio } from '@mastra/node-audio'
const voiceAgent = new Agent({
id: 'voice-agent',
name: 'Voice Agent',
instructions: 'You are a voice assistant that can help users with their tasks.',
model: 'openai/gpt-5.6-sol',
voice: new OpenAIVoice(),
})
const { text } = await voiceAgent.generate('What color is the sky?')
// Convert text to speech to an Audio Stream
const audioStream = await voiceAgent.voice.speak(text, {
speaker: 'default', // Optional: specify a speaker
responseFormat: 'wav', // Optional: specify a response format
})
playAudio(audioStream)
如需 OpenAI Voice Provider 的詳細資訊,請參閱 OpenAI Voice 參考文件。
import { Agent } from '@mastra/core/agent'
import { AzureVoice } from '@mastra/voice-azure'
import { playAudio } from '@mastra/node-audio'
const voiceAgent = new Agent({
id: 'voice-agent',
name: 'Voice Agent',
instructions: 'You are a voice assistant that can help users with their tasks.',
model: 'openai/gpt-5.6-sol',
voice: new AzureVoice(),
})
const { text } = await voiceAgent.generate('What color is the sky?')
// Convert text to speech to an Audio Stream
const audioStream = await voiceAgent.voice.speak(text, {
speaker: 'en-US-JennyNeural', // Optional: specify a speaker
})
playAudio(audioStream)
如需 Azure Voice Provider 的詳細資訊,請參閱 Azure Voice 參考文件。
import { Agent } from '@mastra/core/agent'
import { ElevenLabsVoice } from '@mastra/voice-elevenlabs'
import { playAudio } from '@mastra/node-audio'
const voiceAgent = new Agent({
id: 'voice-agent',
name: 'Voice Agent',
instructions: 'You are a voice assistant that can help users with their tasks.',
model: 'openai/gpt-5.6-sol',
voice: new ElevenLabsVoice(),
})
const { text } = await voiceAgent.generate('What color is the sky?')
// Convert text to speech to an Audio Stream
const audioStream = await voiceAgent.voice.speak(text, {
speaker: 'default', // Optional: specify a speaker
})
playAudio(audioStream)
如需 ElevenLabs Voice Provider 的詳細資訊,請參閱 ElevenLabs Voice 參考文件。
import { Agent } from '@mastra/core/agent'
import { PlayAIVoice } from '@mastra/voice-playai'
import { playAudio } from '@mastra/node-audio'
const voiceAgent = new Agent({
id: 'voice-agent',
name: 'Voice Agent',
instructions: 'You are a voice assistant that can help users with their tasks.',
model: 'openai/gpt-5.6-sol',
voice: new PlayAIVoice(),
})
const { text } = await voiceAgent.generate('What color is the sky?')
// Convert text to speech to an Audio Stream
const audioStream = await voiceAgent.voice.speak(text, {
speaker: 'default', // Optional: specify a speaker
})
playAudio(audioStream)
如需 PlayAI Voice Provider 的詳細資訊,請參閱 PlayAI Voice 參考文件。
import { Agent } from '@mastra/core/agent'
import { GoogleVoice } from '@mastra/voice-google'
import { playAudio } from '@mastra/node-audio'
const voiceAgent = new Agent({
id: 'voice-agent',
name: 'Voice Agent',
instructions: 'You are a voice assistant that can help users with their tasks.',
model: 'openai/gpt-5.6-sol',
voice: new GoogleVoice(),
})
const { text } = await voiceAgent.generate('What color is the sky?')
// Convert text to speech to an Audio Stream
const audioStream = await voiceAgent.voice.speak(text, {
speaker: 'en-US-Studio-O', // Optional: specify a speaker
})
playAudio(audioStream)
如需 Google Voice Provider 的詳細資訊,請參閱 Google Voice 參考文件。
import { Agent } from '@mastra/core/agent'
import { CloudflareVoice } from '@mastra/voice-cloudflare'
import { playAudio } from '@mastra/node-audio'
const voiceAgent = new Agent({
id: 'voice-agent',
name: 'Voice Agent',
instructions: 'You are a voice assistant that can help users with their tasks.',
model: 'openai/gpt-5.6-sol',
voice: new CloudflareVoice(),
})
const { text } = await voiceAgent.generate('What color is the sky?')
// Convert text to speech to an Audio Stream
const audioStream = await voiceAgent.voice.speak(text, {
speaker: 'default', // Optional: specify a speaker
})
playAudio(audioStream)
如需 Cloudflare Voice Provider 的詳細資訊,請參閱 Cloudflare Voice 參考文件。
import { Agent } from '@mastra/core/agent'
import { DeepgramVoice } from '@mastra/voice-deepgram'
import { playAudio } from '@mastra/node-audio'
const voiceAgent = new Agent({
id: 'voice-agent',
name: 'Voice Agent',
instructions: 'You are a voice assistant that can help users with their tasks.',
model: 'openai/gpt-5.6-sol',
voice: new DeepgramVoice(),
})
const { text } = await voiceAgent.generate('What color is the sky?')
// Convert text to speech to an Audio Stream
const audioStream = await voiceAgent.voice.speak(text, {
speaker: 'aura-english-us', // Optional: specify a speaker
})
playAudio(audioStream)
如需 Deepgram Voice Provider 的詳細資訊,請參閱 Deepgram Voice 參考文件。
import { Agent } from '@mastra/core/agent'
import { InworldVoice } from '@mastra/voice-inworld'
import { playAudio } from '@mastra/node-audio'
const voiceAgent = new Agent({
id: 'voice-agent',
name: 'Voice Agent',
instructions: 'You are a voice assistant that can help users with their tasks.',
model: 'openai/gpt-5.6-sol',
voice: new InworldVoice(),
})
const { text } = await voiceAgent.generate('What color is the sky?')
// Convert text to speech to an Audio Stream
const audioStream = await voiceAgent.voice.speak(text, {
speaker: 'Dennis', // Optional: specify a speaker
})
playAudio(audioStream)
如需 Inworld Voice Provider 的詳細資訊,請參閱 Inworld Voice 參考文件。
import { Agent } from '@mastra/core/agent'
import { SpeechifyVoice } from '@mastra/voice-speechify'
import { playAudio } from '@mastra/node-audio'
const voiceAgent = new Agent({
id: 'voice-agent',
name: 'Voice Agent',
instructions: 'You are a voice assistant that can help users with their tasks.',
model: 'openai/gpt-5.6-sol',
voice: new SpeechifyVoice(),
})
const { text } = await voiceAgent.generate('What color is the sky?')
// Convert text to speech to an Audio Stream
const audioStream = await voiceAgent.voice.speak(text, {
speaker: 'matthew', // Optional: specify a speaker
})
playAudio(audioStream)
如需 Speechify Voice Provider 的詳細資訊,請參閱 Speechify Voice 參考文件。
import { Agent } from '@mastra/core/agent'
import { SarvamVoice } from '@mastra/voice-sarvam'
import { playAudio } from '@mastra/node-audio'
const voiceAgent = new Agent({
id: 'voice-agent',
name: 'Voice Agent',
instructions: 'You are a voice assistant that can help users with their tasks.',
model: 'openai/gpt-5.6-sol',
voice: new SarvamVoice(),
})
const { text } = await voiceAgent.generate('What color is the sky?')
// Convert text to speech to an Audio Stream
const audioStream = await voiceAgent.voice.speak(text, {
speaker: 'shubh', // Optional: specify a bulbul:v3 speaker
})
playAudio(audioStream)
如需 Sarvam Voice Provider 的詳細資訊,請參閱 Sarvam Voice 參考文件。
import { Agent } from '@mastra/core/agent'
import { MurfVoice } from '@mastra/voice-murf'
import { playAudio } from '@mastra/node-audio'
const voiceAgent = new Agent({
id: 'voice-agent',
name: 'Voice Agent',
instructions: 'You are a voice assistant that can help users with their tasks.',
model: 'openai/gpt-5.6-sol',
voice: new MurfVoice(),
})
const { text } = await voiceAgent.generate('What color is the sky?')
// Convert text to speech to an Audio Stream
const audioStream = await voiceAgent.voice.speak(text, {
speaker: 'default', // Optional: specify a speaker
})
playAudio(audioStream)
如需 Murf Voice Provider 的詳細資訊,請參閱 Murf Voice 參考文件。
語音轉文字(STT)「語音轉文字(STT)」的直接連結
使用 OpenAI、ElevenLabs 等 Provider 轉錄語音內容。如需詳細設定選項與其他資訊,請參閱語音轉文字。
你可以從此處下載音訊範例檔案。
- OpenAI
- Azure
- ElevenLabs
- Cloudflare
- Deepgram
- Inworld
- Sarvam
import { Agent } from '@mastra/core/agent'
import { OpenAIVoice } from '@mastra/voice-openai'
import { createReadStream } from 'fs'
const voiceAgent = new Agent({
id: 'voice-agent',
name: 'Voice Agent',
instructions: 'You are a voice assistant that can help users with their tasks.',
model: 'openai/gpt-5.6-sol',
voice: new OpenAIVoice(),
})
// Use an audio file from a URL
const audioStream = await createReadStream('./how_can_i_help_you.mp3')
// Convert audio to text
const transcript = await voiceAgent.voice.listen(audioStream)
console.log(`User said: ${transcript}`)
// Generate a response based on the transcript
const { text } = await voiceAgent.generate(transcript)
如需 OpenAI Voice Provider 的詳細資訊,請參閱 OpenAI Voice 參考文件。
import { createReadStream } from 'fs'
import { Agent } from '@mastra/core/agent'
import { AzureVoice } from '@mastra/voice-azure'
import { createReadStream } from 'fs'
const voiceAgent = new Agent({
id: 'voice-agent',
name: 'Voice Agent',
instructions: 'You are a voice assistant that can help users with their tasks.',
model: 'openai/gpt-5.6-sol',
voice: new AzureVoice(),
})
// Use an audio file from a URL
const audioStream = await createReadStream('./how_can_i_help_you.mp3')
// Convert audio to text
const transcript = await voiceAgent.voice.listen(audioStream)
console.log(`User said: ${transcript}`)
// Generate a response based on the transcript
const { text } = await voiceAgent.generate(transcript)
如需 Azure Voice Provider 的詳細資訊,請參閱 Azure Voice 參考文件。
import { Agent } from '@mastra/core/agent'
import { ElevenLabsVoice } from '@mastra/voice-elevenlabs'
import { createReadStream } from 'fs'
const voiceAgent = new Agent({
id: 'voice-agent',
name: 'Voice Agent',
instructions: 'You are a voice assistant that can help users with their tasks.',
model: 'openai/gpt-5.6-sol',
voice: new ElevenLabsVoice(),
})
// Use an audio file from a URL
const audioStream = await createReadStream('./how_can_i_help_you.mp3')
// Convert audio to text
const transcript = await voiceAgent.voice.listen(audioStream)
console.log(`User said: ${transcript}`)
// Generate a response based on the transcript
const { text } = await voiceAgent.generate(transcript)
如需 ElevenLabs Voice Provider 的詳細資訊,請參閱 ElevenLabs Voice 參考文件。
import { Agent } from '@mastra/core/agent'
import { GoogleVoice } from '@mastra/voice-google'
import { createReadStream } from 'fs'
const voiceAgent = new Agent({
id: 'voice-agent',
name: 'Voice Agent',
instructions: 'You are a voice assistant that can help users with their tasks.',
model: 'openai/gpt-5.6-sol',
voice: new GoogleVoice(),
})
// Use an audio file from a URL
const audioStream = await createReadStream('./how_can_i_help_you.mp3')
// Convert audio to text
const transcript = await voiceAgent.voice.listen(audioStream)
console.log(`User said: ${transcript}`)
// Generate a response based on the transcript
const { text } = await voiceAgent.generate(transcript)
如需 Google Voice Provider 的詳細資訊,請參閱 Google Voice 參考文件。
import { Agent } from '@mastra/core/agent'
import { CloudflareVoice } from '@mastra/voice-cloudflare'
import { createReadStream } from 'fs'
const voiceAgent = new Agent({
id: 'voice-agent',
name: 'Voice Agent',
instructions: 'You are a voice assistant that can help users with their tasks.',
model: 'openai/gpt-5.6-sol',
voice: new CloudflareVoice(),
})
// Use an audio file from a URL
const audioStream = await createReadStream('./how_can_i_help_you.mp3')
// Convert audio to text
const transcript = await voiceAgent.voice.listen(audioStream)
console.log(`User said: ${transcript}`)
// Generate a response based on the transcript
const { text } = await voiceAgent.generate(transcript)
如需 Cloudflare Voice Provider 的詳細資訊,請參閱 Cloudflare Voice 參考文件。
import { Agent } from '@mastra/core/agent'
import { DeepgramVoice } from '@mastra/voice-deepgram'
import { createReadStream } from 'fs'
const voiceAgent = new Agent({
id: 'voice-agent',
name: 'Voice Agent',
instructions: 'You are a voice assistant that can help users with their tasks.',
model: 'openai/gpt-5.6-sol',
voice: new DeepgramVoice(),
})
// Use an audio file from a URL
const audioStream = await createReadStream('./how_can_i_help_you.mp3')
// Convert audio to text
const transcript = await voiceAgent.voice.listen(audioStream)
console.log(`User said: ${transcript}`)
// Generate a response based on the transcript
const { text } = await voiceAgent.generate(transcript)
如需 Deepgram Voice Provider 的詳細資訊,請參閱 Deepgram Voice 參考文件。
import { Agent } from '@mastra/core/agent'
import { InworldVoice } from '@mastra/voice-inworld'
import { createReadStream } from 'fs'
const voiceAgent = new Agent({
id: 'voice-agent',
name: 'Voice Agent',
instructions: 'You are a voice assistant that can help users with their tasks.',
model: 'openai/gpt-5.6-sol',
voice: new InworldVoice(),
})
// Use an audio file from a URL
const audioStream = await createReadStream('./how_can_i_help_you.mp3')
// Convert audio to text
const transcript = await voiceAgent.voice.listen(audioStream)
console.log(`User said: ${transcript}`)
// Generate a response based on the transcript
const { text } = await voiceAgent.generate(transcript)
如需 Inworld Voice Provider 的詳細資訊,請參閱 Inworld Voice 參考文件。
import { Agent } from '@mastra/core/agent'
import { SarvamVoice } from '@mastra/voice-sarvam'
import { createReadStream } from 'fs'
const voiceAgent = new Agent({
id: 'voice-agent',
name: 'Voice Agent',
instructions: 'You are a voice assistant that can help users with their tasks.',
model: 'openai/gpt-5.6-sol',
voice: new SarvamVoice(),
})
// Use an audio file from a URL
const audioStream = await createReadStream('./how_can_i_help_you.mp3')
// Convert audio to text
const transcript = await voiceAgent.voice.listen(audioStream)
console.log(`User said: ${transcript}`)
// Generate a response based on the transcript
const { text } = await voiceAgent.generate(transcript)
如需 Sarvam Voice Provider 的詳細資訊,請參閱 Sarvam Voice 參考文件。
語音轉語音(STS)「語音轉語音(STS)」的直接連結
運用語音轉語音功能建立對話體驗。統一 API 能讓使用者與 AI Agent 進行即時語音互動。 如需詳細設定選項與進階功能,請參閱語音轉語音。
- OpenAI
- AWS Nova Sonic
- Inworld Realtime
- xAI
import { Agent } from '@mastra/core/agent'
import { playAudio, getMicrophoneStream } from '@mastra/node-audio'
import { OpenAIRealtimeVoice } from '@mastra/voice-openai-realtime'
const voiceAgent = new Agent({
id: 'voice-agent',
name: 'Voice Agent',
instructions: 'You are a voice assistant that can help users with their tasks.',
model: 'openai/gpt-5.6-sol',
voice: new OpenAIRealtimeVoice(),
})
// Listen for agent audio responses
voiceAgent.voice.on('speaker', ({ audio }) => {
playAudio(audio)
})
// Initiate the conversation
await voiceAgent.voice.speak('How can I help you today?')
// Send continuous audio from the microphone
const micStream = getMicrophoneStream()
await voiceAgent.voice.send(micStream)
如需 OpenAI Voice Provider 的詳細資訊,請參閱 OpenAI Voice 參考文件。
import { Agent } from '@mastra/core/agent'
import { playAudio, getMicrophoneStream } from '@mastra/node-audio'
import { GeminiLiveVoice } from '@mastra/voice-google-gemini-live'
const voiceAgent = new Agent({
id: 'voice-agent',
name: 'Voice Agent',
instructions: 'You are a voice assistant that can help users with their tasks.',
model: 'openai/gpt-5.6-sol',
voice: new GeminiLiveVoice({
// Live API mode
apiKey: process.env.GOOGLE_API_KEY,
model: 'gemini-2.0-flash-exp',
speaker: 'Puck',
debug: true,
// Vertex AI alternative:
// vertexAI: true,
// project: 'your-gcp-project',
// location: 'us-central1',
// serviceAccountKeyFile: '/path/to/service-account.json',
}),
})
// Connect before using speak/send
await voiceAgent.voice.connect()
// Listen for agent audio responses
voiceAgent.voice.on('speaker', ({ audio }) => {
playAudio(audio)
})
// Listen for text responses and transcriptions
voiceAgent.voice.on('writing', ({ text, role }) => {
console.log(`${role}: ${text}`)
})
// Initiate the conversation
await voiceAgent.voice.speak('How can I help you today?')
// Send continuous audio from the microphone
const micStream = getMicrophoneStream()
await voiceAgent.voice.send(micStream)
如需 Google Gemini Live Voice Provider 的詳細資訊,請參閱 Google Gemini Live 參考文件。
import { Agent } from '@mastra/core/agent'
import { playAudio, getMicrophoneStream } from '@mastra/node-audio'
import { NovaSonicVoice } from '@mastra/voice-aws-nova-sonic'
const voiceAgent = new Agent({
id: 'voice-agent',
name: 'Voice Agent',
instructions: 'You are a voice assistant that can help users with their tasks.',
model: 'openai/gpt-5.6-sol',
voice: new NovaSonicVoice({
region: 'us-east-1',
speaker: 'matthew',
// Static credentials are optional. The default AWS credential
// provider chain is used when none are passed.
}),
})
// Connect before using speak/send
await voiceAgent.voice.connect()
// Listen for assistant audio (Int16Array PCM)
voiceAgent.voice.on('speaking', ({ audioData }) => {
if (audioData) playAudio(audioData)
})
// Listen for transcribed text
voiceAgent.voice.on('writing', ({ text, role }) => {
console.log(`${role}: ${text}`)
})
// Initiate the conversation
await voiceAgent.voice.speak('How can I help you today?')
// Send continuous audio from the microphone
const micStream = getMicrophoneStream()
await voiceAgent.voice.send(micStream)
如需 AWS Nova Sonic Voice Provider 的詳細資訊,請參閱 AWS Nova Sonic 參考文件。
import { Agent } from '@mastra/core/agent'
import { playAudio, getMicrophoneStream } from '@mastra/node-audio'
import { InworldRealtimeVoice } from '@mastra/voice-inworld'
const voiceAgent = new Agent({
id: 'voice-agent',
name: 'Voice Agent',
instructions: 'You are a voice assistant that can help users with their tasks.',
model: 'openai/gpt-5.6-sol',
voice: new InworldRealtimeVoice({
apiKey: process.env.INWORLD_API_KEY,
model: 'inworld/models/gemma-4-26b-a4b-it',
speaker: 'Sarah',
}),
})
// Connect before using speak/send
await voiceAgent.voice.connect()
// Listen for agent audio (PCM stream)
voiceAgent.voice.on('speaker', stream => {
playAudio(stream)
})
// Listen for text responses and transcriptions
voiceAgent.voice.on('writing', ({ text, role }) => {
console.log(`${role}: ${text}`)
})
// Initiate the conversation
await voiceAgent.voice.speak('How can I help you today?')
// Send continuous audio from the microphone
const micStream = getMicrophoneStream()
await voiceAgent.voice.send(micStream)
如需 Inworld Realtime Voice Provider 的詳細資訊,請參閱 Inworld Realtime 參考文件。
import { Agent } from '@mastra/core/agent'
import { playAudio, getMicrophoneStream } from '@mastra/node-audio'
import { XAIRealtimeVoice } from '@mastra/voice-xai-realtime'
const voiceAgent = new Agent({
id: 'voice-agent',
name: 'Voice Agent',
instructions: 'You are a voice assistant that can help users with their tasks.',
model: 'xai/grok-4.3',
voice: new XAIRealtimeVoice({
apiKey: process.env.XAI_API_KEY,
model: 'grok-voice-think-fast-1.0',
speaker: 'eve',
turnDetection: { type: 'server_vad' },
}),
})
// Connect before using speak/send
await voiceAgent.voice.connect()
// Listen for agent audio responses
voiceAgent.voice.on('speaker', audioStream => {
playAudio(audioStream)
})
// Listen for text responses and transcriptions
voiceAgent.voice.on('writing', ({ text, role }) => {
console.log(`${role}: ${text}`)
})
// Initiate the conversation
await voiceAgent.voice.speak('How can I help you today?')
// Send continuous audio from the microphone
const micStream = getMicrophoneStream()
await voiceAgent.voice.send(micStream)
如需 xAI Voice Provider 的詳細資訊,請參閱 xAI Realtime Voice 參考文件。
即時 Voice「即時 Voice」的直接連結
在瀏覽器或電話上執行使用者可隨時插話的即時通話。Mastra 將音訊迴圈交由 LiveKit 處理語音活動偵測、語意回合偵測與插話,而 Agent 則使用自己的模型、Tool 與 Memory 產生每次回應。設定方式與選項請參閱即時 Voice。
Voice 設定「Voice 設定」的直接連結
每個 Voice Provider 都能設定不同模型與選項。以下列出所有支援 Provider 的詳細設定選項:
- OpenAI
- Azure
- ElevenLabs
- PlayAI
- Cloudflare
- Deepgram
- Inworld
- Speechify
- Sarvam
- Murf
- OpenAI Realtime
- xAI Realtime
- Google Gemini Live
- AWS Nova Sonic
- Inworld Realtime
- AI SDK
// OpenAI Voice Configuration
const voice = new OpenAIVoice({
speechModel: {
name: 'gpt-3.5-turbo', // Example model name
apiKey: process.env.OPENAI_API_KEY,
language: 'en-US', // Language code
voiceType: 'neural', // Type of voice model
},
listeningModel: {
name: 'whisper-1', // Example model name
apiKey: process.env.OPENAI_API_KEY,
language: 'en-US', // Language code
format: 'wav', // Audio format
},
speaker: 'alloy', // Example speaker name
})
如需 OpenAI Voice Provider 的詳細資訊,請參閱 OpenAI Voice 參考文件。
// Azure Voice Configuration
const voice = new AzureVoice({
speechModel: {
name: 'en-US-JennyNeural', // Example model name
apiKey: process.env.AZURE_SPEECH_KEY,
region: process.env.AZURE_SPEECH_REGION,
language: 'en-US', // Language code
style: 'cheerful', // Voice style
pitch: '+0Hz', // Pitch adjustment
rate: '1.0', // Speech rate
},
listeningModel: {
name: 'en-US', // Example model name
apiKey: process.env.AZURE_SPEECH_KEY,
region: process.env.AZURE_SPEECH_REGION,
format: 'simple', // Output format
},
})
如需 Azure Voice Provider 的詳細資訊,請參閱 Azure Voice 參考文件。
// ElevenLabs Voice Configuration
const voice = new ElevenLabsVoice({
speechModel: {
voiceId: 'your-voice-id', // Example voice ID
model: 'eleven_multilingual_v2', // Example model name
apiKey: process.env.ELEVENLABS_API_KEY,
language: 'en', // Language code
emotion: 'neutral', // Emotion setting
},
// ElevenLabs may not have a separate listening model
})
如需 ElevenLabs Voice Provider 的詳細資訊,請參閱 ElevenLabs Voice 參考文件。
// PlayAI Voice Configuration
const voice = new PlayAIVoice({
speechModel: {
name: 'playai-voice', // Example model name
speaker: 'emma', // Example speaker name
apiKey: process.env.PLAYAI_API_KEY,
language: 'en-US', // Language code
speed: 1.0, // Speech speed
},
// PlayAI may not have a separate listening model
})
如需 PlayAI Voice Provider 的詳細資訊,請參閱 PlayAI Voice 參考文件。
// Google Voice Configuration
const voice = new GoogleVoice({
speechModel: {
name: 'en-US-Studio-O', // Example model name
apiKey: process.env.GOOGLE_API_KEY,
languageCode: 'en-US', // Language code
gender: 'FEMALE', // Voice gender
speakingRate: 1.0, // Speaking rate
},
listeningModel: {
name: 'en-US', // Example model name
sampleRateHertz: 16000, // Sample rate
},
})
如需 Google Voice Provider 的詳細資訊,請參閱 Google Voice 參考文件。
// Cloudflare Voice Configuration
const voice = new CloudflareVoice({
speechModel: {
name: 'cloudflare-voice', // Example model name
accountId: process.env.CLOUDFLARE_ACCOUNT_ID,
apiToken: process.env.CLOUDFLARE_API_TOKEN,
language: 'en-US', // Language code
format: 'mp3', // Audio format
},
// Cloudflare may not have a separate listening model
})
如需 Cloudflare Voice Provider 的詳細資訊,請參閱 Cloudflare Voice 參考文件。
// Deepgram Voice Configuration
const voice = new DeepgramVoice({
speechModel: {
name: 'nova-2', // Example model name
speaker: 'aura-english-us', // Example speaker name
apiKey: process.env.DEEPGRAM_API_KEY,
language: 'en-US', // Language code
tone: 'formal', // Tone setting
},
listeningModel: {
name: 'nova-2', // Example model name
format: 'flac', // Audio format
},
})
如需 Deepgram Voice Provider 的詳細資訊,請參閱 Deepgram Voice 參考文件。
// Inworld Voice Configuration
const voice = new InworldVoice({
speechModel: {
name: 'inworld-tts-2',
apiKey: process.env.INWORLD_API_KEY,
},
listeningModel: {
name: 'groq/whisper-large-v3',
apiKey: process.env.INWORLD_API_KEY,
},
speaker: 'Dennis',
audioEncoding: 'MP3',
sampleRateHertz: 48000,
language: 'en-US',
})
// Per-call options: `deliveryMode` is honored only by `inworld-tts-2`.
const audioStream = await voice.speak('Hello!', {
deliveryMode: 'BALANCED', // 'STABLE' | 'BALANCED' | 'CREATIVE'
language: 'en-US', // BCP-47 per-call override
})
如需 Inworld Voice Provider 的詳細資訊,請參閱 Inworld Voice 參考文件。
// Speechify Voice Configuration
const voice = new SpeechifyVoice({
speechModel: {
name: 'speechify-voice', // Example model name
speaker: 'matthew', // Example speaker name
apiKey: process.env.SPEECHIFY_API_KEY,
language: 'en-US', // Language code
speed: 1.0, // Speech speed
},
// Speechify may not have a separate listening model
})
如需 Speechify Voice Provider 的詳細資訊,請參閱 Speechify Voice 參考文件。
// Sarvam Voice Configuration
const voice = new SarvamVoice({
speechModel: {
model: 'bulbul:v3', // TTS model (bulbul:v2 or bulbul:v3)
apiKey: process.env.SARVAM_API_KEY,
language: 'en-IN', // BCP-47 language code
},
listeningModel: {
model: 'saarika:v2.5', // STT model (saarika:v2.5 or saaras:v3)
apiKey: process.env.SARVAM_API_KEY,
},
speaker: 'shubh', // Default bulbul:v3 speaker
})
如需 Sarvam Voice Provider 的詳細資訊,請參閱 Sarvam Voice 參考文件。
// Murf Voice Configuration
const voice = new MurfVoice({
speechModel: {
name: 'murf-voice', // Example model name
apiKey: process.env.MURF_API_KEY,
language: 'en-US', // Language code
emotion: 'happy', // Emotion setting
},
// Murf may not have a separate listening model
})
如需 Murf Voice Provider 的詳細資訊,請參閱 Murf Voice 參考文件。
// OpenAI Realtime Voice Configuration
const voice = new OpenAIRealtimeVoice({
speechModel: {
name: 'gpt-3.5-turbo', // Example model name
apiKey: process.env.OPENAI_API_KEY,
language: 'en-US', // Language code
},
listeningModel: {
name: 'whisper-1', // Example model name
apiKey: process.env.OPENAI_API_KEY,
format: 'ogg', // Audio format
},
speaker: 'alloy', // Example speaker name
})
如需 OpenAI Realtime Voice Provider 的詳細資訊,請參閱 OpenAI Realtime Voice 參考文件。
// xAI Realtime Voice Configuration
const voice = new XAIRealtimeVoice({
apiKey: process.env.XAI_API_KEY,
model: 'grok-voice-think-fast-1.0',
speaker: 'eve',
instructions: 'You are a concise voice assistant.',
turnDetection: {
type: 'server_vad',
threshold: 0.85,
silence_duration_ms: 1000,
prefix_padding_ms: 333,
},
audio: {
input: { format: { type: 'audio/pcm', rate: 24000 } },
output: { format: { type: 'audio/pcm', rate: 24000 } },
},
serverTools: [
{ type: 'web_search' },
{
type: 'mcp',
server_url: 'https://mcp.example.com/mcp',
server_label: 'business-tools',
},
],
})
如需 xAI realtime Voice Provider 的詳細資訊,請參閱 xAI Realtime Voice 參考文件。
// Google Gemini Live Voice Configuration
const voice = new GeminiLiveVoice({
speechModel: {
name: 'gemini-2.0-flash-exp', // Example model name
apiKey: process.env.GOOGLE_API_KEY,
},
speaker: 'Puck', // Example speaker name
// Google Gemini Live is a realtime bidirectional API without separate speech and listening models
})
如需 Google Gemini Live Voice Provider 的詳細資訊,請參閱 Google Gemini Live 參考文件。
// AWS Nova Sonic Voice Configuration
const voice = new NovaSonicVoice({
region: 'us-east-1',
speaker: 'matthew',
sessionConfig: {
inferenceConfiguration: {
temperature: 0.7,
maxTokens: 1024,
},
turnDetectionConfiguration: {
endpointingSensitivity: 'MEDIUM',
},
},
// AWS Nova Sonic is a realtime bidirectional API without separate speech and listening models
})
如需 AWS Nova Sonic Voice Provider 的詳細資訊,請參閱 AWS Nova Sonic 參考文件。
// Inworld Realtime Voice Configuration
const voice = new InworldRealtimeVoice({
apiKey: process.env.INWORLD_API_KEY,
model: 'inworld/models/gemma-4-26b-a4b-it',
speaker: 'Sarah',
// Typed Inworld realtime knobs (semantic VAD, playback speed, MCP tool routing, ...)
session: {
audio: {
output: { speed: 1.1 },
input: { turn_detection: { type: 'semantic_vad', eagerness: 'high' } },
},
},
})
如需 Inworld Realtime Voice Provider 的詳細資訊,請參閱 Inworld Realtime 參考文件。
// AI SDK Voice Configuration
import { CompositeVoice } from '@mastra/core/voice'
import { openai } from '@ai-sdk/openai'
import { elevenlabs } from '@ai-sdk/elevenlabs'
// Use AI SDK models directly - no need to install separate packages
const voice = new CompositeVoice({
input: openai.transcription('whisper-1'), // AI SDK transcription
output: elevenlabs.speech('eleven_turbo_v2'), // AI SDK speech
})
// Works seamlessly with your agent
const voiceAgent = new Agent({
id: 'aisdk-voice-agent',
name: 'AI SDK Voice Agent',
instructions: 'You are a helpful assistant with voice capabilities.',
model: 'openai/gpt-5.6-sol',
voice,
})
使用多個 Voice Provider「使用多個 Voice Provider」的直接連結
此範例示範如何在 Mastra 中建立及使用兩個不同的 Voice Provider:使用 OpenAI 進行語音轉文字(STT),並使用 PlayAI 進行文字轉語音(TTS)。
首先建立 Voice Provider 執行個體,並加入所需設定。
import { OpenAIVoice } from '@mastra/voice-openai'
import { PlayAIVoice } from '@mastra/voice-playai'
import { CompositeVoice } from '@mastra/core/voice'
import { playAudio, getMicrophoneStream } from '@mastra/node-audio'
// Initialize OpenAI voice for STT
const input = new OpenAIVoice({
listeningModel: {
name: 'whisper-1',
apiKey: process.env.OPENAI_API_KEY,
},
})
// Initialize PlayAI voice for TTS
const output = new PlayAIVoice({
speechModel: {
name: 'playai-voice',
apiKey: process.env.PLAYAI_API_KEY,
},
})
// Combine the providers using CompositeVoice
const voice = new CompositeVoice({
input,
output,
})
// Implement voice interactions using the combined voice provider
const audioStream = getMicrophoneStream() // Assume this function gets audio input
const transcript = await voice.listen(audioStream)
// Log the transcribed text
console.log('Transcribed text:', transcript)
// Convert text to speech
const responseAudio = await voice.speak(`You said: ${transcript}`, {
speaker: 'default', // Optional: specify a speaker,
responseFormat: 'wav', // Optional: specify a response format
})
// Play the audio response
playAudio(responseAudio)
使用 AI SDK 模型 Provider「使用 AI SDK 模型 Provider」的直接連結
你也可以直接搭配 CompositeVoice 使用 AI SDK 模型:
import { CompositeVoice } from '@mastra/core/voice'
import { openai } from '@ai-sdk/openai'
import { elevenlabs } from '@ai-sdk/elevenlabs'
import { playAudio, getMicrophoneStream } from '@mastra/node-audio'
// Use AI SDK models directly - no provider setup needed
const voice = new CompositeVoice({
input: openai.transcription('whisper-1'), // AI SDK transcription
output: elevenlabs.speech('eleven_turbo_v2'), // AI SDK speech
})
// Works the same way as Mastra providers
const audioStream = getMicrophoneStream()
const transcript = await voice.listen(audioStream)
console.log('Transcribed text:', transcript)
// Convert text to speech
const responseAudio = await voice.speak(`You said: ${transcript}`, {
speaker: 'Rachel', // ElevenLabs voice
})
playAudio(responseAudio)
你也可以混合使用 AI SDK 模型與 Mastra Provider:
import { CompositeVoice } from '@mastra/core/voice'
import { PlayAIVoice } from '@mastra/voice-playai'
import { groq } from '@ai-sdk/groq'
const voice = new CompositeVoice({
input: groq.transcription('whisper-large-v3'), // AI SDK for STT
output: new PlayAIVoice(), // Mastra provider for TTS
})
如需 CompositeVoice 的詳細資訊,請參閱 CompositeVoice 參考文件。