본문으로 건너뛰기

마스트라의 목소리

Mastra의 음성 시스템은 음성 상호 작용을 위한 통합 인터페이스를 제공하여 애플리케이션에서 TTS(텍스트 음성 변환), STT(음성 변환) 및 실시간 STS(음성 변환) 기능을 활성화합니다.

Agent에게 음성 추가
Agent에게 음성 추가에 대한 직접 링크

다음을 사용하여 음성 제공자를 Agent에게 전달합니다.voice 속성을 사용합니다. 구성한 Provider에 따라 동일한 속성에서 텍스트 음성 변환(TTS), 음성 텍스트 변환(STT), 실시간 음성 간 변환(STS)을 지원합니다.

import { Agent } from '@mastra/core/agent'
import { OpenAIVoice } from '@mastra/voice-openai'

// Initialize OpenAI voice for TTS

const voiceAgent = new Agent({
id: 'voice-agent',
name: 'Voice Agent',
instructions: 'You are a voice assistant that can help users with their tasks.',
model: 'openai/gpt-5.6-sol',
voice: new OpenAIVoice(),
})

그런 다음 다음 음성 기능을 사용할 수 있습니다.

텍스트 음성 변환(TTS)
텍스트 음성 변환(TTS)에 대한 직접 링크

Mastra의 TTS 기능을 사용하여 Agent의 응답을 자연스러운 음성으로 변환하세요. OpenAI, ElevenLabs 등과 같은 여러 Provider 중에서 선택하세요.

자세한 구성 옵션과 고급 기능을 알아보려면 당사를 확인하세요.Text-to-Speech guide.

import { Agent } from '@mastra/core/agent'
import { OpenAIVoice } from '@mastra/voice-openai'
import { playAudio } from '@mastra/node-audio'

const voiceAgent = new Agent({
id: 'voice-agent',
name: 'Voice Agent',
instructions: 'You are a voice assistant that can help users with their tasks.',
model: 'openai/gpt-5.6-sol',
voice: new OpenAIVoice(),
})

const { text } = await voiceAgent.generate('What color is the sky?')

// Convert text to speech to an Audio Stream
const audioStream = await voiceAgent.voice.speak(text, {
speaker: 'default', // Optional: specify a speaker
responseFormat: 'wav', // Optional: specify a response format
})

playAudio(audioStream)

방문OpenAI Voice Reference 에서 OpenAI 음성 Provider에 관한 자세한 정보를 확인하세요.

음성을 텍스트로 변환(STT)
음성을 텍스트로 변환(STT)에 대한 직접 링크

OpenAI, ElevenLabs 등과 같은 Provider를 사용하여 음성 콘텐츠를 텍스트로 변환하세요. 자세한 구성 옵션 등을 확인하려면 다음을 확인하세요.Speech to Text.

다음에서 샘플 오디오 파일을 다운로드할 수 있습니다.here.


import { Agent } from '@mastra/core/agent'
import { OpenAIVoice } from '@mastra/voice-openai'
import { createReadStream } from 'fs'

const voiceAgent = new Agent({
id: 'voice-agent',
name: 'Voice Agent',
instructions: 'You are a voice assistant that can help users with their tasks.',
model: 'openai/gpt-5.6-sol',
voice: new OpenAIVoice(),
})

// Use an audio file from a URL
const audioStream = await createReadStream('./how_can_i_help_you.mp3')

// Convert audio to text
const transcript = await voiceAgent.voice.listen(audioStream)
console.log(`User said: ${transcript}`)

// Generate a response based on the transcript
const { text } = await voiceAgent.generate(transcript)

방문OpenAI Voice Reference 에서 OpenAI 음성 Provider에 관한 자세한 정보를 확인하세요.

음성 대 음성(STS)
음성 대 음성(STS)에 대한 직접 링크

음성 대 음성 변환 기능으로 대화 환경을 조성하세요. 통합 API를 사용하면 사용자와 AI Agent 간의 실시간 음성 상호작용이 가능합니다. 자세한 구성 옵션과 고급 기능을 확인하려면 다음을 확인하세요.Speech to Speech.

import { Agent } from '@mastra/core/agent'
import { playAudio, getMicrophoneStream } from '@mastra/node-audio'
import { OpenAIRealtimeVoice } from '@mastra/voice-openai-realtime'

const voiceAgent = new Agent({
id: 'voice-agent',
name: 'Voice Agent',
instructions: 'You are a voice assistant that can help users with their tasks.',
model: 'openai/gpt-5.6-sol',
voice: new OpenAIRealtimeVoice(),
})

// Listen for agent audio responses
voiceAgent.voice.on('speaker', ({ audio }) => {
playAudio(audio)
})

// Initiate the conversation
await voiceAgent.voice.speak('How can I help you today?')

// Send continuous audio from the microphone
const micStream = getMicrophoneStream()
await voiceAgent.voice.send(micStream)

방문OpenAI Voice Reference 에서 OpenAI 음성 Provider에 관한 자세한 정보를 확인하세요.

실시간 음성
실시간 음성에 대한 직접 링크

사용자가 브라우저나 전화를 통해 대화할 수 있는 실시간 통화를 실행하세요. Mastra는 음성 활동 감지, 의미론적 차례 감지 및 참여를 다루는 오디오 루프를 LiveKit에 전달하고 Agent는 자체 Model, Tool 및 Memory를 사용하여 각 응답을 생성합니다. 설정 및 구성 옵션을 확인하려면 다음을 확인하세요.Realtime voice.

음성 구성
음성 구성에 대한 직접 링크

각 음성 Provider는 다양한 Model과 옵션으로 구성될 수 있습니다. 다음은 지원되는 모든 공급자에 대한 자세한 구성 옵션입니다.

// OpenAI Voice Configuration
const voice = new OpenAIVoice({
speechModel: {
name: 'gpt-3.5-turbo', // Example model name
apiKey: process.env.OPENAI_API_KEY,
language: 'en-US', // Language code
voiceType: 'neural', // Type of voice model
},
listeningModel: {
name: 'whisper-1', // Example model name
apiKey: process.env.OPENAI_API_KEY,
language: 'en-US', // Language code
format: 'wav', // Audio format
},
speaker: 'alloy', // Example speaker name
})

방문OpenAI Voice Reference 에서 OpenAI 음성 Provider에 관한 자세한 정보를 확인하세요.

여러 음성 Provider 사용
여러 음성 Provider 사용에 대한 직접 링크

이 예에서는 Mastra에서 STT(음성-텍스트)용 OpenAI와 TTS(텍스트 음성 변환)용 PlayAI라는 두 가지 음성 공급자를 만들고 사용하는 방법을 보여줍니다.

필요한 구성을 사용하여 음성 공급자의 인스턴스를 만드는 것부터 시작하세요.

import { OpenAIVoice } from '@mastra/voice-openai'
import { PlayAIVoice } from '@mastra/voice-playai'
import { CompositeVoice } from '@mastra/core/voice'
import { playAudio, getMicrophoneStream } from '@mastra/node-audio'

// Initialize OpenAI voice for STT
const input = new OpenAIVoice({
listeningModel: {
name: 'whisper-1',
apiKey: process.env.OPENAI_API_KEY,
},
})

// Initialize PlayAI voice for TTS
const output = new PlayAIVoice({
speechModel: {
name: 'playai-voice',
apiKey: process.env.PLAYAI_API_KEY,
},
})

// Combine the providers using CompositeVoice
const voice = new CompositeVoice({
input,
output,
})

// Implement voice interactions using the combined voice provider
const audioStream = getMicrophoneStream() // Assume this function gets audio input
const transcript = await voice.listen(audioStream)

// Log the transcribed text
console.log('Transcribed text:', transcript)

// Convert text to speech
const responseAudio = await voice.speak(`You said: ${transcript}`, {
speaker: 'default', // Optional: specify a speaker,
responseFormat: 'wav', // Optional: specify a response format
})

// Play the audio response
playAudio(responseAudio)

AI SDK Model 공급자 사용
AI SDK Model 공급자 사용에 대한 직접 링크

AI SDK Model을 직접 사용할 수도 있습니다.CompositeVoice:

import { CompositeVoice } from '@mastra/core/voice'
import { openai } from '@ai-sdk/openai'
import { elevenlabs } from '@ai-sdk/elevenlabs'
import { playAudio, getMicrophoneStream } from '@mastra/node-audio'

// Use AI SDK models directly - no provider setup needed
const voice = new CompositeVoice({
input: openai.transcription('whisper-1'), // AI SDK transcription
output: elevenlabs.speech('eleven_turbo_v2'), // AI SDK speech
})

// Works the same way as Mastra providers
const audioStream = getMicrophoneStream()
const transcript = await voice.listen(audioStream)

console.log('Transcribed text:', transcript)

// Convert text to speech
const responseAudio = await voice.speak(`You said: ${transcript}`, {
speaker: 'Rachel', // ElevenLabs voice
})

playAudio(responseAudio)

AI SDK Model을 Mastra 공급자와 혼합할 수도 있습니다.

import { CompositeVoice } from '@mastra/core/voice'
import { PlayAIVoice } from '@mastra/voice-playai'
import { groq } from '@ai-sdk/groq'

const voice = new CompositeVoice({
input: groq.transcription('whisper-large-v3'), // AI SDK for STT
output: new PlayAIVoice(), // Mastra provider for TTS
})

CompositeVoice에 대한 자세한 내용은CompositeVoice Reference.

더 많은 리소스
더 많은 리소스에 대한 직접 링크