Inference
透過 Mastra 的模型路由器存取 9 個 Inference 模型。系統會自動使用 INFERENCE_API_KEY 環境變數完成身分驗證。
如需進一步了解,請參閱Inference 文件。
.env
INFERENCE_API_KEY=your-api-key
src/mastra/agents/my-agent.ts
import { Agent } from '@mastra/core/agent'
const agent = new Agent({
id: 'my-agent',
name: 'My Agent',
instructions: 'You are a helpful assistant',
model: 'inference/google/gemma-3',
})
// Generate a response
const response = await agent.generate('Hello!')
// Stream a response
const stream = await agent.stream('Tell me a story')
for await (const chunk of stream) {
console.log(chunk)
}
資訊
Mastra 使用與 OpenAI 相容的 /chat/completions endpoint。部分 Provider 特有功能可能無法使用,詳情請參閱Inference 文件。
模型「模型」的直接連結
| Model | Context | Tools | Reasoning | Image | Audio | Video | Input $/1M | Output $/1M |
|---|---|---|---|---|---|---|---|---|
inference/google/gemma-3 | 125K | $0.15 | $0.30 | |||||
inference/meta/llama-3.1-8b-instruct | 16K | $0.03 | $0.03 | |||||
inference/meta/llama-3.2-11b-vision-instruct | 16K | $0.06 | $0.06 | |||||
inference/meta/llama-3.2-1b-instruct | 16K | $0.01 | $0.01 | |||||
inference/meta/llama-3.2-3b-instruct | 16K | $0.02 | $0.02 | |||||
inference/mistral/mistral-nemo-12b-instruct | 16K | $0.04 | $0.10 | |||||
inference/osmosis/osmosis-structure-0.6b | 4K | $0.10 | $0.50 | |||||
inference/qwen/qwen-2.5-7b-vision-instruct | 125K | $0.20 | $0.20 | |||||
inference/qwen/qwen3-embedding-4b | 32K | $0.01 | — |
進階設定「進階設定」的直接連結
自訂標頭「自訂標頭」的直接連結
src/mastra/agents/my-agent.ts
const agent = new Agent({
id: 'custom-agent',
name: 'custom-agent',
model: {
url: 'https://inference.net/v1',
id: 'inference/google/gemma-3',
apiKey: process.env.INFERENCE_API_KEY,
headers: {
'X-Custom-Header': 'value',
},
},
})
動態選擇模型「動態選擇模型」的直接連結
src/mastra/agents/my-agent.ts
const agent = new Agent({
id: 'dynamic-agent',
name: 'Dynamic Agent',
model: ({ requestContext }) => {
const useAdvanced = requestContext.task === 'complex'
return useAdvanced ? 'inference/qwen/qwen3-embedding-4b' : 'inference/google/gemma-3'
},
})