- Text-to-Speech
GPT Audio is an OpenAI model that can understand and generate natural-sounding speech in real time. It is notable for combining strong language understanding with fast,…
Powered by Mistral
Voxtral Mini TTS is Mistral’s 4B-parameter text-to-speech model that generates expressive, low-latency speech and supports multilingual, zero-shot voice cloning. It is available via the Mistral API and as open weights for self-hosting.
Output tokens per second · Higher is better
Seconds · Lower is better
USD per 1M tokens (blended) · Lower is better
About the model
Voxtral Mini TTS is a 4B-parameter text-to-speech model from Mistral that converts text into natural, expressive speech with multilingual support and voice cloning from very short audio samples. It is mainly used to build voice agents and assistants that respond in real time with low-latency audio, and to generate high-quality synthetic voices for applications like content narration, product voices, and accessibility tools. It also serves use cases that require cloning or reusing consistent speaker identities across many utterances, such as branded voice experiences and character dialogue. The model is part of Mistral’s Voxtral audio family, alongside Voxtral Mini and Voxtral Small transcription and audio-understanding models.
Model capabilities
Generates natural-sounding speech audio from written text, suitable for dialogue, narration, and interface responses in multiple scenarios.
Produces speech tailored for interactive assistants, enabling clear, responsive spoken dialogue aligned with conversational AI systems’ outputs.
Supports speech generation in multiple languages, allowing applications to vocalize content for diverse linguistic audiences and use cases.
Can power screen readers or accessibility tools by converting on-screen text into intelligible, continuous spoken audio output.
Provides synthesized voices for videos, podcasts, or interactive media, enabling scalable voiceover creation without human recording sessions.
Use cases
Transparent pricing
Up to ~70% cheaper and lower-latency than comparable TTS APIs
| Provider | Region | Latency | Throughput | Uptime | Input ($/1M) | Output ($/1M) | Context |
|---|---|---|---|---|---|---|---|
| LLM API BEST | Global | 80ms | 120 req/s | 99.99% | $0.004/min | $0.004/min | ~15 min audio |
| Mistral | EU West | ~140ms | ~45 req/s | ~99.9% | ~$0.010/min | ~$0.010/min | ~10 min audio |
| OpenAI | Global | ~150ms | ~60 req/s | 99.9% | ~$0.015/min | ~$0.015/min | ~15 min audio |
| Azure AI Speech | Global | ~180ms | ~80 req/s | 99.9% | ~$0.016/min | ~$0.016/min | ~10 min audio |
| Google Cloud Text-to-Speech | Global | ~170ms | ~70 req/s | 99.9% | ~$0.014/min | ~$0.014/min | ~10 min audio |
Performance benchmarks
| Metric | Voxtral Mini TTS | OpenAI gpt-4o-mini TTS | Google Chirp TTS (small) |
|---|---|---|---|
| Avg Latency | ~180ms | ~200ms | ~220ms |
| Languages Supported | ~25 | ~30 | ~20 |
| Price per 1M chars | ~$0.70 | ~$1.00 | ~$0.80 |
| Max Input Length | ~4K chars | ~8K chars | ~5K chars |
| Sample Rate | 24 kHz | 24 kHz | 22.05 kHz |
| Voices / Styles | ~20 | ~30 | ~15 |
| Uptime | 99.9% | 99.9% | 99.5% |
30-day usage via LLM API
Architecture & Integration
One unified API. Every major model. Built-in reliability, cost control, and observability.
Automatically route each request to the best-fit model across providers based on cost, latency, or quality—without changing your code or client integration.
One endpoint, every model.Define cost ceilings and model preferences, then let LLM.API optimize per-call spend so you can scale usage without surprise bills or manual tuning.
More usage, less spend.When a provider times out, errors, or rate-limits, LLM.API seamlessly retries on backup models so your production flows stay reliable and resilient.
No single point of failure.Get unified logs, metrics, traces, and payload samples across all models and providers, making debugging, performance tuning, and governance radically simpler.
See every token, everywhere.Describe tasks like chat, generation, tools, or RAG once and let LLM.API translate them into provider-specific calls, so you avoid brittle model-specific code.
Code to tasks, not models.Send massive batches of prompts through a single API call, with automatic chunking, retries, and concurrency controls to maximize throughput across providers.
Process thousands in one go.Decision guide
FAQ
Voxtral Mini TTS is a Mistral text-to-speech model focused on fast, lightweight voice synthesis for applications that need low-latency audio generation.
It is best for real-time or near real-time speech generation in interactive apps, voice assistants, and low-resource environments.
Pricing is usage-based per generated character or token, with exact rates defined in the LLM.API model pricing table.
The model accepts short to moderate text prompts suitable for speech synthesis, with exact character limits determined by LLM.API configuration.
Voxtral Mini TTS is optimized for low latency, typically returning audio quickly enough for responsive user experiences in interactive applications.
It supports text-to-speech only, taking text input and returning synthesized audio output.
Call the LLM.API generation endpoint with the Voxtral Mini TTS model identifier, passing text input and any audio configuration parameters supported by the API.
Compared to larger TTS models, it trades some maximum quality and configurability for lower cost, faster inference, and smaller resource requirements.
Limitations can include less natural prosody on complex texts, language coverage constraints, and quality degradation on very long inputs.
Streaming availability depends on LLM.API’s implementation; check the streaming or response_mode options for this specific model.
Compare
GPT Audio is an OpenAI model that can understand and generate natural-sounding speech in real time. It is notable for combining strong language understanding with fast,…
GPT-4o Mini TTS is a text-to-speech variant of OpenAI’s lightweight GPT-4o Mini model, designed to generate natural-sounding spoken audio from text with low latency and efficient…
Kokoro 82M is an open-weight, 82‑million‑parameter text‑to‑speech model from hexgrad that focuses on natural, multilingual speech with low latency and low resource usage. It is notable…