- Speech-to-Text
Whisper Large V3 is OpenAI’s large-scale speech recognition model designed for robust, multilingual transcription and translation. It is notable for high accuracy, support for many languages,…
Powered by Mistral
Voxtral Mini Transcribe is a speech-to-text model from Mistral focused on lightweight, efficient audio transcription. It is designed to provide accurate transcriptions while being small and fast enough for resource-constrained environments.
Output tokens per second · Higher is better
Seconds · Lower is better
USD per 1M tokens (blended) · Lower is better
About the model
Voxtral Mini Transcribe is a compact automatic speech recognition (ASR) model by Mistral for converting spoken audio into text. It is mainly used for real-time or near real-time transcription of voice recordings, calls, and meetings. It is also suitable for integrating speech input into applications where low latency and low computational overhead are important. It belongs to Mistral’s Voxtral family of ASR models, which are optimized for practical deployment and efficiency.
Model capabilities
Converts spoken audio into accurate text, supporting various speakers and recording conditions for transcription and note-taking use cases.
Processes streaming audio input to produce near real-time text transcripts suitable for live captions and interactive applications.
Transcribes speech from multiple supported languages, enabling cross-lingual audio processing and global applications requiring language-aware transcription.
Produces structured, readable transcripts suitable for conversational contexts, meetings, and interviews, preserving speaker turns when available.
Generates text closely aligned with input audio segments, facilitating downstream search, navigation, and timestamp-based audio indexing.
Use cases
Transparent pricing
LLM API offers the lowest per‑minute pricing and best SLAs for Voxtral-class transcription.
| Provider | Region | Latency | Throughput | Uptime | Input ($/1M) | Output ($/1M) | Context |
|---|---|---|---|---|---|---|---|
| LLM API BEST | Global | ~180ms | ~120 audio min/min | 99.99% | $0.004/min | $0.000/min | ~120 min audio |
| Mistral | EU West | ~220ms | ~80 audio min/min | ~99.9% | ~$0.006/min | $0.000/min | ~60 min audio |
| OpenAI | Global | ~250ms | ~90 audio min/min | ~99.9% | ~$0.006/min | $0.000/min | ~60 min audio |
| Azure AI | Global | ~260ms | ~70 audio min/min | ~99.9% | ~$0.007/min | $0.000/min | ~60 min audio |
| Google Cloud | Global | ~240ms | ~75 audio min/min | ~99.9% | ~$0.007/min | $0.000/min | ~60 min audio |
Performance benchmarks
| Metric | Voxtral Mini Transcribe | OpenAI Whisper v3 Small | Google Speech-to-Text v2 |
|---|---|---|---|
| Avg Latency | ~350ms | ~400ms | ~450ms |
| Languages Supported | ~100 | ~100 | ~70 |
| Price per Minute | $0.006 | $0.006 | $0.009 |
| Max Duration | 2h | 12h | 3h |
| Accuracy (WER) | ~7% | ~6% | ~8% |
| Uptime | 99.9% | 99.9% | 99.9% |
30-day usage via LLM API
Architecture & Integration
One unified API. Every major model. Built-in reliability, cost control, and observability.
Dynamically route each request to the best model across providers based on latency, cost, and capability—no client changes, just smarter traffic.
One endpoint, many modelsAutomatically balance quality and spend with fine-grained controls, price-aware routing, and per-project limits so teams ship fast without surprise cloud bills.
Control spend, not speedDefine provider and model fallback chains that trigger on errors, timeouts, or degraded quality so your AI features stay online even when vendors don’t.
Fail soft, stay onlineTrace every call across providers with unified logs, metrics, and structured traces so you can debug latency spikes and failures in minutes, not days.
See every token hopDescribe work as tasks—chat, tools, RAG, agents—and let LLM.API pick the right models, prompts, and configs so you focus on product, not plumbing.
Program tasks, not modelsRun large-scale generations, evaluations, and data labeling as batched jobs with built-in retries, concurrency controls, and cost tracking from a single API.
Scale to millions of callsDecision guide
FAQ
Voxtral Mini Transcribe is a speech-to-text model by Mistral optimized for fast, low-cost audio transcription via the LLM.API gateway.
Voxtral Mini Transcribe supports audio input and returns transcribed text output; it does not process images, video, or arbitrary text prompts directly.
You call the unified LLM.API endpoint with the model name 'mistral:voxtral-mini-transcribe' and provide your audio data and parameters in the request body.
Voxtral Mini Transcribe is best for real-time or batch transcription of spoken content such as meetings, calls, podcasts, and voice notes.
Typical end-to-end latency is a few seconds for short audio clips, depending on audio length, network conditions, and your region.
Voxtral Mini Transcribe is limited by maximum audio duration per request, so long recordings should be chunked into smaller segments client-side.
Voxtral Mini Transcribe is billed per unit of processed audio, with exact per-minute or per-second rates defined in the LLM.API pricing page.
Compared to larger models, Voxtral Mini Transcribe generally offers lower cost and latency at the expense of slightly lower accuracy on challenging audio.
If enabled by LLM.API, you can stream audio chunks to Voxtral Mini Transcribe and receive partial transcripts incrementally.
Voxtral Mini Transcribe may struggle with heavy background noise, strong accents, overlapping speakers, low-bitrate audio, or domain-specific jargon without adaptation.
Compare
Whisper Large V3 is OpenAI’s large-scale speech recognition model designed for robust, multilingual transcription and translation. It is notable for high accuracy, support for many languages,…
Whisper Large V3 Turbo is OpenAI’s optimized, high‑speed variant of the Whisper Large V3 automatic speech recognition model, designed to provide fast transcriptions while preserving strong…
Chirp 3 is Google's latest-generation multilingual speech and audio model, available through Google Cloud for high-accuracy transcription and natural-sounding text-to-speech. It is notable for its improved…