- Instruction Following
Claude Opus 4.8 (Fast) is Anthropic’s flagship Claude Opus 4.8 model running in a special fast mode that delivers significantly higher output token throughput at premium…
Powered by Mistral
Voxtral Small 24B 2507 is a 24-billion-parameter audio-language model from Mistral that extends Mistral Small 3 with advanced speech understanding. It is notable for strong, cost-efficient performance on transcription, translation, and audio-informed text tasks across multiple languages.
Output tokens per second · Higher is better
Seconds · Lower is better
USD per 1M tokens (blended) · Lower is better
About the model
Voxtral Small 24B 2507 is an open-source speech understanding and language model from Mistral that combines text generation with state-of-the-art audio input capabilities. It is mainly used for high-quality speech transcription and translation directly from audio in many languages. It is also applied to tasks like Q&A, summarization, and general chat where audio context must be understood alongside text. It belongs to the Voxtral family and is built as an enhancement of the Mistral Small 3 series.
Model capabilities
Handles multi-turn text conversations with strong general reasoning, instruction following, and tool-use support in many domains and scenarios.
Transcribes spoken audio into accurate text using a dedicated speech transcription mode optimized for high-quality automatic speech recognition.
Performs speech-to-text translation across multiple languages, enabling multilingual audio translation and cross-lingual understanding within one model.
Analyzes audio beyond transcription, supporting audio-based question answering, summarization, and semantic comprehension of spoken content.
Integrates with various inference and observability platforms, supporting structured outputs, tools, and deployment in managed environments.
Use cases
Transparent pricing
LLM API offers the lowest cost and best performance for Voxtral Small–class 24B models.
| Provider | Region | Latency | Throughput | Uptime | Input ($/1M) | Output ($/1M) | Context |
|---|---|---|---|---|---|---|---|
| LLM API BEST | Global | 90ms | 80 tps | 99.99% | $0.40 | $0.40 | 256K |
| Mistral | EU West | ~140ms | ~45 tps | 99.9% | ~$0.60 | ~$0.60 | ~128K |
| Together AI | US East | ~160ms | ~40 tps | 99.9% | ~$0.55 | ~$0.55 | ~128K |
| Fireworks AI | US West | ~150ms | ~42 tps | 99.9% | ~$0.58 | ~$0.58 | ~200K |
Performance benchmarks
| Metric | Voxtral Small 24B 2507 (Mistral) | Mistral Large 2 123B | GPT-4.1 Mini |
|---|---|---|---|
| Avg Latency | ~220ms | ~280ms | ~180ms |
| Context Window | 128K | 128K | 128K |
| Input Price ($/1M) | $0.40 | $2.00 | $0.15 |
| Output Price ($/1M) | $1.20 | $6.00 | $0.60 |
| Max Output Tokens | 4K | 4K | 4K |
| Throughput | ~60 tps | ~40 tps | ~80 tps |
| Uptime | 99.9% | 99.9% | 99.9% |
30-day usage via LLM API
Architecture & Integration
One unified API. Every major model. Built-in reliability, cost control, and observability.
Dynamically route each request to the best model based on latency, cost, and capabilities—using one stable API contract instead of per-provider plumbing.
One endpoint, every model.Automatically balance price and quality with configurable cost ceilings, tiered model selection, and real-time usage controls so you never get surprised by your AI bill.
Control spend by design.Define automatic failover chains across providers and models. Survive outages and rate limits without rewriting application logic or degrading user experience.
Never fail on 500sTrace every request with structured logs, metrics, and latency breakdowns by provider and model. Debug production issues and tune routes using real data.
See every tokenDescribe high-level tasks like chat, extraction, or generation once. LLM.API maps them to best-fit models, so you decouple product logic from vendors.
Think tasks, not modelsProcess thousands of prompts in parallel with backpressure, retries, and cost controls built in. Ideal for reindexing, evaluations, and large content migrations.
Scale jobs, not codeDecision guide
FAQ
Voxtral Small 24B 2507 is a 24B-parameter Mistral model exposed via LLM.API, targeting high-quality, general-purpose text generation for developers.
Voxtral Small 24B 2507 is a text-only language model, supporting text input and text output through the LLM.API endpoints.
Voxtral Small 24B 2507 uses LLM.API’s unified per-token pricing; check your LLM.API dashboard or pricing docs for current input and output rates.
Voxtral Small 24B 2507 supports a multi‑kilotoken context window suitable for long prompts and conversations; see the LLM.API model card for exact limits.
Voxtral Small 24B 2507 is optimized for low-latency inference with streaming responses, but actual speed depends on prompt length and concurrency.
Voxtral Small 24B 2507 is best for general chat, code assistance, reasoning over medium-length documents, and building production assistants with predictable cost.
Use the standard LLM.API chat or completions endpoint and set the model field to "Voxtral Small 24B 2507" in your request payload.
Voxtral Small 24B 2507 targets a balance of quality and throughput comparable to other ~20–30B open models, exposed under a unified LLM.API interface.
Voxtral Small 24B 2507 can hallucinate, lacks real-time knowledge or browsing, and should not be used as a sole source for critical decisions.
If enabled in LLM.API, Voxtral Small 24B 2507 can follow tool or function-calling schemas, but behavior depends on your request format and routing configuration.
Compare
Claude Opus 4.8 (Fast) is Anthropic’s flagship Claude Opus 4.8 model running in a special fast mode that delivers significantly higher output token throughput at premium…
MoonshotAI Kimi Latest is the most recent version of MoonshotAI’s Kimi conversational large language model, designed for fast, web-connected chat and practical assistance in Chinese and…
LFM2.5-1.2B-Instruct (free) is a 1.2B-parameter, instruction-tuned hybrid language model from LiquidAI, optimized for fast, on-device inference with a ~32k token context window. It offers general-purpose conversational…