- Speech-to-Text
Whisper Large V3 Turbo is OpenAI’s optimized, high‑speed variant of the Whisper Large V3 automatic speech recognition model, designed to provide fast transcriptions while preserving strong…
Powered by OpenAI
Whisper 1 is OpenAI’s hosted automatic speech recognition model based on the open-source Whisper family, designed for high-quality transcription and translation of audio. It is notable for robust multilingual speech-to-text performance and language identification across diverse audio conditions.
Output tokens per second · Higher is better
Seconds · Lower is better
USD per 1M tokens (blended) · Lower is better
About the model
Whisper 1 is an OpenAI speech recognition model served via API for converting spoken audio into text. It is mainly used for automatic transcription of recordings such as meetings, podcasts, or voice notes, and for generating captions or searchable text from spoken content. It is also widely used to translate non‑English speech into English transcripts and to detect the spoken language in audio. Whisper 1 belongs to the Whisper model family and is based on the large-v2 variant of OpenAI’s open-source Whisper models.
Model capabilities
Converts spoken audio into accurate text transcriptions across many languages, handling varied accents, recording conditions, and speaking styles.
Transcribes speech in multiple supported languages, preserving original language content while coping with diverse pronunciations and vocabularies.
Translates spoken language in audio into written text in another language, enabling cross-lingual understanding and communication.
Extracts spoken content from audio or video files, effectively performing OCR-like text extraction for voice-based information.
Provides text outputs describing spoken segments in audio, supporting captioning and subtitling workflows for media content.
Use cases
Transparent pricing
LLM API offers the lowest Whisper‑class transcription cost and latency across major providers.
| Provider | Region | Latency | Throughput | Uptime | Input ($/1M) | Output ($/1M) | Context |
|---|---|---|---|---|---|---|---|
| LLM API BEST | Global | ~180ms | ~120 audio min/s | 99.99% | ~$0.003/min | $0.00 | ~4 hour audio |
| OpenAI | Global | ~250ms | ~60 audio min/s | 99.9% | $0.006/min | $0.006/min | 30 min audio |
| Azure OpenAI | US East | ~450ms | ~45 min/s | 99.9% | ~$0.0065/min | ~$0.0065/min | 30 min audio |
| Google Cloud Speech-to-Text | Global | ~500ms | ~40 min/s | 99.9% | ~$0.009/min | ~$0.009/min | 30 min audio |
| Amazon Transcribe | US East | ~550ms | ~35 min/s | 99.9% | ~$0.008/min | ~$0.008/min | 30 min audio |
Performance benchmarks
| Metric | Whisper 1 (OpenAI) | Google Speech-to-Text v2 | Amazon Transcribe |
|---|---|---|---|
| Avg Latency | ~300ms | ~350ms | ~400ms |
| Languages Supported | ~99 | ~73 | ~79 |
| Price per Minute | $0.006 | $0.012 | $0.015 |
| Max Duration per Request | 60 min | 480 min | 240 min |
| Accuracy (WER) | ~7% | ~8% | ~9% |
| Uptime | 99.9% | 99.9% | 99.9% |
| Streaming Support | Yes | Yes | Yes |
30-day usage via LLM API
Architecture & Integration
One unified API. Every major model. Built-in reliability, cost control, and observability.
Automatically direct each request to the optimal model across providers using latency, cost, and quality signals, so you ship faster without hardcoding vendor logic.
One endpoint, smart routing.Control spend with per-route budgets, price-aware model selection, and real-time usage insights, so you can scale traffic without surprise bills or manual tuning.
Optimize cost, not code.Define automatic failover chains across models and providers, so outages, rate limits, or degraded quality don’t take your features offline.
Stay online under stress.Inspect every request with traces, metrics, and structured logs across providers, making it easy to debug prompts, compare models, and tune performance in production.
See every token flow.Describe tasks like chat, RAG, or extraction once, then swap models or providers without rewriting business logic, keeping your app code clean and future-proof.
Code to tasks, not models.Process large workloads with parallelized, provider-agnostic batching and automatic retries, reducing latency and unit cost for bulk jobs and backfills.
Batch at production scale.Decision guide
FAQ
Whisper 1 is OpenAI’s automatic speech recognition (ASR) model for transcribing and translating audio into text.
Whisper 1 supports audio input and returns text output for transcription and translation tasks through LLM.API.
Whisper 1 is best for accurate speech-to-text transcription, multilingual audio transcription, and speech translation to English.
Whisper 1 is typically billed per minute of processed audio; consult LLM.API’s pricing page for exact current rates.
Whisper 1 generally supports long-form audio, but maximum duration may be capped by LLM.API request size and timeout limits.
Whisper 1 usually processes audio close to or faster than real time, but actual latency depends on audio length and LLM.API infrastructure.
You select the Whisper 1 model identifier in your LLM.API request and send audio data in the supported format and encoding.
Whisper 1 is generally more accurate, robust, and cost-efficient for transcription than using general-purpose text-only LLMs with external audio preprocessing.
Yes, Whisper 1 supports many languages for transcription and can translate non-English speech into English text.
Whisper 1 typically supports common formats like MP3, MP4, WAV, and FLAC with standard speech sample rates such as 16 kHz.
Real-time streaming support depends on LLM.API features; if streaming endpoints are provided, they can expose Whisper 1 for low-latency use.
Whisper 1 may struggle with heavy background noise, strong accents, overlapping speakers, domain-specific jargon, and very low-quality recordings.
Compare
Whisper Large V3 Turbo is OpenAI’s optimized, high‑speed variant of the Whisper Large V3 automatic speech recognition model, designed to provide fast transcriptions while preserving strong…
Voxtral Mini Transcribe is a speech-to-text model from Mistral focused on lightweight, efficient audio transcription. It is designed to provide accurate transcriptions while being small and…
GPT-4o Mini Transcribe is an OpenAI model specialized for converting spoken language in audio into accurate text. It is optimized for lightweight, fast transcription while maintaining…