- Text Generation
Mercury 2 is a proprietary, diffusion-based large language model (dLLM) from Inception designed for extremely fast reasoning and text generation with a long 128K-token context window.
Powered by OpenAI
GPT-4o Transcribe is an OpenAI model specialized for converting audio into accurate, time-aligned text transcripts. It is notable for handling natural speech, varied accents, and real-world audio conditions with high reliability.
Output tokens per second · Higher is better
Seconds · Lower is better
USD per 1M tokens (blended) · Lower is better
About the model
GPT-4o Transcribe is a transcription-focused variant of OpenAI’s GPT-4o model designed to turn spoken audio into structured text. It is mainly used for tasks like meeting notes, call and podcast transcription, caption generation, and transforming voice recordings into searchable documents. It also supports workflows that combine transcription with light understanding, such as summarizing or tagging segments of speech. It belongs to the GPT-4o family of multimodal OpenAI models adapted for speech-to-text transcription workloads.
Model capabilities
Converts spoken audio into accurate, punctuated text transcripts, handling diverse speakers, accents, and recording conditions in real time.
Transcribes speech from multiple languages into text, preserving language-specific characters, names, and terminology where supported.
Generates structured transcripts for dialogues, meetings, and interviews, distinguishing speakers when metadata or channel separation is available.
Produces text captions from audio tracks in videos or podcasts, supporting workflows for accessibility, search, and content indexing.
Supports near real-time transcription for live audio streams, enabling monitoring, compliance checks, and rapid downstream processing.
Use cases
Transparent pricing
LLM API offers the lowest per‑minute transcription cost with best‑in‑class latency and uptime.
| Provider | Region | Latency | Throughput | Uptime | Input ($/1M) | Output ($/1M) | Context |
|---|---|---|---|---|---|---|---|
| LLM API BEST | Global | ~150ms | ~120 min/s | 99.99% | ~$0.004/min | ~$0.004/min | ~4 hour audio |
| OpenAI | Global | ~400ms | ~40 min/s | 99.9% | $0.006/min | $0.006/min | ~3 hour audio |
| Azure OpenAI | US East | ~450ms | ~35 min/s | 99.9% | ~$0.007/min | ~$0.007/min | ~3 hour audio |
| Google Cloud (Speech-to-Text via Gemini) | Global | ~500ms | ~30 min/s | 99.9% | ~$0.010/min | ~$0.010/min | ~2 hour audio |
| Amazon Web Services (Transcribe-like) | US East | ~550ms | ~25 min/s | 99.9% | ~$0.014/min | ~$0.014/min | ~2 hour audio |
Performance benchmarks
| Metric | GPT-4o Transcribe (OpenAI) | Whisper v3 (OpenAI) | Deepgram Nova-2 (Deepgram) |
|---|---|---|---|
| Avg Latency | ~180ms | ~250ms | ~220ms |
| Languages Supported | ~100+ | ~90+ | ~60+ |
| Price per Minute | ~$0.006 | ~$0.006 | ~$0.010 |
| Max Duration per Request | ~60 min | ~60 min | ~300 min |
| Accuracy (WER, English) | ~6–8% | ~7–9% | ~8–11% |
| Real-time Streaming Support | Yes | Yes | Yes |
| Throughput | ~50× RT | ~30× RT | ~60× RT |
| Uptime SLA | ~99.9% | ~99.9% | ~99.9% |
30-day usage via LLM API
Architecture & Integration
One unified API. Every major model. Built-in reliability, cost control, and observability.
Automatically route each request to the best model across providers based on latency, cost, and quality—without changing your integration or retraining clients.
One endpoint, any modelUse pricing-aware routing, quotas, and policies to keep spend predictable while still hitting quality and latency targets across models and clouds.
Control spend by designDefine failover rules once and let LLM.API retry on alternate providers or models when timeouts, rate limits, or provider outages occur.
Resilience built inGet traces, metrics, and structured logs for every request so you can debug prompts, compare providers, and tune performance in production.
See every tokenDescribe tasks like chat, extraction, or tooling once and let LLM.API pick and tune models behind the scenes, simplifying integration and future upgrades.
Code to tasks, not modelsSubmit large jobs as batches with automatic chunking, retries, and aggregation to slash costs and saturate throughput without writing glue code.
Scale jobs, not codeDecision guide
FAQ
GPT-4o Transcribe is an OpenAI GPT-4o-based model on LLM.API specialized for accurate, low-latency speech-to-text transcription.
GPT-4o Transcribe is best for real-time or batch transcription of meetings, calls, podcasts, and other audio into structured text.
GPT-4o Transcribe accepts audio input and returns text output; it is not intended for direct image or video understanding.
GPT-4o Transcribe is billed on LLM.API per unit of audio processed, typically metered in minutes or seconds rather than tokens.
GPT-4o Transcribe effectively handles long audio segments, but downstream text usage is constrained by the GPT-4o text context window.
GPT-4o Transcribe is optimized for low latency and can stream partial transcriptions for near real-time use cases.
You invoke GPT-4o Transcribe by specifying the model name in LLM.API audio endpoints and sending your audio file or stream payload.
GPT-4o Transcribe focuses on audio-to-text accuracy and efficiency, while general GPT-4o chat models focus on multi-turn natural language reasoning.
GPT-4o Transcribe supports multilingual transcription, but accuracy can vary by language and audio quality.
GPT-4o Transcribe may struggle with heavy background noise, overlapping speakers, very low-quality audio, or highly domain-specific jargon.
Compare
Mercury 2 is a proprietary, diffusion-based large language model (dLLM) from Inception designed for extremely fast reasoning and text generation with a long 128K-token context window.
all-mpnet-base-v2 is a widely used English sentence-embedding model from Sentence Transformers that maps text to 768-dimensional vectors for semantic similarity tasks. It is built on Microsoft’s…
Sora 2 Pro is an OpenAI model name that has been mentioned publicly, but as of now OpenAI has not released authoritative technical details or documentation…