- Text-to-Speech
GPT-4o Mini TTS is a text-to-speech variant of OpenAI’s lightweight GPT-4o Mini model, designed to generate natural-sounding spoken audio from text with low latency and efficient…
Powered by OpenAI
GPT Audio is an OpenAI model that can understand and generate natural-sounding speech in real time. It is notable for combining strong language understanding with fast, conversational audio input and output.
Output tokens per second · Higher is better
Seconds · Lower is better
USD per 1M tokens (blended) · Lower is better
About the model
GPT Audio is an OpenAI model designed for real-time speech understanding and generation. It is mainly used to power voice-based assistants, enabling spoken conversations that include tasks like answering questions, controlling applications, and assisting with productivity. It is also used for interactive experiences such as hands-free interfaces, accessibility tools, and multimodal applications where speech is combined with text or other inputs. GPT Audio is part of OpenAI’s GPT family of generative models, extending them from text and images into low-latency voice interaction.
Model capabilities
Engages in natural, low-latency spoken dialogue, handling interruptions and back-and-forth conversation while reasoning about user intent.
Converts spoken language in audio into accurate text transcripts, supporting multiple speakers and diverse recording conditions.
Generates natural-sounding speech from text input, enabling interactive voice experiences and read-aloud functionality.
Listens to speech in one language and outputs translated text or speech in another, preserving meaning and conversational flow.
Interprets audio content beyond transcription, using it as context for reasoning, answering questions, or following spoken instructions.
Use cases
Transparent pricing
LLM API offers the lowest audio prices and latency for GPT Audio–class models.
| Provider | Region | Latency | Throughput | Uptime | Input ($/1M) | Output ($/1M) | Context |
|---|---|---|---|---|---|---|---|
| LLM API BEST | Global | ~150ms | ~120 req/s | 99.99% | ~$0.10/hr | ~$0.10/hr | ~10 hr audio |
| OpenAI | Global | ~400ms | ~40 req/s | 99.9% | ~$0.36/hr | ~$0.36/hr | ~4 hr audio |
| Azure OpenAI | US East | ~450ms | ~35 req/s | 99.9% | ~$0.40/hr | ~$0.40/hr | ~4 hr audio |
| Google Cloud (Speech/Audio Gen) | US Central | ~500ms | ~30 req/s | 99.9% | ~$0.50/hr | ~$0.50/hr | ~3 hr audio |
| Amazon Web Services (Bedrock Audio) | US West | ~550ms | ~25 req/s | 99.9% | ~$0.55/hr | ~$0.55/hr | ~3 hr audio |
Performance benchmarks
| Metric | GPT Audio (OpenAI) | Whisper v3 (OpenAI) | Google Speech-to-Text v2 |
|---|---|---|---|
| Avg Latency | ~180ms | ~250ms | ~300ms |
| Languages Supported | ~50+ | ~50+ | ~70+ |
| Price per Minute | $0.015 | $0.010 | $0.016 |
| Max Duration per Request | 60 min | 60 min | 60 min |
| Accuracy (WER, English clean) | ~5.0% | ~6.0% | ~7.5% |
| Accuracy (WER, noisy) | ~9.5% | ~11.0% | ~12.5% |
| Uptime SLA | 99.9% | 99.9% | 99.5% |
30-day usage via LLM API
Architecture & Integration
One unified API. Every major model. Built-in reliability, cost control, and observability.
Automatically route each request to the optimal model across providers based on latency, cost, or quality. One API, zero vendor lock-in, instant flexibility.
One endpoint, any modelControl spend with smart routing, tiered models, and granular usage limits. Optimize every token without rewriting application logic or duplicating integration work.
More performance, less spendDefine automatic fallbacks when a provider throttles, fails, or degrades. Keep your production workloads online without custom retry code per vendor.
Stay up when models failTrace every request across providers with logs, metrics, and structured events. Debug latency, errors, and drift from a single, provider-agnostic dashboard.
See every token’s pathCall models by task—chat, tools, embeddings, rerank—through a consistent API. Swap providers or upgrade models without touching downstream application code.
Code to tasks, not vendorsRun massive batch jobs across providers with automatic chunking, retries, and aggregation. Process millions of inputs efficiently without hand-rolled job infrastructure.
Ship batch at scaleDecision guide
FAQ
GPT Audio is an OpenAI model on LLM.API that adds low-latency, bidirectional audio input and output to the GPT language capabilities.
GPT Audio supports text input, audio input, and audio or text output, enabling real-time voice assistants and conversational interfaces.
You call the unified LLM.API endpoint with the GPT Audio model name, sending text or audio input and receiving streaming audio or text responses.
GPT Audio is best for real-time voice agents, interactive assistants, and applications needing natural, low-latency spoken conversations.
GPT Audio inherits the underlying GPT model’s context window, typically up to 128K tokens depending on the configured base model.
GPT Audio is optimized for sub-second token-level streaming, allowing responses to start playing almost immediately after user speech.
GPT Audio is billed per input and output token, with audio tokens counted similarly to text tokens according to LLM.API’s OpenAI pricing schedule.
Compared to text-only GPT models, GPT Audio adds speech recognition and speech synthesis, enabling end-to-end voice experiences without separate ASR or TTS services.
GPT Audio can handle interactive conversational streams, but very long uninterrupted audio may require chunking and session management in your application.
GPT Audio may struggle with heavy background noise, highly technical jargon, or strict real-time requirements below typical network round-trip latencies.
Compare
GPT-4o Mini TTS is a text-to-speech variant of OpenAI’s lightweight GPT-4o Mini model, designed to generate natural-sounding spoken audio from text with low latency and efficient…
Kokoro 82M is an open-weight, 82‑million‑parameter text‑to‑speech model from hexgrad that focuses on natural, multilingual speech with low latency and low resource usage. It is notable…
Grok Voice TTS 1.0 is xAI’s text-to-speech model that turns Grok’s language outputs into natural-sounding, expressive audio with multilingual support and fine-grained control over delivery. It…