- Speech-to-Text
GPT-4o Mini Transcribe is an OpenAI model specialized for converting spoken language in audio into accurate text. It is optimized for lightweight, fast transcription while maintaining…
Powered by Google
Chirp 3 is Google's latest-generation multilingual speech and audio model, available through Google Cloud for high-accuracy transcription and natural-sounding text-to-speech. It is notable for its improved accuracy, speed, and support for advanced features like diarization, automatic language detection, and custom voices.
Output tokens per second · Higher is better
Seconds · Lower is better
USD per 1M tokens (blended) · Lower is better
About the model
Chirp 3 is a multilingual Automatic Speech Recognition and audio generation model from Google that powers Speech-to-Text and Text-to-Speech capabilities in Google Cloud. It is used for accurate real-time and batch audio transcription across many languages, including support for speaker diarization and language-agnostic transcription. It is also used to generate high-fidelity synthetic speech, including instant custom voice models built from high-quality recordings. Chirp 3 succeeds earlier Chirp models as part of Google’s Chirp family of speech and audio foundation models.
Model capabilities
Engages in natural, multi-turn voice conversations, understanding user intent and context to provide relevant, coherent spoken responses.
Converts spoken language in audio input into accurate text, supporting real-time or near real-time voice transcription scenarios.
Translates spoken language from one language to another, enabling cross-lingual voice conversations and real-time interpretation use cases.
Processes and monitors audio streams for commands or triggers, enabling responsive voice-driven applications and interactive systems.
Can be integrated with image-capable systems to associate spoken descriptions with visual content, supporting multimodal user experiences.
Use cases
Transparent pricing
LLM API offers the lowest cost and highest performance for Chirp 3‑class speech models.
| Provider | Region | Latency | Throughput | Uptime | Input ($/1M) | Output ($/1M) | Context |
|---|---|---|---|---|---|---|---|
| LLM API BEST | Global | 80ms | 220 min/s | 99.99% | $0.008/min | $0.008/min | ~480 min audio |
| Global | ~150ms | ~150 min/s | 99.9% | ~$0.012/min | ~$0.012/min | ~300 min audio | |
| Azure | Global | ~170ms | ~120 min/s | 99.9% | ~$0.013/min | ~$0.013/min | ~240 min audio |
| Amazon Web Services | Global | ~190ms | ~100 min/s | 99.9% | ~$0.014/min | ~$0.014/min | ~240 min audio |
Performance benchmarks
| Metric | Chirp 3 (Google) | Whisper v3 (OpenAI) | NeMo ASR Large (NVIDIA) |
|---|---|---|---|
| Avg Latency | ~250ms | ~300ms | ~350ms |
| Languages Supported | ~100+ | ~100+ | ~30+ |
| Price per Minute | ~$0.006 | ~$0.006 | ~$0.005 |
| Max Duration | ~2 hours | ~2 hours | ~3 hours |
| Accuracy (WER) | ~6% | ~5% | ~7% |
| Uptime | 99.9% | 99.9% | 99.9% |
| Real-time Throughput | ~60x RT | ~50x RT | ~40x RT |
30-day usage via LLM API
Architecture & Integration
One unified API. Every major model. Built-in reliability, cost control, and observability.
Automatically route each request to the optimal model across providers based on latency, cost, and quality. One endpoint, dynamic policies, no SDK sprawl.
One endpoint, any modelControl spend with price-aware routing, per-project limits, and transparent metering across vendors. Swap models without rewiring billing or touching client code.
Cut cost, keep qualityDesign multi-provider fallback trees that auto-retry on failures, timeouts, or quota limits. Keep production workloads online even when a vendor has issues.
Never ship single-vendor SPOFGet unified traces, logs, and metrics for every request across providers. Inspect prompts, latencies, and errors in one place to debug faster and tune confidently.
Single pane of AI truthDescribe tasks like “chat”, “embed”, or “moderate” instead of binding to model names. LLM.API maps tasks to the best capabilities behind a stable interface.
Code to tasks, not modelsShip bulk workloads with streaming-safe, rate-aware batching. Push thousands of prompts per job while LLM.API handles chunking, retries, and provider limits.
Batch at production scaleDecision guide
FAQ
Chirp 3 is a Google speech model focused on automatic speech recognition with strong multilingual performance and robustness to noisy, real‑world audio.
Chirp 3 is best for high‑accuracy, large‑scale transcription of calls, meetings, videos, and user‑generated audio across many languages and accents.
Through LLM.API, Chirp 3 supports audio input and text output for speech‑to‑text workloads, without image or text‑generation capabilities.
Chirp 3 is typically billed per processed audio minute or second via LLM.API; check your LLM.API pricing page for exact current rates.
Chirp 3 supports long‑form audio transcription, but maximum duration and effective context depend on LLM.API limits and configuration for streaming or batch mode.
Chirp 3 generally operates near real time for short clips, with latency mainly determined by audio length and LLM.API region and network conditions.
You select the Google Chirp 3 model in your LLM.API request, provide audio bytes or a URL, and receive transcribed text in the response.
Compared with general text LLMs, Chirp 3 is specialized, usually cheaper and more accurate for speech recognition but cannot perform text‑only reasoning.
If enabled by LLM.API, Chirp 3 can consume audio chunks incrementally and return partial transcripts for low‑latency streaming experiences.
Chirp 3 is limited to speech recognition, may struggle with extremely noisy audio, rare languages, domain‑specific jargon, and does not generate or understand images.
Compare
GPT-4o Mini Transcribe is an OpenAI model specialized for converting spoken language in audio into accurate text. It is optimized for lightweight, fast transcription while maintaining…
Parakeet TDT 0.6B v3 is NVIDIA’s 600M-parameter multilingual automatic speech recognition (ASR) model built on the FastConformer-TDT architecture, optimized for high-throughput speech-to-text across European languages.
Whisper 1 is OpenAI’s hosted automatic speech recognition model based on the open-source Whisper family, designed for high-quality transcription and translation of audio. It is notable…