- Speech-to-Text
Voxtral Mini Transcribe is a speech-to-text model from Mistral focused on lightweight, efficient audio transcription. It is designed to provide accurate transcriptions while being small and…
Powered by NVIDIA
Parakeet TDT 0.6B v3 is NVIDIA’s 600M-parameter multilingual automatic speech recognition (ASR) model built on the FastConformer-TDT architecture, optimized for high-throughput speech-to-text across European languages.
Output tokens per second · Higher is better
Seconds · Lower is better
USD per 1M tokens (blended) · Lower is better
About the model
Parakeet TDT 0.6B v3 is a 600-million-parameter multilingual speech-to-text model from NVIDIA based on the FastConformer-TDT architecture and trained on over 670,000 hours of audio from the Granary dataset. It is primarily used for real-time and batch transcription of audio and video, offering automatic language detection across roughly 25 European or EU languages and returning text with punctuation and timestamps. It is also adopted in cost-efficient pipelines and offline tools as an alternative to Whisper-class ASR for multilingual captioning, dictation, and media indexing. Parakeet TDT 0.6B v3 is part of NVIDIA’s Parakeet family and follows earlier Parakeet TDT 0.6B v2 and related multilingual ASR work.
Model capabilities
Performs automatic speech recognition across 25 European languages, converting spoken audio into accurate text transcripts.
Automatically identifies the spoken language in input audio among supported European languages before transcribing.
Processes long-form recordings up to several hours using FastConformer local attention while maintaining throughput and stability.
Generates transcripts with punctuation and segment-level timestamps suitable for indexing, search, and subtitle generation.
Supports consistent transcription quality across diverse European languages, enabling unified multilingual speech-to-text workflows.
Use cases
Transparent pricing
LLM API offers the lowest cost and latency for Parakeet‑class TDT models.
| Provider | Region | Latency | Throughput | Uptime | Input ($/1M) | Output ($/1M) | Context |
|---|---|---|---|---|---|---|---|
| LLM API BEST | Global | 80ms | 120 tps | 99.99% | $0.08 per 1M tokens | $0.08 per 1M tokens | 128K tokens |
| NVIDIA (Parakeet TDT 0.6B v3 via NIM) | US West | ~140ms | ~60 tps | ~99.9% | ~$0.20 per 1M tokens | ~$0.20 per 1M tokens | ~32K tokens |
| AWS Bedrock (Parakeet‑equivalent small TDT) | US East | ~160ms | ~40 tps | 99.9% | ~$0.30 per 1M tokens | ~$0.30 per 1M tokens | ~32K tokens |
| Azure AI (Parakeet‑class small TDT) | EU West | ~170ms | ~35 tps | 99.9% | ~$0.32 per 1M tokens | ~$0.32 per 1M tokens | ~32K tokens |
| GCP Vertex AI (Parakeet‑class small TDT) | Global | ~180ms | ~30 tps | ~99.9% | ~$0.35 per 1M tokens | ~$0.35 per 1M tokens | ~32K tokens |
Performance benchmarks
| Metric | Parakeet TDT 0.6B v3 | Parakeet TDT 1.1B v3 | Parakeet TDT 0.6B |
|---|---|---|---|
| Model Type | Transducer / TDT ASR | Transducer / TDT ASR | Transducer ASR |
| Parameter Count | 0.6B | 1.1B | 0.6B |
| Latency | — | — | — |
| Languages Supported | English | English | English |
| Price per Minute | — | — | — |
| Max Audio Duration | — | — | — |
| Accuracy (WER) | — | — | — |
| Uptime | — | — | — |
30-day usage via LLM API
Architecture & Integration
One unified API. Every major model. Built-in reliability, cost control, and observability.
Automatically route each request to the optimal model across providers based on latency, cost, or quality—without changing your integration or redeploying code.
One endpoint, every model.Balance quality and spend by mixing premium and budget models, enforcing per-project limits, and getting clear cost insights per call, team, and environment.
Cut spend, keep quality.Define automatic fallbacks across models and providers so timeouts, rate limits, or regional outages degrade gracefully instead of breaking your production workloads.
No single point of failure.Trace every request across models with logs, metrics, and structured events to debug latency, drift, and errors directly from your AI gateway, not scattered dashboards.
See every token hop.Describe tasks like chat, tools, RAG, or moderation once and let LLM.API pick the best implementation details and models for each provider behind the scenes.
Think tasks, not APIs.Process millions of inferences efficiently with provider-optimized batching, backoff, and concurrency controls that maximize throughput while staying within rate and budget limits.
Scale from 10 to millions.Decision guide
FAQ
Parakeet TDT 0.6B v3 is a 0.6B-parameter NVIDIA speech model focused on fast, lightweight transcription and diarization tasks.
It is best for real-time or near–real-time speech-to-text, turn detection, and diarization in resource-constrained or high-throughput environments.
Pricing is usage-based per audio duration processed; check the Parakeet TDT 0.6B v3 entry on LLM.API’s pricing page for current rates.
LLM.API caps the maximum audio duration per request; refer to the model’s documentation for the latest per-call audio length limit.
Thanks to its small 0.6B size, it offers low latency and is suitable for streaming or interactive speech applications via LLM.API.
Parakeet TDT 0.6B v3 accepts audio input and produces text outputs, including speaker and turn information where applicable.
Use the LLM.API speech endpoint with the model identifier for Parakeet TDT 0.6B v3, passing audio data and configuration in the request body.
Compared to larger Parakeet variants, it trades some accuracy for significantly lower latency, memory usage, and compute cost.
If enabled by LLM.API, you can use streaming mode for incremental transcripts; check the API reference for streaming availability details.
Its smaller size may reduce accuracy on noisy, highly accented, or domain-specific speech compared with larger, more capable speech models.
Compare
Voxtral Mini Transcribe is a speech-to-text model from Mistral focused on lightweight, efficient audio transcription. It is designed to provide accurate transcriptions while being small and…
Whisper Large V3 Turbo is OpenAI’s optimized, high‑speed variant of the Whisper Large V3 automatic speech recognition model, designed to provide fast transcriptions while preserving strong…
Whisper 1 is OpenAI’s hosted automatic speech recognition model based on the open-source Whisper family, designed for high-quality transcription and translation of audio. It is notable…