- Text Generation
Seedance 1.5 Pro is ByteDance’s flagship native joint audio‑video generation model, focused on high‑quality, lip‑synced video with synchronized sound. It is notable for producing short, production‑ready…
Powered by Qwen
Qwen3 ASR Flash is Qwen’s high-accuracy, multilingual automatic speech recognition (ASR) service optimized for real-time transcription of short audio. It is built on the Qwen3-Omni foundation model and trained on tens of millions of hours of multimodal speech data for robust performance across noisy and varied environments.
Output tokens per second · Higher is better
Seconds · Lower is better
USD per 1M tokens (blended) · Lower is better
About the model
Qwen3 ASR Flash is an automatic speech recognition model and cloud service from Qwen (Alibaba) designed for fast, accurate transcription of short audio segments. It is mainly used to convert speech to text in real time for applications such as live captioning, meeting or call transcription, and voice-driven interfaces. It is also used as a backend ASR component in broader multimodal and translation pipelines, including tools that extend it to long-form audio transcription. The model is part of the Qwen3-ASR family and is built on the Qwen3-Omni multimodal model within the broader Qwen3 model ecosystem.
Model capabilities
Performs low-latency automatic speech recognition, transcribing spoken audio to text in real time for interactive applications.
Converts prerecorded audio files into accurate text transcripts, supporting efficient processing of long-form speech content.
Recognizes and transcribes speech across multiple languages, enabling global voice-powered applications and multilingual audio processing.
Enables voice-driven control and command interfaces by reliably turning spoken instructions into structured text for downstream handling.
Handles diverse acoustic conditions and speaking styles to robustly capture and transcribe speech in real-world noisy environments.
Use cases
Transparent pricing
LLM API offers the lowest ASR minute pricing and best overall performance for Qwen3 ASR-class models.
| Provider | Region | Latency | Throughput | Uptime | Input ($/1M) | Output ($/1M) | Context |
|---|---|---|---|---|---|---|---|
| LLM API BEST | Global | 120ms | 120 audio min/s | 99.99% | $0.004/min | $0.00/min | ~4 hr audio |
| Qwen | Global | ~180ms | ~80 audio min/s | ~99.9% | ~$0.006/min | $0.00/min | ~3 hr audio |
| Alibaba Cloud | APAC | ~220ms | ~60 audio min/s | ~99.9% | ~$0.007/min | $0.00/min | ~2 hr audio |
| Replicate | Global | ~250ms | ~40 audio min/s | ~99.5% | ~$0.010/min | $0.00/min | ~2 hr audio |
| Fireworks AI | US East | ~200ms | ~70 audio min/s | ~99.9% | ~$0.008/min | $0.00/min | ~3 hr audio |
Performance benchmarks
| Metric | Qwen3 ASR Flash | Whisper Large v3 (OpenAI API) | Deepgram Nova-2 |
|---|---|---|---|
| Avg Latency (Streaming) | — | Real‑time or better on GPU (varies by provider) | — |
| Languages Supported | Multilingual (exact count —) | ≈99 languages | Multilingual (English + others; exact count —) |
| Price per Minute (Hosted API) | ~$0.0019/min | $0.006/min | $0.0043/min (pre‑recorded baseline) |
| Max Audio Duration per Request | — | ~25 MB per request via OpenAI Whisper-1; v3 limits vary by host | Typical API up to multi‑hour files; hard limit — |
| Accuracy (WER, clean English) | State‑of‑the‑art vs Whisper v3 (exact WER —) | ≈2.7% WER on clean audio | Higher accuracy than Nova and Whisper v2; vs Whisper v3 — |
| Model Type / Architecture | All‑in‑one ASR, non‑autoregressive alignment; Qwen3‑based | Encoder–decoder Transformer ASR | End‑to‑end neural ASR (Deepgram Nova family) |
| Deployment / Availability | Cloud API via Alibaba/Qwen; open weights for some variants | Open‑source weights + multiple hosted APIs | Proprietary hosted API (Deepgram cloud) |
| Licensing | Apache‑2.0 for open‑weight variants; commercial terms for cloud | MIT license (open weights); commercial API terms | Commercial, closed‑source |
30-day usage via LLM API
Architecture & Integration
One unified API. Every major model. Built-in reliability, cost control, and observability.
Dynamically route each request to the best model across providers based on latency, cost, and quality—no client changes, no redeploys, just smarter defaults.
One endpoint, any modelSet hard budgets, price caps, and model tiers so teams can experiment freely while finance stays in control of spend across every AI provider.
Predictable AI spendDefine provider and model failover rules so traffic transparently shifts on errors or outages—keeping your AI features online without manual intervention.
Resilience by defaultTrace every request, token, error, and latency across providers with unified logs, metrics, and alerts so you can debug, tune, and prove ROI in one place.
See every tokenExpress higher-level tasks—chat, tools, RAG, evaluation—through a single abstraction that hides provider quirks, simplifying complex AI workflows into clean, testable units.
One API for tasksSubmit massive batches of generations or evaluations with built-in chunking, retries, and concurrency control to saturate throughput limits without blowing up rate caps.
Scale jobs, not codeDecision guide
FAQ
Qwen3 ASR Flash is a fast automatic speech recognition model by Qwen optimized for low-latency transcription via API.
Qwen3 ASR Flash accepts audio as input and outputs text transcripts.
Qwen3 ASR Flash prioritizes speed and low cost over maximum accuracy or advanced language understanding found in larger general-purpose Qwen models.
Qwen3 ASR Flash supports long-form audio segments, but you should chunk very long recordings client-side to manage latency and partial failures.
Yes, Qwen3 ASR Flash is designed for low-latency use cases like real-time or near real-time transcription where speed is critical.
Qwen3 ASR Flash may struggle with heavy background noise, very low-resource languages, domain-specific jargon, or tasks requiring deep semantic understanding beyond transcription.
LLM.API exposes Qwen3 ASR Flash with usage-based pricing per audio duration; check the LLM.API pricing page for the latest exact rates.
Qwen3 ASR Flash is tuned for high throughput and low latency, typically returning transcripts much faster than the input audio duration.
You specify the provider as Qwen and the model name as Qwen3 ASR Flash in your LLM.API request, sending audio content in the supported format.
Qwen3 ASR Flash supports multilingual transcription, but accuracy varies by language and is generally best for its highest-resource languages.
Compare
Seedance 1.5 Pro is ByteDance’s flagship native joint audio‑video generation model, focused on high‑quality, lip‑synced video with synchronized sound. It is notable for producing short, production‑ready…
Cogito v2.1 671B is Deep Cogito’s flagship 671B-parameter open-weight Mixture-of-Experts language model optimized for efficient hybrid reasoning. It delivers frontier-level performance while using significantly shorter reasoning…
Orpheus 3B is a 3-billion-parameter English text-to-speech model from Canopy Labs, optimized for natural prosody, expressive delivery, and real-time streaming speech generation. It is notable for…