- Text Generation
Gemini 3.1 Pro Preview Custom Tools is a preview large language model from Google’s Gemini 3.1 Pro line that supports integration with user-defined tools and APIs.…
Powered by Xiaomi
MiMo-V2-Flash is an open-source Mixture-of-Experts language model from Xiaomi optimized for fast, long-context reasoning and coding. It combines a 309B-parameter MoE architecture with only 15B active parameters to deliver high performance at low cost.
Output tokens per second · Higher is better
Seconds · Lower is better
USD per 1M tokens (blended) · Lower is better
About the model
MiMo-V2-Flash is a Xiaomi open-source foundation language model using a Mixture-of-Experts architecture with 309B total parameters and 15B active parameters, designed for efficient high-speed inference. It is mainly used for complex reasoning tasks, code generation, and agent-style workflows where both quality and latency matter. With its 256K–262K token context window, it also serves long-form text generation and analysis use cases such as documentation, data processing, and interactive applications. It belongs to Xiaomi’s MiMo-V2 family of models, alongside variants like MiMo-V2-Pro and MiMo-V2-Omni.
Model capabilities
Performs strong logical and analytical reasoning, achieving competitive results on complex benchmarks and decision-making tasks at low cost.
Generates, debugs, and explains source code, performing competitively on software engineering benchmarks like SWE-Bench and related tasks.
Acts as a foundation for AI agents, handling tool invocation, planning, and multi-step task execution in practical applications.
Supports extended conversational sessions with very large context windows, maintaining coherence across long interactions and documents.
Understands and generates text in multiple languages, suitable for cross-language interactions and globally-deployed Xiaomi ecosystem products.
Use cases
Transparent pricing
LLM API offers the lowest latency and cost for MiMo‑V2‑Flash–class models.
| Provider | Region | Latency | Throughput | Uptime | Input ($/1M) | Output ($/1M) | Context |
|---|---|---|---|---|---|---|---|
| LLM API BEST | Global | 80ms | 70 tps | 99.99% | $0.06 | $0.06 | 64K tokens |
| Xiaomi | Global | ~150ms | ~40 tps | ~99.9% | ~$0.08 | ~$0.08 | ~32K tokens |
| OpenAI | US East | ~110ms | ~50 tps | ~99.9% | ~$0.10 | ~$0.10 | ~128K tokens |
| Google Cloud | EU West | ~130ms | ~45 tps | ~99.9% | ~$0.09 | ~$0.09 | ~64K tokens |
| Azure | US West | ~140ms | ~42 tps | ~99.95% | ~$0.11 | ~$0.11 | ~128K tokens |
Performance benchmarks
| Metric | MiMo-V2-Flash | Xiaomi MiMo-V2 | Huawei PanGu-Flash |
|---|---|---|---|
| Avg Latency | ~180ms | ~240ms | ~220ms |
| Context Window | 128K | 64K | 128K |
| Input Price ($/1M tokens) | $0.25 | $0.30 | $0.28 |
| Output Price ($/1M tokens) | $0.75 | $0.90 | $0.85 |
| Max Output Tokens | 4K | 4K | 8K |
| Throughput | 60 tps | 40 tps | 50 tps |
| Uptime | 99.9% | 99.5% | 99.7% |
30-day usage via LLM API
Architecture & Integration
One unified API. Every major model. Built-in reliability, cost control, and observability.
Intelligently route each request across providers and models based on performance, latency, or cost. One API, pluggable policies, no client rewrites.
One endpoint, any modelDynamically balance premium and budget models with per-project guardrails. Ship features faster while keeping AI spend predictable and auditable.
Control your AI billRecover from provider outages and timeouts with built-in failover to backup models. Your AI features keep working, even when vendors don’t.
Resilient by defaultTrace every request across providers with metrics, logs, and structured events. Debug prompts, tune routing, and prove reliability with real data.
See every tokenDefine tasks like chat, tools, RAG, or vision once, then swap underlying models freely. Keep business logic stable as the model landscape shifts.
Code to tasks, not modelsProcess millions of requests cost-effectively with batch APIs optimized for concurrency and retries. Perfect for backfills, evaluations, and bulk content generation.
Scale workloads cheaplyDecision guide
FAQ
MiMo-V2-Flash is a Xiaomi multimodal large language model accessible through LLM.API, tuned for fast, low-latency generation on text and images.
MiMo-V2-Flash is best for interactive apps needing quick responses, such as chatbots, lightweight agents, and image-aware assistants with rapid turn-around.
MiMo-V2-Flash supports a context window up to 8,000 tokens via LLM.API, suitable for moderately long conversations and documents.
MiMo-V2-Flash is optimized for low latency, typically streaming first tokens within a few hundred milliseconds under normal load on LLM.API.
MiMo-V2-Flash supports text input and output, plus image input for vision-language tasks like captioning, classification, and grounded Q&A.
MiMo-V2-Flash uses LLM.API’s unified token-based pricing, billed per input and output token according to the Xiaomi MiMo-V2-Flash rate tier.
Use the LLM.API chat or completion endpoint with the model identifier "xiaomi/mimo-v2-flash" and pass your prompts as usual JSON payloads.
Compared to similar flash models, MiMo-V2-Flash emphasizes low latency and solid multimodal capabilities, trading off some reasoning depth for speed.
MiMo-V2-Flash may underperform larger models on complex multi-step reasoning, long-context synthesis, and highly specialized domain knowledge.
Yes, MiMo-V2-Flash supports streaming, allowing tokens to be delivered incrementally for faster perceived latency in interactive applications.
Compare
Gemini 3.1 Pro Preview Custom Tools is a preview large language model from Google’s Gemini 3.1 Pro line that supports integration with user-defined tools and APIs.…
KAT-Coder-Pro V2 is Kwaipilot's second-generation flagship agentic coding model with a 256K-token context window, optimized for complex software engineering and large-codebase tasks. It is designed for…
Qwen3 VL 30B A3B Thinking is a large multimodal Qwen model with around 30 billion parameters, designed for vision-language reasoning with extended “thinking” capabilities. It is…