- Instruction Following
MoonshotAI Kimi Latest is the most recent version of MoonshotAI’s Kimi conversational large language model, designed for fast, web-connected chat and practical assistance in Chinese and…
Powered by Baidu
ERNIE 4.5 21B A3B Thinking is Baidu’s upgraded lightweight MoE language model optimized for deep reasoning, with a context window around 131K tokens and competitive pricing for large-scale use.
Output tokens per second · Higher is better
Seconds · Lower is better
USD per 1M tokens (blended) · Lower is better
About the model
ERNIE 4.5 21B A3B Thinking is a 21B-parameter sparse Mixture-of-Experts language model from Baidu, designed to activate about 3B parameters per token for efficient high-quality reasoning. It is mainly used for complex multi-step logical reasoning, math and science problem solving, and expert-level academic or benchmark tasks. It is also applied to coding assistance and advanced text generation where long-context (≈131K tokens) understanding is required at relatively low cost per token. The model belongs to Baidu’s ERNIE 4.5 family as a reasoning-enhanced successor to earlier ERNIE 4.x and ERNIE 3.x variants.
Model capabilities
Performs complex multi-step reasoning for logical puzzles, math, science, and academic-style problems using an MoE thinking architecture.
Acts as a conversational chat model, generating coherent, context-aware responses for interactive dialogue and assistant-style applications.
Produces long-form, structured written content and explanations over very long contexts up to around 128K–131K tokens.
Understands and generates text in both Chinese and English, suitable for bilingual tasks and cross-language information access.
Provides efficient tool usage capabilities, supporting structured interactions like function or tool calling in complex workflows.
Use cases
Transparent pricing
LLM API offers the lowest token prices and latency for ERNIE 4.5–class models.
| Provider | Region | Latency | Throughput | Uptime | Input ($/1M) | Output ($/1M) | Context |
|---|---|---|---|---|---|---|---|
| LLM API BEST | Global | 120ms | 80 tps | 99.99% | $0.40 | $1.20 | 200K |
| Baidu | China | ~280ms | ~40 tps | 99.9% | ~$0.80 | ~$2.40 | ~128K |
| Alibaba Cloud | APAC East | ~260ms | ~35 tps | 99.9% | ~$0.90 | ~$2.70 | ~128K |
| Tencent Cloud | APAC North | ~300ms | ~30 tps | 99.9% | ~$0.95 | ~$2.85 | ~100K |
Performance benchmarks
| Metric | ERNIE 4.5 21B A3B Thinking | GPT-4o (128K) | Gemini 1.5 Pro |
|---|---|---|---|
| Avg Latency | ~900ms | ~700ms | ~800ms |
| Context Window | 128K | 128K | 1M |
| Input Price ($/1M) | $0.90 | $5.00 | $3.50 |
| Output Price ($/1M) | $3.00 | $15.00 | $10.50 |
| Max Output Tokens | 4K | 4K | 8K |
| Throughput | ~60 tps | ~40 tps | ~35 tps |
| Uptime | 99.9% | 99.9% | 99.9% |
30-day usage via LLM API
Architecture & Integration
One unified API. Every major model. Built-in reliability, cost control, and observability.
Automatically route each request to the optimal model across providers using policies, latency, and quality signals—no client changes or new integrations required.
One endpoint, every modelEnforce budgets, pick cheaper equivalents, and downgrade gracefully under load so you control spend without touching application code or sacrificing SLAs.
Lower costs, same outputDefine provider and model failover chains so requests auto-retry on alternates, shielding your app from outages, rate limits, and sudden model deprecations.
Stay online, automaticallyTrace every call across providers with logs, metrics, and structured events to debug failures, compare models, and tune prompts from one unified view.
See every token hopDeclare tasks, tools, and constraints once; LLM.API handles planning, multi-step execution, and provider selection for robust, reusable AI workflows.
From prompts to workflowsShip millions of inferences as managed batches with automatic chunking, retries, and aggregation so you can backfill datasets or experiments at scale.
Crank up the throughputDecision guide
FAQ
ERNIE 4.5 21B A3B Thinking is a 21-billion-parameter Baidu large language model focused on reasoning-heavy text generation tasks.
It is best for multi-step reasoning, complex code understanding, tool-using agents, and analytical workflows where chain-of-thought quality matters more than raw speed.
ERNIE 4.5 21B A3B Thinking supports up to a 32K token context window on LLM.API, including prompt and generated tokens.
LLM.API charges per 1,000 input and output tokens for this model; check your LLM.API pricing page for current rates.
Typical first-token latency is a few hundred milliseconds to a couple of seconds, depending on load, with streamed tokens arriving progressively.
On LLM.API, ERNIE 4.5 21B A3B Thinking currently supports text input and text output only.
Use the standard LLM.API chat or completion endpoint and set the model field to the ERNIE 4.5 21B A3B Thinking identifier.
Compared with similar-sized models, it emphasizes stronger step-by-step reasoning but may be slower and more expensive per request.
It can hallucinate facts, has no real-time web access, and may struggle with highly specialized domain knowledge without careful prompting.
Yes, you can pair it with LLM.API's tool-calling mechanisms, but tool schemas and orchestration logic must be implemented on your side.
Compare
MoonshotAI Kimi Latest is the most recent version of MoonshotAI’s Kimi conversational large language model, designed for fast, web-connected chat and practical assistance in Chinese and…
INTELLECT-3 is an AI model from Prime Intellect, but publicly available technical details about its architecture, capabilities, and benchmarks are not documented. Information about its specific…
Grok Build 0.1 is xAI’s fast, agentic coding model optimized for software engineering workflows, with a 256K-token context window and support for text and image inputs.