- Text Generation
Body Builder (beta) is an OpenRouter model that converts natural language descriptions into structured OpenRouter API request objects, enabling automated construction of complex, multi-model calls.
Powered by Inception
Mercury 2 is a proprietary, diffusion-based large language model (dLLM) from Inception designed for extremely fast reasoning and text generation with a long 128K-token context window.
Output tokens per second · Higher is better
Seconds · Lower is better
USD per 1M tokens (blended) · Lower is better
About the model
Mercury 2 is a commercial-scale diffusion-based language model by Inception optimized for high-speed reasoning and generation. It is primarily used for code generation, analytical reasoning, and complex automation workflows where low latency is critical. It is also applied in AI agents, search, and business applications that benefit from rapid, large-context processing. Mercury 2 belongs to Inception’s Mercury family of diffusion-based LLMs, succeeding earlier Mercury models and specialized variants such as Mercury Coder.
Model capabilities
Engages in multi-turn dialogue, answering questions and following instructions while maintaining context across user interactions.
Processes images to identify objects and scenes, enabling descriptions and basic reasoning about visual content.
Translates written content between multiple languages while attempting to preserve meaning and tone.
Extracts machine-readable text from images or scanned documents, supporting downstream search or analysis.
Assists in monitoring streams of textual data for specific topics or issues using pattern matching and basic analysis.
Use cases
Transparent pricing
LLM API offers the lowest cost and highest performance for Mercury 2–class models.
| Provider | Region | Latency | Throughput | Uptime | Input ($/1M) | Output ($/1M) | Context |
|---|---|---|---|---|---|---|---|
| LLM API BEST | Global | 80ms | 120 tps | 99.99% | $0.40 | $0.80 | 128K tokens |
| Inception | Global | ~150ms | ~60 tps | ~99.9% | ~$0.80 | ~$1.60 | ~64K tokens |
| OpenAI | Global | ~160ms | ~70 tps | ~99.9% | ~$1.00 | ~$2.00 | ~128K tokens |
| Anthropic | US East | ~170ms | ~50 tps | ~99.9% | ~$1.20 | ~$2.40 | ~200K tokens |
Performance benchmarks
| Metric | Mercury 2 (Inception) | GPT-4.1 (OpenAI) | Claude 3.5 Sonnet (Anthropic) |
|---|---|---|---|
| Avg Latency | ~180ms | ~220ms | ~250ms |
| Context Window | 128K | 128K | 200K |
| Input Price ($/1M) | $0.80 | $5.00 | $3.00 |
| Output Price ($/1M) | $2.40 | $15.00 | $15.00 |
| Max Output Tokens | 8K | 4K | 8K |
| Throughput | 80 tps | 40 tps | 50 tps |
| Uptime | 99.9% | 99.9% | 99.9% |
30-day usage via LLM API
Architecture & Integration
One unified API. Every major model. Built-in reliability, cost control, and observability.
Automatically route each request to the best model across providers based on latency, cost, and quality—without changing your integration or redeploying code.
One endpoint, every modelControl and predict spend with per-route pricing policies, budget guards, and automatic downshifts to cheaper models when quality thresholds are still met.
Optimize every tokenSurvive provider outages and rate limits with built-in multi-region, multi-model failover so your app keeps responding even when an upstream service doesn’t.
Always-on reliabilityTrace every request across models and providers with logs, latency breakdowns, and error analytics to debug faster and continuously tune your routing rules.
See every token hopUse high-level task APIs for chat, RAG, tools, and more so you can swap underlying models without rewriting prompts or business logic.
Tasks, not raw callsProcess massive workloads efficiently with parallelized, rate-limit-aware batch execution, automatic retries, and deduplicated inputs for lower cost and higher throughput.
Ship at batch scaleDecision guide
FAQ
Mercury 2 is an Inception large language model accessible via LLM.API, designed for fast, cost-efficient general-purpose text generation and reasoning.
Mercury 2 is best for code generation, step-by-step reasoning, chatbot-style conversations, and structured text transformations like summarization or extraction.
Mercury 2 supports a 32K token context window, allowing it to handle long documents, multi-step tools, and extended conversations reliably.
Mercury 2 is optimized for low p95 latency and high token throughput, making it suitable for interactive applications and high-traffic backends.
Mercury 2 currently supports text input and text output only, with no native image, audio, or video processing.
Mercury 2 uses LLM.API’s unified token-based pricing, with separate rates for input and output tokens configurable per project in your LLM.API dashboard.
Use the chat or completions endpoint with `model` set to `inception/mercury-2`, passing your prompt, optional system instructions, and any tool definitions.
Mercury 2 targets a balance of quality and speed, typically trading slightly lower peak capability for materially lower cost and latency.
Mercury 2 supports JSON-structured outputs and standard tool or function-calling semantics via LLM.API’s unified tool-calling interface.
Mercury 2 can hallucinate facts, lacks real-time knowledge or browsing, and is not suitable for safety-critical or compliance-required decision-making without human review.
Compare
Body Builder (beta) is an OpenRouter model that converts natural language descriptions into structured OpenRouter API request objects, enabling automated construction of complex, multi-model calls.
Claude Opus 4.7 is Anthropic’s most capable generally available large language model, designed for advanced coding, long-horizon agentic workflows, and high-resolution vision tasks. It emphasizes stronger…
Palmyra X5 is Writer's most advanced enterprise large language model, featuring an extremely long context window and adaptive reasoning for complex business workflows. It is purpose-built…