- Text Generation
MiniMax M2 is an open‑weight Mixture‑of‑Experts large language model from MiniMax, designed to deliver high coding and agentic workflow performance with low latency and cost. It…
Powered by Anthropic
Claude Sonnet 4.6 is Anthropic’s most capable Sonnet‑tier large language model, offering Opus‑class performance in coding, computer use, and long‑context reasoning with a 1 million token context window in beta.
Output tokens per second · Higher is better
Seconds · Lower is better
USD per 1M tokens (blended) · Lower is better
About the model
Claude Sonnet 4.6 is a multimodal large language model from Anthropic designed to balance high intelligence with speed and cost efficiency. It is used for software development and debugging, long‑horizon knowledge work and planning, and interacting with real computer environments by navigating applications and documents. It also supports design, analysis, and other general assistant tasks over very long contexts. Claude Sonnet 4.6 belongs to the Claude Sonnet family in Anthropic’s Claude model series, succeeding earlier Sonnet 4.x generations such as Sonnet 4.5.
Model capabilities
Engages in multi-turn dialogue, following complex instructions, maintaining context, and adapting tone for assistance, analysis, and brainstorming.
Interprets images by identifying objects, text, layout, and visual relationships to support descriptions, analysis, and reasoning tasks.
Translates between major languages, preserving meaning and style for instructions, explanations, and general-purpose multilingual communication.
Extracts machine-readable text from images or document photos, enabling downstream search, summarization, and editing workflows.
Helps interpret logs, metrics, and alerts conceptually, supporting troubleshooting and analysis of technical systems when given textual telemetry.
Use cases
Transparent pricing
Save up to ~55% vs. standard Claude Sonnet 4.6 API pricing.
| Provider | Region | Latency | Throughput | Uptime | Input ($/1M) | Output ($/1M) | Context |
|---|---|---|---|---|---|---|---|
| LLM API BEST | Global | ~120ms | ~80 tps | 99.99% | $0.80 | $4.00 | 200K |
| Anthropic | US East | ~220ms | ~40 tps | 99.9% | ~$1.80 | ~$9.00 | 200K |
| AWS Bedrock | US West | ~260ms | ~35 tps | 99.9% | ~$2.00 | ~$10.00 | 200K |
| Google Cloud Vertex AI | Global | ~250ms | ~30 tps | 99.9% | ~$2.10 | ~$10.50 | 200K |
Performance benchmarks
| Metric | Claude Sonnet 4.6 | GPT-4.1 Mini | Gemini 1.5 Flash |
|---|---|---|---|
| Avg Latency | ~180ms | ~220ms | ~250ms |
| Context Window | 200K | 128K | 1M |
| Input Price ($/1M) | $0.20 | $0.15 | $0.20 |
| Output Price ($/1M) | $0.80 | $0.60 | $0.60 |
| Max Output Tokens | 8K | 4K | 8K |
| Throughput | 80 tps | 100 tps | 90 tps |
| Uptime | 99.9% | 99.9% | 99.9% |
30-day usage via LLM API
Architecture & Integration
One unified API. Every major model. Built-in reliability, cost control, and observability.
Automatically route each request to the optimal model across providers based on latency, cost, and quality—without changing your integration or client code.
One endpoint. Any model.Control spend with price caps, smart model selection, and usage controls so you can experiment freely while keeping production costs predictable and optimized.
Optimize quality per dollar.Define automatic failover chains so requests recover from provider outages, rate limits, or timeouts—without shipping new code or impacting end users.
Stay up, even when they’re down.Get full visibility into every call—latency, errors, cost, and provider breakdowns—so you can debug faster, tune prompts, and prove performance to stakeholders.
See every token’s journey.Use high-level task APIs for chat, tools, RAG, and structured outputs instead of wiring raw providers, cutting boilerplate while keeping full config control.
Think in tasks, not providers.Run large-scale generations, evaluations, and enrichments via optimized batch execution with concurrency controls, retries, and cost tracking built in.
Scale from one to millions.Decision guide
FAQ
Claude Sonnet 4.6 is an Anthropic large language model optimized for balanced cost, quality, and speed across coding, chat, and analysis tasks.
Claude Sonnet 4.6 excels at multi-step reasoning, code generation and refactoring, data analysis, and high-quality conversational agents with moderate latency and cost.
Claude Sonnet 4.6 supports context windows up to 200,000 tokens, enabling long documents, multi-file codebases, and complex workflows in a single request.
Through LLM.API, Claude Sonnet 4.6 supports text input and output, and image inputs for vision-language tasks where enabled by your LLM.API plan.
Claude Sonnet 4.6 generally returns first tokens within a few hundred milliseconds to a couple seconds, depending on prompt size and LLM.API region.
Claude Sonnet 4.6 uses a per-token billing model on LLM.API, with separate input and output token rates defined in LLM.API’s pricing schedule.
You select the model identifier for Claude Sonnet 4.6 in your LLM.API completion or chat endpoint request and send prompts using the standard JSON schema.
Claude Sonnet 4.6 typically offers lower cost and latency than flagship Claude models, with slightly reduced peak reasoning depth and creativity.
Claude Sonnet 4.6 can still hallucinate facts, mishandle very domain-specific edge cases, and should not be used without human review for high-stakes decisions.
Yes, Claude Sonnet 4.6 can stream tokens incrementally through LLM.API by enabling the streaming option on your request.
Compare
MiniMax M2 is an open‑weight Mixture‑of‑Experts large language model from MiniMax, designed to deliver high coding and agentic workflow performance with low latency and cost. It…
Ministral 3 14B 2512 is a 14-billion-parameter AI language model from Mistral’s Ministral 3 series, configured with a 2,512-dimensional internal representation. It is designed to provide…
Gemma 4 26B A4B is a 26-billion-parameter multimodal Mixture-of-Experts model from Google’s Gemma 4 family, optimized for high-throughput reasoning with long context windows. It supports text…