- Text Generation
LFM2-24B-A2B is LiquidAI’s largest LFM2-series hybrid Mixture-of-Experts language model, designed to deliver high-quality text generation while remaining efficient enough to run on consumer hardware.
Powered by ~Anthropic
Anthropic Claude Sonnet Latest refers to the most recent mid-tier Claude Sonnet language model from Anthropic, designed to balance strong intelligence with speed and cost-efficiency. It is commonly used as Anthropic’s default general-purpose assistant model in the Claude product and API.
Output tokens per second · Higher is better
Seconds · Lower is better
USD per 1M tokens (blended) · Lower is better
About the model
Anthropic Claude Sonnet Latest is a production-grade large language model in Anthropic’s Claude Sonnet series, positioned as the balanced, mid-tier option between smaller Haiku and larger Opus models. It is mainly used for general-purpose chat assistants, writing and analysis, and knowledge work that require strong reasoning at lower latency and cost than flagship frontier models. It is also widely used for coding, tool use, and enterprise applications that need long-context processing and robust safety at scale. It belongs to Anthropic’s Claude model family, which is organized into Opus (flagship), Sonnet (balanced), and Haiku (lightweight) tiers that have evolved through multiple generations such as Claude 3.x and 4.x Sonnet.
Model capabilities
Engages in multi-turn, context-aware conversations, following complex instructions and maintaining coherent, helpful dialogue across diverse topics.
Interprets images to identify objects, scenes, and relationships, supporting tasks like description, comparison, and visual context reasoning.
Translates between multiple languages, preserving meaning and tone for general-purpose content, instructions, and user queries.
Extracts and structures text from images or document photos, enabling search, summarization, and downstream processing of visual text content.
Understands and writes code, reasons step-by-step, and coordinates use of external tools or APIs when integrated into applications.
Use cases
Transparent pricing
LLM API offers the lowest cost and fastest access to Claude Sonnet–class models.
| Provider | Region | Latency | Throughput | Uptime | Input ($/1M) | Output ($/1M) | Context |
|---|---|---|---|---|---|---|---|
| LLM API BEST | Global | ~180ms | ~120 tps | 99.99% | $0.60 | $1.80 | 200K |
| Anthropic | US East | ~350ms | ~60 tps | 99.9% | ~$3.00 | ~$15.00 | 200K |
| Amazon Bedrock (Anthropic Claude Sonnet equivalent) | US West | ~420ms | ~45 tps | 99.9% | ~$3.20 | ~$16.00 | 200K |
| Google Cloud (Anthropic Claude Sonnet equivalent) | Global | ~400ms | ~50 tps | 99.9% | ~$3.40 | ~$17.00 | 200K |
| Azure (Anthropic Claude Sonnet equivalent) | EU West | ~380ms | ~55 tps | 99.9% | ~$3.60 | ~$18.00 | 200K |
Performance benchmarks
| Metric | Anthropic Claude Sonnet Latest | OpenAI GPT-4.1 Mini | Google Gemini 1.5 Flash |
|---|---|---|---|
| Avg Latency | ~250ms | ~220ms | ~260ms |
| Context Window | 200K | 128K | 1M |
| Input Price ($/1M tokens) | $0.80 | $0.30 | $0.35 |
| Output Price ($/1M tokens) | $4.00 | $1.25 | $1.50 |
| Max Output Tokens | 4K | 4K | 8K |
| Throughput | 45 tps | 50 tps | 40 tps |
| Uptime | 99.9% | 99.9% | 99.9% |
30-day usage via LLM API
Architecture & Integration
One unified API. Every major model. Built-in reliability, cost control, and observability.
Automatically route each request to the best-fit model across providers based on latency, cost, and quality—no client changes required as your stack evolves.
One endpoint, every modelControl and predict spend with transparent pricing, per-provider budgets, and cost-based routing policies that keep experiments fast while production remains under budget.
Optimize every tokenDesign multi-step failover strategies so if a provider degrades or times out, requests automatically retry on backup models without impacting your application.
Never drop a requestGet centralized traces, metrics, and logs for every call across all providers, enabling rapid debugging, performance tuning, and regression detection from a single dashboard.
See every token hopDefine high-level tasks like chat, tools, or embeddings once, then swap underlying models or providers freely without rewriting business logic or prompt plumbing.
Code to tasks, not modelsRun large-scale inference workloads with parallelized, rate-aware batching that maximizes throughput, minimizes costs, and abstracts provider-specific batch quirks.
Ship batch at scaleDecision guide
FAQ
Anthropic Claude Sonnet Latest is a balanced, general-purpose Claude 3.5 family model from ~Anthropic, exposed through the LLM.API unified gateway.
Anthropic Claude Sonnet Latest supports up to a 200K token context window, suitable for long documents, multi-step tools, and complex conversations.
It excels at high‑quality reasoning, coding assistance, multi-step problem solving, and robust general chat while offering better cost‑performance than flagship models.
Pricing is metered per 1,000 tokens for input and output; check the LLM.API pricing page for the latest Anthropic Claude Sonnet rates.
Latency depends on load and request size, but Sonnet typically offers mid‑range response times faster than Opus‑class models and slower than Haiku‑class models.
Anthropic Claude Sonnet Latest supports text input and output, and can process images when configured for multimodal use via compatible LLM.API endpoints.
Use the LLM.API endpoint with the model identifier for Anthropic Claude Sonnet Latest, passing your prompt, optional system instructions, and tool configuration if needed.
Sonnet generally offers similar reasoning quality at lower cost and latency than Opus‑class models but with slightly reduced peak capability on the hardest tasks.
Yes, when configured in LLM.API, it can consume structured tool definitions and return arguments for function calls to integrate external tools or APIs.
It can still hallucinate, lacks real‑time internet access without tools, and may underperform specialized or larger models on highly technical or domain‑specific tasks.
Compare
LFM2-24B-A2B is LiquidAI’s largest LFM2-series hybrid Mixture-of-Experts language model, designed to deliver high-quality text generation while remaining efficient enough to run on consumer hardware.
Claude Opus 4.7 is Anthropic’s most capable generally available large language model, designed for advanced coding, long-horizon agentic workflows, and high-resolution vision tasks. It emphasizes stronger…
Relace Search is a text-only large language model from Relace optimized for agentic multi-step search over large codebases, using parallel file-inspection tools to return highly relevant…