- Text Generation
Video v3.0 Standard by Kling is a text-to-video and image-to-video generation model that produces cinematic, multi-shot clips with optional native audio. It offers up to roughly…
Powered by LiquidAI
LFM2-24B-A2B is LiquidAI’s largest LFM2-series hybrid Mixture-of-Experts language model, designed to deliver high-quality text generation while remaining efficient enough to run on consumer hardware.
Output tokens per second · Higher is better
Seconds · Lower is better
USD per 1M tokens (blended) · Lower is better
About the model
LFM2-24B-A2B is a 24B-parameter sparse Mixture-of-Experts hybrid language model from LiquidAI, with about 2B active parameters per token and a context window of around 128K tokens. It is primarily used for general-purpose text generation tasks such as drafting, summarization, and chat-style assistance, with a focus on low-cost inference. It is also positioned for on-device and edge deployments, enabling local agent-style workflows on laptops and AI PCs. It belongs to the LFM2 family of models, extending the series from smaller variants (e.g., LFM2-350M and mid-sized LFM2 models) up to this largest 24B configuration.
Model capabilities
Engages in multi-turn dialogue, answering questions, following instructions, and adapting responses to user context and intent.
Analyzes images to identify objects, scenes, and relationships, enabling visual question answering and descriptive explanations.
Translates written content between multiple languages while preserving meaning, tone, and stylistic nuance as closely as possible.
Extracts machine-readable text from documents and images, enabling downstream search, summarization, and content analysis workflows.
Supports monitoring-style tasks such as interpreting logs, alerts, and metrics to assist with diagnostics and incident summaries.
Use cases
Transparent pricing
LLM API offers the lowest cost and highest performance for LFM2-24B-A2B-class models.
| Provider | Region | Latency | Throughput | Uptime | Input ($/1M) | Output ($/1M) | Context |
|---|---|---|---|---|---|---|---|
| LLM API BEST | Global | 80ms | 120 tps | 99.99% | $0.40 | $0.80 | 256K |
| LiquidAI | US East | ~140ms | ~70 tps | ~99.9% | ~$0.65 | ~$1.30 | ~128K |
| OpenAI (comparable 20–30B model) | Global | ~200ms | ~60 tps | ~99.9% | ~$1.00 | ~$2.00 | ~128K |
| Anthropic (comparable 20–30B model) | US West | ~190ms | ~55 tps | ~99.9% | ~$1.10 | ~$2.20 | ~200K |
| Azure AI (LiquidAI-compatible deployment) | EU West | ~210ms | ~50 tps | ~99.95% | ~$0.90 | ~$1.80 | ~128K |
Performance benchmarks
| Metric | LFM2-24B-A2B (LiquidAI) | GPT-4.1-mini (OpenAI) | Claude 3.5 Haiku (Anthropic) |
|---|---|---|---|
| Avg Latency | ~180ms | ~220ms | ~250ms |
| Context Window | 128K | 128K | 200K |
| Input Price ($/1M) | $0.10 | $0.15 | $0.25 |
| Output Price ($/1M) | $0.40 | $0.60 | $0.80 |
| Max Output Tokens | 8K | 4K | 8K |
| Throughput | 120 tps | 100 tps | 90 tps |
| Uptime | 99.9% | 99.9% | 99.9% |
30-day usage via LLM API
Architecture & Integration
One unified API. Every major model. Built-in reliability, cost control, and observability.
Dynamically route each request across providers by latency, price, and quality. One endpoint abstracts vendor lock-in and keeps workloads on the best option automatically.
One endpoint, every modelAutomatically balance quality and spend with per-request cost controls, usage caps, and cheaper alternates. Ship rich AI features without blowing your infrastructure budget.
Optimize cost per tokenDefine provider and model fallbacks that trigger instantly on timeouts, rate limits, or errors. Keep user-facing experiences stable even when vendors fail.
Failures auto-reroutedTrace every request across models and providers with logs, metrics, and latency breakdowns. Debug production issues fast and tune routing using real traffic data.
See every token hopDescribe tasks—not models—and let LLM.API pick the right tools, prompts, and providers. Standardize patterns like chat, tools, and RAG behind one API.
Program tasks, not modelsSend large batches of requests in a single call with concurrency controls and retry policies. Maximize throughput and minimize overhead for heavy workloads.
Scale up without thrashDecision guide
FAQ
LFM2-24B-A2B is a 24B-parameter LiquidAI language model available through LLM.API, designed for high-quality text generation and reasoning tasks.
LFM2-24B-A2B is best for complex code generation, multi-step reasoning, data transformation, and longer-form content where quality matters more than minimal latency.
LFM2-24B-A2B is a text-only model that accepts text prompts and returns text completions.
LFM2-24B-A2B supports up to a 32K-token context window via LLM.API, including input and output tokens combined.
LFM2-24B-A2B targets stronger reasoning and coding quality than typical 7–14B models, with higher cost but better performance on complex tasks.
LFM2-24B-A2B has moderate first-token latency typical of 20–30B models, but streams tokens quickly enough for interactive applications.
LFM2-24B-A2B uses a per-token pricing model on LLM.API, with separate input and output token rates defined in the LLM.API pricing page.
Specify the model ID "LFM2-24B-A2B" in your LLM.API completion or chat endpoint request, along with your API key and usual parameters.
LFM2-24B-A2B can be prompted to emit structured JSON, but native function-calling semantics depend on LLM.API’s tooling layer, not the model itself.
LFM2-24B-A2B can hallucinate facts, lacks real-time knowledge, and may struggle with highly specialized domain data without careful prompting or retrieval.
Compare
Video v3.0 Standard by Kling is a text-to-video and image-to-video generation model that produces cinematic, multi-shot clips with optional native audio. It offers up to roughly…
Body Builder (beta) is an OpenRouter model that converts natural language descriptions into structured OpenRouter API request objects, enabling automated construction of complex, multi-model calls.
all-MiniLM-L12-v2 is a compact Sentence Transformers model that generates high-quality sentence embeddings for efficient semantic search and similarity tasks. It is notable for its strong performance-to-size…