- Text Generation
Riverflow V2 Pro is Sourceful’s most powerful Riverflow 2.0 model, focused on high-quality, controllable image generation and perfect text rendering.
Powered by xAI
Grok 4.20 is xAI’s flagship large language model designed for high-speed inference, low hallucination rates, and strong agentic tool-calling for complex tasks.
Output tokens per second · Higher is better
Seconds · Lower is better
USD per 1M tokens (blended) · Lower is better
About the model
Grok 4.20 is a flagship large language model from xAI focused on fast, reliable reasoning with multiple internal agents to improve answer quality. It is primarily used for advanced chat-based assistance, complex reasoning tasks, and agentic workflows where it can orchestrate tools and APIs. It is also deployed in enterprise and developer platforms via APIs and partner integrations for building applications that need structured output, function calling, and multimodal (text and image) understanding. It succeeds earlier Grok 4-series models and builds on the broader Grok family of xAI language models.
Model capabilities
Engages in multi-turn dialogue, answering questions and following instructions with contextual awareness and controllable tone and style.
Understands and generates code snippets, and can reason about using external tools or APIs when appropriately integrated.
Interprets images to identify objects and visual patterns, supporting question answering and basic visual understanding tasks.
Translates between multiple major languages while maintaining meaning and style across diverse informal and formal text inputs.
Extracts readable text and structured information from documents or images, enabling downstream processing and analysis.
Use cases
Transparent pricing
LLM API offers the lowest cost and best performance for Grok‑class models.
| Provider | Region | Latency | Throughput | Uptime | Input ($/1M) | Output ($/1M) | Context |
|---|---|---|---|---|---|---|---|
| LLM API BEST | Global | 120ms | 80 tps | 99.99% | $0.30 | $0.60 | 128K |
| xAI | Global | ~450ms | ~35 tps | ~99.9% | ~$0.80 | ~$1.60 | ~128K |
| OpenAI | Global | ~400ms | ~40 tps | ~99.9% | ~$0.75 | ~$1.50 | ~128K |
| Anthropic | US East | ~420ms | ~30 tps | ~99.9% | ~$0.85 | ~$1.70 | ~200K |
Performance benchmarks
| Metric | Grok 4.20 (xAI) | GPT-4.1 (OpenAI) | Claude 3.5 Sonnet (Anthropic) |
|---|---|---|---|
| Avg Latency | ~700ms | ~900ms | ~850ms |
| Context Window | 128K | 128K | 200K |
| Input Price ($/1M) | $2.00 | $5.00 | $3.00 |
| Output Price ($/1M) | $5.00 | $15.00 | $15.00 |
| Max Output Tokens | 8K | 4K | 4K |
| Throughput | 40 tps | 30 tps | 25 tps |
| Uptime | 99.5% | 99.9% | 99.9% |
30-day usage via LLM API
Architecture & Integration
One unified API. Every major model. Built-in reliability, cost control, and observability.
Automatically route each request to the best model across providers based on latency, cost, and quality—without changing your integration or application code.
One endpoint, all modelsEnforce budgets, compare provider pricing, and downgrade to cheaper models when quality thresholds are met so you never overspend on inference again.
Control spend by designDefine failover chains so requests automatically retry on alternative models or providers, preventing downtime and degraded UX when a single vendor has issues.
No single point of failureCapture structured logs, metrics, and traces for every call across providers, making it easy to debug failures, tune prompts, and optimize performance in production.
See every token, everywhereDescribe tasks like chat, completion, tools, or rerank once and let LLM.API pick the right models and parameters for each use case automatically.
Think in tasks, not modelsRun massive, parallel LLM workloads with built-in queuing, rate-limit handling, and retries so you can process millions of items reliably and cost-effectively.
Scale jobs without glue codeDecision guide
FAQ
Grok 4.20 is an xAI large language model accessible via LLM.API, designed for fast, general-purpose code, chat, and analysis workloads.
Grok 4.20 is best for conversational agents, code assistance, data analysis, and iterative reasoning where low latency and strong general capabilities matter.
Grok 4.20 supports up to a 128K token context window when accessed through LLM.API.
Grok 4.20 pricing is set by LLM.API and may differ from xAI direct pricing; check your LLM.API dashboard for current per-token rates.
Grok 4.20 is optimized on LLM.API for low p95 latency and streaming responses suitable for interactive applications, subject to your chosen deployment region.
Grok 4.20 supports text input and text output only when used through LLM.API.
Use the LLM.API chat or completions endpoint with the model identifier "grok-4.20" and your LLM.API key in the Authorization header.
Grok 4.20 targets competitive reasoning and coding quality at generally lower cost and latency than many flagship general-purpose models on LLM.API.
Grok 4.20 can hallucinate facts, lacks real-time knowledge, and should not be solely relied on for safety-critical, legal, or medical decisions.
Yes, Grok 4.20 can use LLM.API’s tool or function-calling interface when you define tools in the request schema.
Yes, Grok 4.20 can be used for batch processing through LLM.API, but you must respect rate limits and maximum tokens per request.
Compare
Riverflow V2 Pro is Sourceful’s most powerful Riverflow 2.0 model, focused on high-quality, controllable image generation and perfect text rendering.
Qwen3 VL 30B A3B Instruct is a 30B-parameter Mixture-of-Experts vision-language model from Qwen, offering strong multimodal understanding and generation with a 262K-token context window. It is…
Nano Banana 2 (Gemini 3.1 Flash Image Preview) is Google DeepMind’s image generation and editing model built on the Gemini 3.1 Flash architecture, optimized for fast,…