- Instruction Following
Free Models Router is an OpenRouter meta-model that automatically routes requests to compatible free models, providing no-cost inference across multiple underlying LLMs. It filters candidates based…
Powered by xAI
Grok Build 0.1 is xAI’s fast, agentic coding model optimized for software engineering workflows, with a 256K-token context window and support for text and image inputs.
Output tokens per second · Higher is better
Seconds · Lower is better
USD per 1M tokens (blended) · Lower is better
About the model
Grok Build 0.1 is xAI’s coding-focused language model designed for agentic software development tasks such as web development, debugging, and multi-step code planning. It is mainly used to power interactive coding agents that perform planning, tool use, and function calling over long contexts, and can also serve as a cost‑effective general-purpose model for structured outputs and automation workflows. Grok Build 0.1 succeeds earlier xAI coding models like grok-code-fast-1 and belongs to the broader Grok model family.
Model capabilities
Specialized in autonomous, multi-step coding workflows including planning, implementing, refactoring, and iterating on software projects and features.
Generates and updates front-end and back-end web application code, scaffolds projects, and helps integrate common frameworks, libraries, and APIs.
Analyzes error messages, stack traces, and failure cases to locate bugs, propose fixes, and improve overall code reliability and maintainability.
Supports function calling and MCP-style tool integration, enabling automated interaction with external APIs, services, and developer tooling pipelines.
Accepts images like diagrams, UI mockups, or error screenshots to infer structure and generate or adjust corresponding implementation code.
Use cases
Transparent pricing
LLM API offers the lowest cost and highest performance for Grok-class models.
| Provider | Region | Latency | Throughput | Uptime | Input ($/1M) | Output ($/1M) | Context |
|---|---|---|---|---|---|---|---|
| LLM API BEST | Global | 80ms | 120 tps | 99.99% | $0.20 | $0.60 | 256K |
| xAI | US | ~250ms | ~40 tps | ~99.9% | ~$5.00 | ~$15.00 | ~128K |
| OpenAI (GPT-4o class) | Global | ~300ms | ~35 tps | 99.9% | ~$2.50 | ~$10.00 | 128K |
| Anthropic (Claude 3 class) | US East | ~320ms | ~30 tps | 99.9% | ~$3.00 | ~$15.00 | 200K |
| Google (Gemini 1.5 Pro class) | Global | ~350ms | ~25 tps | ~99.9% | ~$4.00 | ~$12.00 | ~128K |
Performance benchmarks
| Metric | Grok Build 0.1 (xAI) | GPT-4o mini (OpenAI) | Claude 3.5 Haiku (Anthropic) |
|---|---|---|---|
| Model Type | Coding-optimized LLM | General-purpose LLM | General-purpose LLM |
| Context Window | 256K tokens | 128K tokens | 200K tokens |
| Input Price ($/1M tokens) | $1.00 | $0.15 | $0.80 |
| Output Price ($/1M tokens) | $2.00 | $0.60 | $4.00 |
| Max Output Tokens | — | — | 8K |
| Throughput | ≥100 tokens/s | — | — |
| Avg Latency | Low (coding-optimized) | Low | Very low |
| Uptime (API SLA) | — | — | — |
30-day usage via LLM API
Architecture & Integration
One unified API. Every major model. Built-in reliability, cost control, and observability.
Automatically route each request to the best model by cost, latency, and capability. One endpoint abstracts away provider churn and manual model selection.
One endpoint, smart routingOptimize spend with per-request cost controls, price-aware routing, and detailed usage insights so you can ship faster without surprise bills or manual tuning.
Control and cut costsDefine automatic fallbacks across providers and models to survive outages and rate limits, keeping your production workloads online without custom retry code.
Built-in reliability layerTrace every call across providers with logs, metrics, and latency breakdowns, so you can debug, optimize, and benchmark models from a single pane of glass.
See every token, everywhereWork at the level of tasks—chat, tools, RAG, workflows—instead of raw APIs, so you can swap underlying models without rewriting application logic.
Code to tasks, not APIsRun massive batch inference across providers with automatic chunking, retries, and progress tracking, maximizing throughput while staying within limits and budgets.
Scale jobs, not scriptsDecision guide
FAQ
Grok Build 0.1 is an xAI language model accessible through LLM.API for fast, general-purpose text generation and reasoning tasks.
Grok Build 0.1 is best for rapid prototyping, chat-style assistants, and tools requiring concise reasoning over medium-length inputs.
Grok Build 0.1 supports a 32K token context window through LLM.API for prompts plus generated output combined.
Grok Build 0.1 generally returns first tokens within a few hundred milliseconds, with full responses depending on output length and load.
Grok Build 0.1 currently supports text input and text output only when accessed via LLM.API.
Grok Build 0.1 is billed per token through LLM.API, with separate rates for input and output tokens defined in your LLM.API pricing plan.
Set the model parameter to "xai:grok-build-0.1" in your LLM.API request and authenticate with your LLM.API key as usual.
Grok Build 0.1 typically offers lower cost and latency than frontier models but with reduced peak reasoning depth and nuanced instruction following.
Grok Build 0.1 can hallucinate facts, struggle with very long multi-step reasoning, and should not be used as a sole source for critical decisions.
Yes, Grok Build 0.1 supports server-sent event streaming on LLM.API when you enable streaming in the request options.
Compare
Free Models Router is an OpenRouter meta-model that automatically routes requests to compatible free models, providing no-cost inference across multiple underlying LLMs. It filters candidates based…
Qwen3.7 Max is a large language model from Qwen optimized for powerful, general-purpose reasoning and coding assistance. It is designed to handle complex, multi-step tasks with…
Step 3.5 Flash is StepFun’s sparse Mixture-of-Experts language model that delivers frontier-level reasoning and agentic capabilities while remaining highly efficient and fast for production use.