- Instruction Following
Google Gemini Flash Latest is a fast, cost‑optimized variant of Google’s Gemini family, designed to deliver high-throughput, low-latency multimodal reasoning for everyday and agent-style workloads. It…
Powered by Google
Gemma 4 31B is Google DeepMind’s largest Gemma 4 open-weight dense multimodal model, featuring around 31 billion parameters and strong performance on text and image understanding tasks. It is notable for competitive reasoning quality among open models while remaining Apache-licensed and developer-friendly.
Output tokens per second · Higher is better
Seconds · Lower is better
USD per 1M tokens (blended) · Lower is better
About the model
Gemma 4 31B is a 31-billion-parameter dense multimodal large language model from Google DeepMind that processes text and images with text outputs. It is primarily used for advanced assistant-style chat, coding help, and analytical reasoning tasks that benefit from long-context understanding. It is also applied to multimodal use cases such as image-grounded question answering and document understanding where both text and images must be interpreted together. It belongs to the Gemma 4 family of open models, which span multiple sizes from edge-oriented variants to this largest 31B configuration.
Model capabilities
Performs complex, step-by-step reasoning for difficult tasks, benefiting from an explicit thinking mode in instruction-tuned variants.
Processes text and images together, supporting tasks like document parsing, UI comprehension, charts, and general visual understanding.
Acts as a strong conversational assistant, following instructions, maintaining context, and supporting agentic workflows and tool use.
Generates, completes, and debugs source code in multiple languages, suitable for software development and technical scripting tasks.
Handles multilingual input and output across many languages, enabling translation-style tasks and cross-lingual reasoning over long context.
Use cases
Transparent pricing
LLM API offers Gemma 4 31B access at significantly lower cost and latency than major cloud providers.
| Provider | Region | Latency | Throughput | Uptime | Input ($/1M) | Output ($/1M) | Context |
|---|---|---|---|---|---|---|---|
| LLM API BEST | Global | ~140ms | ~120 tps | ~99.99% | ~$0.12 per 1M tokens | ~$0.24 per 1M tokens | ~256K tokens |
| Global | ~220ms | ~80 tps | ~99.9% | ~$0.35 per 1M tokens | ~$0.70 per 1M tokens | ~128K tokens | |
| Vertex AI (Google Cloud) | US East | ~260ms | ~60 tps | ~99.9% | ~$0.38 per 1M tokens | ~$0.76 per 1M tokens | ~128K tokens |
| AWS Bedrock (3rd‑party Gemma‑equivalent) | US East | ~250ms | ~70 tps | ~99.9% | ~$0.40 per 1M tokens | ~$0.80 per 1M tokens | ~128K tokens |
| Anthropic (Claude Sonnet‑class alternative) | Global | ~230ms | ~75 tps | ~99.9% | ~$0.50 per 1M tokens | ~$1.00 per 1M tokens | ~200K tokens |
Performance benchmarks
| Metric | Gemma 4 31B (Google) | GPT-4.1 (OpenAI) | Claude 3.5 Sonnet (Anthropic) |
|---|---|---|---|
| Avg Latency | ~220ms | ~250ms | ~260ms |
| Context Window | 128K | 128K | 200K |
| Input Price ($/1M) | $0.70 | $5.00 | $3.00 |
| Output Price ($/1M) | $2.10 | $15.00 | $15.00 |
| Max Output Tokens | 8K | 8K | 8K |
| Throughput | 80 tps | 60 tps | 55 tps |
| Uptime | 99.9% | 99.9% | 99.9% |
30-day usage via LLM API
Architecture & Integration
One unified API. Every major model. Built-in reliability, cost control, and observability.
Automatically route each request to the optimal model across providers based on latency, cost, and capability — no client changes required.
One endpoint, every modelOptimize spend with dynamic model selection, rate limiting, and usage controls that keep your AI bill predictable while preserving performance.
Lower cost, same qualityDefine cross-provider failover rules so requests automatically retry on backup models when a provider is down, slow, or throttling.
No single point of failureGet unified logs, metrics, traces, and payload sampling across all providers to debug failures, tune prompts, and monitor performance in one place.
See every token, everywhereCall high-level tasks like chat, RAG, tools, or agents without wiring each provider’s primitives yourself, so you ship features instead of glue code.
APIs speak in tasksRun large-scale inference workloads with parallel execution, retries, and progress tracking built in, without manually managing queues or worker pools.
Scale from 10 to 10MDecision guide
FAQ
Gemma 4 31B is a 31-billion-parameter Google language model focused on strong reasoning, coding, and instruction-following capabilities via the LLM.API gateway.
Gemma 4 31B supports a 32K token context window, allowing relatively long conversations and documents before older tokens are pushed out.
Gemma 4 31B is best for complex reasoning, multi-step agents, advanced coding assistance, and high-quality English writing where accuracy matters.
Gemma 4 31B is a text-only model on LLM.API, supporting text inputs and outputs but not images, audio, or video.
On LLM.API, Gemma 4 31B typically returns first tokens within a few hundred milliseconds and then streams tokens at an interactive rate.
Gemma 4 31B pricing on LLM.API is usage-based per input and output token; check the LLM.API pricing page for up-to-date rates.
You select the Gemma 4 31B model name in your LLM.API request and authenticate with your LLM.API key, without needing direct Google Cloud setup.
Compared to smaller Gemma models, Gemma 4 31B generally offers better reasoning quality and coding ability at the cost of higher latency and price.
Gemma 4 31B can hallucinate facts, lacks real-time web access, may underperform on niche domains, and is restricted to its context window.
Yes, Gemma 4 31B can reliably follow JSON or schema-like formats when prompted clearly and validated by your application logic.
Compare
Google Gemini Flash Latest is a fast, cost‑optimized variant of Google’s Gemini family, designed to deliver high-throughput, low-latency multimodal reasoning for everyday and agent-style workloads. It…
Ling-2.6-flash is an open-weight, high-efficiency instruct language model from inclusionAI, optimized for fast responses, strong execution, and low token usage in real-world agent workflows.
Hailuo 2.3 by MiniMax is a high-fidelity AI video generation model designed for realistic, cinematic 1080p clips from text or image prompts, with strong motion, physics,…