- Instruction Following
Step 3.7 Flash is StepFun’s latest high-efficiency multimodal Mixture-of-Experts vision-language model, optimized for enterprise-scale agentic, coding, and long-context reasoning workloads.
Powered by Google
Gemma 4 26B A4B (free) is a 26-billion-parameter variant in Google’s Gemma 4 family, offered with an A4B quantization profile for more efficient inference. It is accessible for free use, targeting capable reasoning and generation while reducing hardware requirements.
Output tokens per second · Higher is better
Seconds · Lower is better
USD per 1M tokens (blended) · Lower is better
About the model
Gemma 4 26B A4B (free) is a quantized 26B-parameter large language model from Google’s Gemma 4 series, optimized for efficient deployment. It is mainly used for general-purpose chat, code assistance, and text generation tasks where a strong medium‑sized model is suitable. It also supports applications such as prototyping AI agents, educational tools, and lightweight research workflows on constrained compute. It belongs to the Gemma model family, which follows earlier Gemma generations designed as open, efficient LLMs from Google.
Model capabilities
Handles multi-turn conversations, follows instructions, and maintains context to provide coherent, helpful responses on many general topics.
Understands long-form text, summarizes content, extracts key information, and answers questions based on provided documents or prompts.
Helps write, explain, and refactor code snippets in popular programming languages, aiding debugging and implementation of common patterns.
Translates between major natural languages and explains wording choices, tone, and nuances while preserving meaning and style.
Analyzes uploaded images, describing scenes and objects, reading visible text, and supporting reasoning about visual content when available.
Use cases
Transparent pricing
LLM API offers the lowest cost per 1M tokens and fastest Gemma 4–class inference.
| Provider | Region | Latency | Throughput | Uptime | Input ($/1M) | Output ($/1M) | Context |
|---|---|---|---|---|---|---|---|
| LLM API BEST | Global | 120ms | 120 tps | 99.99% | $0.02 | $0.04 | 256K |
| Global | ~220ms | ~70 tps | 99.9% | $0.00 | $0.00 | ~128K | |
| Vertex AI (Google Cloud) | US East | ~260ms | ~60 tps | 99.9% | ~$0.20 | ~$0.40 | ~128K |
| Together AI | US West | ~240ms | ~80 tps | 99.9% | ~$0.15 | ~$0.30 | ~64K |
| Groq | US Central | ~150ms | ~100 tps | 99.9% | ~$0.10 | ~$0.20 | ~32K |
Performance benchmarks
| Metric | Gemma 4 26B A4B (free) | Gemini 2.0 Flash (Google) | GPT-4.1 Mini (OpenAI) |
|---|---|---|---|
| Avg Latency | ~220ms | ~180ms | ~200ms |
| Context Window | 128K | 128K | 128K |
| Input Price ($/1M) | $0.00 | $0.20 | $0.15 |
| Output Price ($/1M) | $0.00 | $0.60 | $0.60 |
| Max Output Tokens | 4K | 4K | 4K |
| Throughput | ~80 tps | ~100 tps | ~90 tps |
| Uptime | 99.5% | 99.9% | 99.9% |
30-day usage via LLM API
Architecture & Integration
One unified API. Every major model. Built-in reliability, cost control, and observability.
Automatically route each request to the optimal model across providers based on latency, cost, and capability—without changing your application code or integration logic.
One endpoint, any modelControl spend with configurable pricing policies, transparent per-request cost breakdowns, and smart routing that prefers cheaper equivalents when quality and latency requirements are met.
Lower cost, same outputDefine automatic provider and model fallbacks so production traffic keeps flowing through alternative backends when a model, region, or vendor has performance or availability issues.
No single point of failureGet end-to-end traces, metrics, and structured logs for every call, enabling fast debugging, performance tuning, and regression detection across all models and providers.
See every token, everywhereDeclare tasks like chat, tools, rerank, or embeddings once, and let LLM.API handle provider-specific quirks, payload shaping, and response normalization automatically.
One task spec, many modelsRun large-scale jobs with parallelized, rate-limit-aware batching, automatic retries, and progress tracking—maximizing throughput while protecting upstream providers and your application.
Ship millions of calls safelyDecision guide
FAQ
Gemma 4 26B A4B (free) is a 26B-parameter Google Gemma 4 language model variant accessible via LLM.API at no usage cost.
It is best for high-quality general text generation, reasoning, and coding assistance when you want a strong model without incurring API charges.
The model is offered with a $0 per-token price, subject to LLM.API’s free-tier rate limits and fair-use policies.
Gemma 4 26B A4B (free) supports a 32K token context window for combined prompt and completion.
Latency is moderate, typically slower than smaller models but acceptable for interactive use, depending on request size and current platform load.
Gemma 4 26B A4B (free) is a text-only model, accepting and producing UTF-8 text but not images, audio, or video.
Use the LLM.API chat or completion endpoint with the provider set to "google" and the model name set exactly to "Gemma 4 26B A4B (free)".
It offers competitive quality for many tasks but may trail frontier paid models in complex reasoning, instruction-following robustness, and safety tuning.
It can hallucinate facts, lacks real-time knowledge, is text-only, and may be subject to strict rate limits due to its free status.
Yes, LLM.API enforces per-minute and daily rate limits for the free model that may throttle or reject excessive traffic.
Compare
Step 3.7 Flash is StepFun’s latest high-efficiency multimodal Mixture-of-Experts vision-language model, optimized for enterprise-scale agentic, coding, and long-context reasoning workloads.
Seed 1.6 Flash is an ultra-fast multimodal "deep thinking" large language model from ByteDance Seed, offering long-context reasoning with support for both text and visual inputs.
GLM 5 is Z.ai’s fifth-generation large language model, a large open-source Mixture-of-Experts foundation model focused on advanced reasoning and long-horizon agent workflows. It is notable for…