- Instruction Following
LFM2.5-1.2B-Instruct (free) is a 1.2B-parameter, instruction-tuned hybrid language model from LiquidAI, optimized for fast, on-device inference with a ~32k token context window. It offers general-purpose conversational…
Powered by Google
Gemini 3.1 Flash Lite is Google’s ultra-fast, low-cost Gemini 3-series language model optimized for high-volume, latency-sensitive applications. It prioritizes speed and cost-efficiency while still supporting multimodal understanding and configurable reasoning depth.
Output tokens per second · Higher is better
Seconds · Lower is better
USD per 1M tokens (blended) · Lower is better
About the model
Gemini 3.1 Flash Lite is a lightweight, highly cost-efficient member of Google’s Gemini 3 family designed for rapid, large-scale text and multimodal inference. It is mainly used for chatbots, customer support, and other interactive applications where low latency and high request throughput are critical. It is also used in background automation, agentic workflows, and tools that need to run many inexpensive calls while occasionally invoking deeper reasoning via its “thinking levels” controls. It belongs to the Gemini Flash lineage as the successor to earlier Flash and Flash-Lite models in the broader Gemini model family.
Model capabilities
Provides quick, low-latency conversational answers suitable for assistants, FAQ bots, and interactive applications with concise, relevant outputs.
Handles everyday text tasks such as drafting, rewriting, and summarizing short content, optimized for speed and efficiency.
Performs lightweight image inspection, enabling simple recognition or extraction tasks where high speed and low resource usage are important.
Translates short text snippets between common languages for quick understanding or localization in low-latency applications.
Extracts readable text from clear images or screenshots to support quick copying, search, or downstream text processing workflows.
Use cases
Transparent pricing
LLM API offers the lowest cost and highest performance for Gemini 3.1 Flash Lite–class models.
| Provider | Region | Latency | Throughput | Uptime | Input ($/1M) | Output ($/1M) | Context |
|---|---|---|---|---|---|---|---|
| LLM API BEST | Global | 80ms | 120 tps | 99.99% | $0.03 | $0.03 | 256K |
| Google AI Studio | Global | ~160ms | ~60 tps | 99.9% | ~$0.075 | ~$0.075 | 128K |
| Google Vertex AI | US & EU | ~190ms | ~50 tps | 99.9% | ~$0.085 | ~$0.085 | 128K |
| OpenRouter (Gemini-equivalent) | Global | ~220ms | ~45 tps | ~99.5% | ~$0.090 | ~$0.090 | ~128K |
| Custom Private Deployment | US East | ~200ms | ~40 tps | ~99.5% | ~$0.110 | ~$0.110 | ~64K |
Performance benchmarks
| Metric | Gemini 3.1 Flash Lite | GPT-4.1-mini | Claude 3.5 Haiku |
|---|---|---|---|
| Avg Latency | ~150ms | ~200ms | ~220ms |
| Context Window | 128K | 128K | 200K |
| Input Price ($/1M) | $0.05 | $0.15 | $0.25 |
| Output Price ($/1M) | $0.15 | $0.60 | $1.25 |
| Max Output Tokens | 4K | 4K | 4K |
| Throughput | ~70 tps | ~50 tps | ~40 tps |
| Uptime | 99.9% | 99.9% | 99.9% |
30-day usage via LLM API
Architecture & Integration
One unified API. Every major model. Built-in reliability, cost control, and observability.
Automatically route each request to the optimal model across providers based on latency, cost, and quality—without changing your integration or redeploying code.
One endpoint, every modelEnforce spend limits, choose cheaper equivalents, and mix premium and budget models per call so you control cost without manually tuning every request.
Cut spend, keep qualityDefine automatic failover chains so timeouts, rate limits, or provider outages transparently retry on backup models—keeping your AI features online and users unblocked.
No single point of failureTrace every request across providers with logs, metrics, and latency breakdowns so you can debug prompts, tune routing, and prove reliability in production.
See every token, everywhereCall high-level tasks—chat, generation, tools, embeddings—through a single, normalized API so you can swap underlying models without rewriting business logic.
Code to tasks, not modelsSend massive batches of requests through one pipeline with concurrency control, retries, and deduping to efficiently process workloads like evals, backfills, and indexing.
Millions of calls, one jobDecision guide
FAQ
Gemini 3.1 Flash Lite is a lightweight, cost-efficient Google model optimized for fast, high-volume inference via the Gemini API and compatible gateways like LLM.API.
It is best for latency-sensitive, high-throughput tasks like chatbots, simple agents, rapid content generation, and bulk text processing where low cost matters.
Gemini 3.1 Flash Lite typically offers a mid-sized context window, suitable for most conversational and task-oriented applications rather than very long-document reasoning.
Gemini 3.1 Flash Lite is tuned for low latency, so responses are generally faster than larger Gemini models when served through LLM.API.
LLM.API usage is billed per-token for input and output; check the LLM.API pricing page for the latest Gemini 3.1 Flash Lite rates.
Through LLM.API, Gemini 3.1 Flash Lite supports text input and output, and may support additional modalities depending on LLM.API’s current Gemini feature mapping.
Specify the model name "gemini-3.1-flash-lite" (or LLM.API’s documented alias) in your completion or chat request to route traffic to this model.
Gemini 3.1 Flash Lite is cheaper and faster but generally less capable than Gemini 3.1 Flash or Pro on complex reasoning, coding, and long-context tasks.
It may struggle with very long contexts, complex multi-step reasoning, nuanced coding tasks, and tasks requiring the strongest Gemini-series accuracy and reliability.
LLM.API typically exposes Gemini models as hosted APIs without user-level fine-tuning; use prompts, system instructions, and retrieval to specialize behavior instead.
Compare
LFM2.5-1.2B-Instruct (free) is a 1.2B-parameter, instruction-tuned hybrid language model from LiquidAI, optimized for fast, on-device inference with a ~32k token context window. It offers general-purpose conversational…
Step 3.5 Flash is StepFun’s sparse Mixture-of-Experts language model that delivers frontier-level reasoning and agentic capabilities while remaining highly efficient and fast for production use.
Gemma 4 31B is Google DeepMind’s largest Gemma 4 open-weight dense multimodal model, featuring around 31 billion parameters and strong performance on text and image understanding…