- Instruction Following
Kimi K2.6 is MoonshotAI’s open-source, 1-trillion-parameter Mixture-of-Experts multimodal model optimized for long-horizon coding, agentic tool use, and image/video understanding. It is notable for its large ~262K-token…
Powered by Google
Gemini 3.1 Flash Lite Preview is a lightweight, cost-efficient Google Gemini 3.1 series model optimized for high-throughput applications with long context and adjustable thinking levels.
Output tokens per second · Higher is better
Seconds · Lower is better
USD per 1M tokens (blended) · Lower is better
About the model
Gemini 3.1 Flash Lite Preview is a preview version of Google’s Gemini 3.1 Flash-Lite large language model, designed to offer fast, inexpensive inference while supporting long-context and multimodal tasks. It is mainly used for large-scale, latency-sensitive workloads such as chatbots, agents, and real-time assistants that need to serve many requests at low cost. It is also used for applications like document and data processing, prompt-based research assistants, and other production AI services that benefit from its long context window and configurable “thinking” budget. It belongs to the Gemini 3.x Flash/Flash-Lite family and succeeds earlier preview models like Gemini 2.5 Flash Lite Preview.
Model capabilities
Handles general-purpose conversational queries and instruction-following with low latency, optimized for high-throughput interactive applications.
Accepts text, image, audio, video, and PDF inputs while producing text outputs, enabling unified reasoning across diverse content types.
Supports executing code via tools, enabling programmatic problem solving, validation of answers, and workflow automation within applications.
Performs large-scale text extraction, summarization, and classification tasks efficiently, suitable for background processing and document workflows.
Provides fast, cost-efficient translation between multiple languages, designed for high-frequency, production-grade localization and communication workloads.
Use cases
Transparent pricing
Save up to ~70% vs major Gemini-compatible providers with consistently lower latency and higher throughput.
| Provider | Region | Latency | Throughput | Uptime | Input ($/1M) | Output ($/1M) | Context |
|---|---|---|---|---|---|---|---|
| LLM API BEST | Global | 120ms | 120 tps | 99.99% | $0.02 | $0.04 | 1M tokens |
| Global | ~220ms | ~40 tps | 99.9% | ~$0.06 | ~$0.12 | ~1M tokens | |
| Vertex AI (Google Cloud) | US East | ~260ms | ~35 tps | 99.9% | ~$0.065 | ~$0.13 | ~1M tokens |
| Third-Party Aggregator A | Global | ~250ms | ~30 tps | 99.9% | ~$0.07 | ~$0.14 | ~512K tokens |
| Third-Party Aggregator B | EU West | ~280ms | ~25 tps | 99.5% | ~$0.075 | ~$0.15 | ~512K tokens |
Performance benchmarks
| Metric | Gemini 3.1 Flash Lite Preview | GPT-4.1 mini (OpenAI) | Claude 3 Haiku (Anthropic) |
|---|---|---|---|
| Avg Latency | ~120ms | ~150ms | ~180ms |
| Context Window | 128K | 128K | 200K |
| Input Price ($/1M) | $0.05 | $0.15 | $0.25 |
| Output Price ($/1M) | $0.15 | $0.60 | $0.80 |
| Max Output Tokens | 4K | 4K | 4K |
| Throughput | ~120 tps | ~100 tps | ~80 tps |
| Uptime | 99.9% | 99.9% | 99.9% |
30-day usage via LLM API
Architecture & Integration
One unified API. Every major model. Built-in reliability, cost control, and observability.
Dynamically route each request to the optimal model across providers based on latency, cost, and quality—without changing your integration or redeploying code.
One endpoint, every modelAutomatically pick the most cost-effective model for each task, enforce budgets, and compare spend across providers from a single, unified billing layer.
Reduce AI spend fastDefine per-request failover chains so outages or rate limits seamlessly roll to backup models, keeping your production workloads stable and always-on.
No single point of failureGet end-to-end traces, latency and error metrics, and payload-level logs for every provider in one place—plus hooks for alerts and custom dashboards.
See every token, everywhereDeclare the job—chat, generation, tools, retrieval, structured outputs—and let LLM.API normalize APIs, schemas, and options across providers for you.
Think tasks, not vendorsSubmit massive workloads as batches with automatic parallelization, retries, and provider-optimized chunking to drive down cost and maximize throughput.
Process millions, reliablyDecision guide
FAQ
Gemini 3.1 Flash Lite Preview is a lightweight, preview-version Gemini model from Google optimized for fast, low-cost generation via the LLM.API gateway.
It is best for high-volume, latency-sensitive tasks like chatbots, simple agents, and lightweight content generation where cost efficiency matters more than peak quality.
Gemini 3.1 Flash Lite Preview supports up to 128K tokens of context via LLM.API, enabling long conversations and documents.
It is tuned for low latency, generally returning first tokens quickly and handling streaming responses efficiently for interactive applications.
Through LLM.API it supports text input and text output, with multimodal features depending on the specific LLM.API integration configuration.
Pricing is usage-based per input and output token, with rates set by LLM.API and typically lower than larger, higher-quality Gemini variants.
You select the model name "google/gemini-3.1-flash-lite-preview" in your LLM.API request and pass messages using the standard chat completions schema.
Flash Lite is generally cheaper and faster but slightly lower in quality and capability than the full Gemini 3.1 Flash model.
It may underperform larger models on complex reasoning, nuanced coding tasks, and highly specialized domains, and is provided as a preview with evolving behavior.
Yes, it can generate and edit code, but for complex or critical programming tasks a more capable Gemini or other advanced model is recommended.
Compare
Kimi K2.6 is MoonshotAI’s open-source, 1-trillion-parameter Mixture-of-Experts multimodal model optimized for long-horizon coding, agentic tool use, and image/video understanding. It is notable for its large ~262K-token…
Llama 3.3 Nemotron Super 49B V1.5 is a 49B-parameter NVIDIA language model optimized for English-centric reasoning and chat, with a long 128K+ context window and support…
GLM 5V Turbo is Z.ai’s native multimodal large language model optimized for vision-based coding and agentic workflows, able to process images, video, and text for complex…