- Instruction Following
Gemini 3.5 Flash is Google’s natively multimodal reasoning model optimized for very low latency and cost while maintaining frontier‑level performance, particularly for coding and agentic workflows.
Powered by Google
Gemini 2.5 Flash Lite Preview 09-2025 is a lightweight preview variant of Google’s Gemini 2.5 Flash-Lite model, optimized for fast, cost-efficient multimodal inference with long-context support. It offers text outputs from text, image, video, audio, and PDF inputs while showcasing improvements over the earlier Flash-Lite release.
Output tokens per second · Higher is better
Seconds · Lower is better
USD per 1M tokens (blended) · Lower is better
About the model
Gemini 2.5 Flash Lite Preview 09-2025 is a Google Gemini API model variant that provides a preview of updated Flash-Lite capabilities as of September 2025. It is mainly used for low-latency, high-throughput applications such as chatbots, agents, and tools that need long-context reasoning over large text or multimodal documents. It also targets developer workloads like batch processing, retrieval-augmented generation, and structured outputs using function calling and file search. It belongs to the Gemini 2.5 Flash-Lite family and is offered alongside the stable gemini-2.5-flash-lite model as a preview version.
Model capabilities
Accepts very long-context inputs across text, code, images, audio, and video while generating coherent text-only responses efficiently.
Handles interactive dialogue, following instructions and maintaining context over extended conversations with low latency and low cost.
Enhances answers using grounding with Google Search, improving factuality and up-to-date knowledge in supported use cases.
Supports many input and output languages, enabling multilingual applications for users across diverse regions and locales.
Analyzes images within multimodal prompts to extract visual details, interpret content, and incorporate findings into generated text.
Use cases
Transparent pricing
Up to ~60% cheaper and faster than comparable Gemini-class LLMs
| Provider | Region | Latency | Throughput | Uptime | Input ($/1M) | Output ($/1M) | Context |
|---|---|---|---|---|---|---|---|
| LLM API BEST | Global | 120ms | 120 tps | 99.99% | $0.05 | $0.10 | 256K |
| Global | ~220ms | ~60 tps | ~99.9% | ~$0.12 | ~$0.24 | ~256K | |
| OpenRouter | Global | ~260ms | ~45 tps | ~99.9% | ~$0.14 | ~$0.28 | ~128K |
| Together AI | US East | ~250ms | ~50 tps | ~99.9% | ~$0.13 | ~$0.26 | ~128K |
| Fireworks AI | US West | ~240ms | ~55 tps | ~99.9% | ~$0.13 | ~$0.25 | ~128K |
Performance benchmarks
| Metric | Gemini 2.5 Flash Lite Preview 09-2025 | GPT-4.1 Mini (OpenAI) | Claude 3.7 Haiku (Anthropic) |
|---|---|---|---|
| Avg Latency | ~180ms | ~220ms | ~250ms |
| Context Window | 128K | 128K | 200K |
| Input Price ($/1M) | $0.03 | $0.15 | $0.25 |
| Output Price ($/1M) | $0.06 | $0.60 | $0.75 |
| Max Output Tokens | 8K | 4K | 8K |
| Throughput | ~500 tps | ~200 tps | ~180 tps |
| Uptime | 99.9% | 99.9% | 99.9% |
30-day usage via LLM API
Architecture & Integration
One unified API. Every major model. Built-in reliability, cost control, and observability.
Automatically route each request to the optimal model across providers based on latency, cost, and capability—without changing your integration or redeploying code.
One endpoint, every model.Set cost ceilings and policies once, then let LLM.API select the cheapest model that still meets your quality and latency requirements in real time.
Optimize spend by default.Define multi-provider fallback chains so when a model or region fails, traffic seamlessly fails over—no downtime, no emergency redeploys.
Stay online, automatically.Get per-request traces, latencies, errors, and token usage across all providers in one place, with structured logs ready for your existing monitoring stack.
See every token, everywhere.Express work as high-level tasks—chat, extraction, tools, agents—while LLM.API handles prompts, models, and retries so your code stays clean and consistent.
Code to tasks, not models.Submit large batches in a single call and let LLM.API handle parallelization, rate limits, retries, and aggregation for massive throughput and lower unit costs.
Scale runs, not complexity.Decision guide
FAQ
Gemini 2.5 Flash Lite Preview 09-2025 is a Google Gemini model variant optimized for low-latency, cost-efficient multimodal generation in public preview.
Gemini 2.5 Flash Lite Preview 09-2025 supports up to 1,048,576 input tokens and 65,535 output tokens, giving it roughly a 1M token context window.
Gemini 2.5 Flash Lite Preview 09-2025 accepts text, code, images, audio, and video as input and generates text-only outputs.
It is best for high-throughput, latency-sensitive applications like chatbots, agents and lightweight multimodal understanding where low cost and speed matter more than peak quality.
Flash Lite Preview is tuned for lower latency and higher throughput than Gemini 2.5 Pro, at slightly lower raw reasoning and generation quality.
On Google Cloud it uses pay-as-you-go token-based billing with discounted input tokens when context caching is used; LLM.API applies its own unified pricing.
Call the LLM.API chat or completion endpoint with the provider set to Google and the model set to "gemini-2.5-flash-lite-preview-09-2025".
The preview model shares the same core architecture but has an earlier lifecycle, fewer supported features, and a scheduled discontinuation date of July 9, 2026.
It does not support Gemini Live API, supervised fine-tuning, or chat-completions endpoints and is constrained by a January 2025 knowledge cutoff.
No, this preview model is not exposed through Gemini Live API, so it cannot be used for real-time streaming audio conversations.
Compare
Gemini 3.5 Flash is Google’s natively multimodal reasoning model optimized for very low latency and cost while maintaining frontier‑level performance, particularly for coding and agentic workflows.
Qwen3 Max is Qwen’s flagship trillion-parameter large language model, offered as a high-end proprietary API model. It is designed to deliver state-of-the-art performance across reasoning, coding,…
Qwen3 VL 235B A22B Instruct is a 235B-parameter Mixture-of-Experts vision-language model from Qwen, offering open-weight, long-context (≈256K) multimodal reasoning over text, images, and video. It is…