- Text Generation
Devstral 2 2512 is a 123B-parameter open-source large language model from Mistral AI, optimized for agentic coding and long-context software engineering workflows. It supports a 256K/262K-token…
Powered by Google
Nano Banana 2 (Gemini 3.1 Flash Image Preview) is Google DeepMind’s image generation and editing model built on the Gemini 3.1 Flash architecture, optimized for fast, cost‑efficient, high‑quality visuals. It balances strong multimodal understanding with 4K-capable output and low latency for both text-to-image and image-edit tasks.
Output tokens per second · Higher is better
Seconds · Lower is better
USD per 1M tokens (blended) · Lower is better
About the model
Nano Banana 2 (Gemini 3.1 Flash Image Preview) is a Google DeepMind model for high-speed, high-quality image understanding, generation, and editing built on the Gemini 3.1 Flash family. It is mainly used for text-to-image generation, enabling rapid creation of detailed images with controllable aspect ratios and resolutions up to 4K. It is also used for image editing workflows, where users supply reference or input images for transformations, variations, and iterative refinements within apps and APIs. As part of the Gemini 3.1 model line, it succeeds earlier Gemini image capabilities and sits alongside the higher-end Nano Banana Pro image models.
Model capabilities
Supports interactive, multi-turn dialogue, following instructions and maintaining context for tasks like Q&A, drafting, and brainstorming.
Interprets images to identify objects, scenes, and visual details, enabling visual question answering and description tasks.
Reads and extracts text from images, including screenshots and documents, enabling search, analysis, and transformation of visual text content.
Helps write and reason about code, and orchestrate external tools or APIs by interpreting structured instructions and outputs.
Translates between multiple natural languages while preserving meaning and tone, supporting cross-lingual understanding and communication tasks.
Use cases
Transparent pricing
LLM API offers the lowest effective cost and latency for Nano Banana 2–class vision models.
| Provider | Region | Latency | Throughput | Uptime | Input ($/1M) | Output ($/1M) | Context |
|---|---|---|---|---|---|---|---|
| LLM API BEST | Global | 80ms | 120 img/min | 99.99% | $0.30/1K img | $0.00 | 16 images / 128K tokens equivalent |
| Global | ~200ms | ~60 img/min | 99.9% | ~$0.60/1K img | $0.00 | ~16 images / 128K tokens equivalent | |
| Vertex AI | US East | ~220ms | ~48 img/min | 99.9% | ~$0.65/1K img | $0.00 | ~16 images / 128K tokens equivalent |
| AWS Bedrock (equivalent vision model) | US East | ~240ms | ~45 img/min | 99.9% | ~$0.80/1K img | $0.00 | ~8 images / 128K tokens equivalent |
| Azure OpenAI (equivalent vision model) | Global | ~230ms | ~50 img/min | 99.9% | ~$0.75/1K img | $0.00 | ~8 images / 128K tokens equivalent |
Performance benchmarks
| Metric | Nano Banana 2 (Gemini 3.1 Flash Image Preview) | GPT-4o mini (Image Preview) | Claude 3.5 Haiku (Vision) |
|---|---|---|---|
| Latency per Image | ~250ms | ~220ms | ~260ms |
| Context Window | 128K | 128K | 200K |
| Max Resolution | 2K | 2K | 2K |
| Price per Image | $0.002 | $0.002 | $0.003 |
| Supported Formats | PNG, JPG, WEBP | PNG, JPG, WEBP | PNG, JPG, WEBP |
| Throughput | 80 img/s | 90 img/s | 70 img/s |
| Uptime | 99.9% | 99.9% | 99.9% |
30-day usage via LLM API
Architecture & Integration
One unified API. Every major model. Built-in reliability, cost control, and observability.
Automatically route each request to the optimal model and provider based on latency, cost, and quality — without changing your integration or redeploying.
One endpoint, every modelDynamically choose cheaper equivalents, downgrade for non-critical paths, and enforce budgets with per-route policies so you never get surprised by your AI bill.
Lower spend, same outputDefine automatic cross-provider fallbacks and retries so traffic fails over seamlessly during outages, rate limits, or model errors — no manual playbooks required.
Zero-downtime AI callsGet unified traces, logs, metrics, and payload sampling across all providers to debug latency, failures, and regressions from a single, model-agnostic view.
See every token, everywhereDescribe tasks like chat, embedding, or tool-calling once and let LLM.API handle provider-specific quirks, parameters, and response formats for you.
Code to tasks, not vendorsBatch thousands of requests per task across providers with built-in retries, rate-limit smoothing, and streaming results to maximize throughput and minimize cost.
Scale to millions of callsDecision guide
FAQ
Nano Banana 2 (Gemini 3.1 Flash Image Preview) is a Google multimodal model optimized for fast, low-cost text and image understanding via LLM.API.
It is best for latency-sensitive applications needing quick image interpretation, lightweight vision-language reasoning, and inexpensive high-volume text processing.
Nano Banana 2 (Gemini 3.1 Flash Image Preview) supports a 32K token context window through LLM.API.
It is tuned for very low end-to-end latency, making it suitable for real-time or interactive user experiences.
The model accepts text and image inputs and returns text outputs, enabling vision-language workflows.
Pricing is usage-based per token and image processed, with exact rates available in the LLM.API Google models pricing table.
Use the LLM.API chat or completions endpoint specifying the Nano Banana 2 model name, sending text and optional image inputs in the request body.
Compared to larger Gemini models, it trades some reasoning depth and creativity for significantly lower latency and cost.
It may struggle with highly complex reasoning, long multi-step problem solving, and domain-expert tasks compared to larger frontier models.
Yes, you can use LLM.API tool-calling interfaces, but the model is mainly optimized for lightweight text and image understanding.
Compare
Devstral 2 2512 is a 123B-parameter open-source large language model from Mistral AI, optimized for agentic coding and long-context software engineering workflows. It supports a 256K/262K-token…
Claude Sonnet 4.5 is an Anthropic large language model optimized for software development, computer use, and agentic workflows, offering strong performance on coding and reasoning tasks…
Anthropic Claude Haiku (Latest) is a lightweight, fast Claude family model optimized for low-latency, cost‑efficient tasks while maintaining strong language understanding. It is notable for offering…