- Image Generation
Recraft V4.1 Utility is a controlled-output image generation model from Recraft, optimized for clean, predictable visuals with flat lighting and front-facing compositions. It is particularly suited…
Powered by OpenAI
GPT-5 Image Mini is an OpenAI model for lightweight image understanding and generation, optimized for speed and efficiency over maximum fidelity. It is designed for everyday visual tasks where quick responses and lower compute costs are important.
Output tokens per second · Higher is better
Seconds · Lower is better
USD per 1M tokens (blended) · Lower is better
About the model
GPT-5 Image Mini is a compact OpenAI vision model focused on fast, cost‑efficient image analysis and generation. It is mainly used for tasks like quick image captioning, simple visual question answering, and basic image-based UI or assistant features. It also supports lightweight creative image generation for mockups, drafts, and low-resolution concepts where turnaround time matters more than photorealism. It follows earlier OpenAI multimodal models in the GPT and image model families, offering a smaller, more efficient option for visual workloads.
Model capabilities
Specialized small-footprint vision model from OpenAI’s GPT-5 family, optimized for fast image-related tasks and integrations.
Extracts readable text from images when present, enabling downstream processing like search, classification, or simple understanding tasks.
Follows concise instructions about images, such as answering simple questions or identifying requested visual elements within them.
Designed for efficient, low-latency use in applications that need quick image understanding without the overhead of larger multimodal models.
Can provide basic labels or short descriptions for visual content that may support multiple languages, depending on tooling configuration.
Use cases
Transparent pricing
LLM API offers the lowest image costs and latency for GPT-5 Image Mini–class models.
| Provider | Region | Latency | Throughput | Uptime | Input ($/1M) | Output ($/1M) | Context |
|---|---|---|---|---|---|---|---|
| LLM API BEST | Global | ~160ms | ~120 img/min | 99.99% | ~$0.0004/img | ~$0.0004/img | ~64K tokens + 8 images |
| OpenAI | Global | ~220ms | ~80 img/min | 99.9% | ~$0.0008/img | ~$0.0008/img | ~32K tokens + 4 images |
| Azure OpenAI | US East | ~250ms | ~70 img/min | 99.9% | ~$0.0009/img | ~$0.0009/img | ~32K tokens + 4 images |
| Amazon Bedrock | US West | ~260ms | ~65 img/min | 99.9% | ~$0.0010/img | ~$0.0010/img | ~32K tokens + 4 images |
| Anthropic | Global | ~240ms | ~75 img/min | 99.9% | ~$0.0011/img | ~$0.0011/img | ~64K tokens + 6 images |
Performance benchmarks
| Metric | GPT-5 Image Mini (OpenAI) | Gemini Flash Vision (Google) | Claude 3.7 Haiku Vision (Anthropic) |
|---|---|---|---|
| Latency per Image | ~180ms | ~220ms | ~250ms |
| Throughput | ~40 img/s | ~30 img/s | ~25 img/s |
| Max Resolution | 4K | 4K | 4K |
| Price per Image | ~$0.0006 | ~$0.0007 | ~$0.0008 |
| Supported Formats | JPG, PNG, WEBP, HEIC | JPG, PNG, WEBP | JPG, PNG, WEBP |
| Uptime | 99.9% | 99.5% | 99.5% |
30-day usage via LLM API
Architecture & Integration
One unified API. Every major model. Built-in reliability, cost control, and observability.
Automatically route each request to the best model across providers based on cost, latency, or quality—no client changes, just smarter traffic decisions.
One endpoint, many LLMsControl spend with dynamic model selection, rate limits, and hard budgets while keeping performance high. Ship fast without losing track of every token.
Cut costs, not coverageDesign multi-provider failover in a few lines: auto-retry on errors, degrade gracefully, and keep production apps online even when vendors break.
Failure-safe by defaultGet full traces, metrics, and logs for every call across all providers. Debug latency, drift, and failures from a single, provider-agnostic dashboard.
See every token hopExpress high-level tasks—chat, tools, RAG, agents—once and let LLM.API pick the right models, parameters, and workflows for each use case.
Tasks, not glue codeRun massive batch generations, evaluations, or embeddings with built-in concurrency controls, retries, and progress tracking—without building custom job infrastructure.
Batch at platform scaleDecision guide
FAQ
GPT-5 Image Mini is an OpenAI model optimized for fast, low-cost image understanding and lightweight vision-language tasks via the LLM.API gateway.
GPT-5 Image Mini is best for quick image captioning, classification, basic visual question answering, and integrating lightweight vision features into applications.
GPT-5 Image Mini usage is billed per input tokens and image units according to LLM.API’s OpenAI pricing tier for this model.
GPT-5 Image Mini supports a context window sized for short to medium prompts, suitable for concise instructions and descriptions alongside images.
GPT-5 Image Mini is optimized for low latency, returning responses quickly enough for interactive applications and real-time user interfaces.
GPT-5 Image Mini accepts image and text inputs and returns text outputs describing, analyzing, or reasoning about the provided images.
Use the LLM.API completion or chat endpoint with the provider set to OpenAI and the model name set to gpt-5-image-mini.
GPT-5 Image Mini is cheaper and faster but less capable on complex reasoning, detailed analysis, and high-stakes vision tasks than larger GPT-5 variants.
No, GPT-5 Image Mini focuses on understanding and describing existing images rather than generating new images from scratch.
Yes, GPT-5 Image Mini can stream text tokens via LLM.API when you enable streaming in the request parameters.
Compare
Recraft V4.1 Utility is a controlled-output image generation model from Recraft, optimized for clean, predictable visuals with flat lighting and front-facing compositions. It is particularly suited…
FLUX.2 Max is Black Forest Labs’ flagship text-to-image and image-editing model, designed for production-grade creative workflows with state-of-the-art photorealism, prompt following, and editing consistency.
Recraft V4.1 Utility Pro is a higher-resolution, general-purpose image generation model from Recraft, designed for controlled, predictable commercial visuals with simple, front-facing compositions. It targets production…