- Image Generation
Recraft V4 is Recraft’s third-generation image generation model, built around professional design workflows and visual taste, with a focus on photorealism, refined composition, and high-quality raster…
Powered by OpenAI
GPT-5 Image is an OpenAI multimodal model variant focused on understanding and generating images, extending GPT-5’s capabilities beyond text to visual content.
Output tokens per second · Higher is better
Seconds · Lower is better
USD per 1M tokens (blended) · Lower is better
About the model
GPT-5 Image is an OpenAI model designed to interpret and create images as part of a broader multimodal AI system. It is mainly used for tasks such as image understanding, description, and visual question answering. It can also support workflows that combine text and imagery, like design ideation, document analysis, and educational content creation. It follows earlier OpenAI vision-capable models such as GPT-4 with vision and other GPT-5 family variants.
Model capabilities
Analyzes uploaded images, identifying objects, people, scenes, and visual relationships, and can answer detailed questions about visual content.
Reads and extracts printed or handwritten text from images, screenshots, documents, and photos, preserving structure where possible.
Engages in interactive conversations mixing text and images, allowing users to reference visuals directly within natural language dialogue.
Assists in monitoring visual and textual content for policy-violating material, supporting safer, compliant applications and user experiences.
Translates text found within images or screenshots into other languages while retaining awareness of surrounding visual context and layout.
Use cases
Transparent pricing
LLM API offers the lowest prices and best limits for GPT-5 Image–class models.
| Provider | Region | Latency | Throughput | Uptime | Input ($/1M) | Output ($/1M) | Context |
|---|---|---|---|---|---|---|---|
| LLM API BEST | Global | ~180ms | ~120 img/min | 99.99% | ~$0.0004/img | ~$0.0004/img | ~64K tokens + 8 images |
| OpenAI | Global | ~220ms | ~80 img/min | 99.9% | ~$0.0010/img | ~$0.0010/img | ~32K tokens + 4 images |
| Azure OpenAI | US East | ~240ms | ~70 img/min | 99.9% | ~$0.0011/img | ~$0.0011/img | ~32K tokens + 4 images |
| Google Cloud (Gemini Vision-equivalent) | Global | ~260ms | ~60 img/min | 99.9% | ~$0.0012/img | ~$0.0012/img | ~32K tokens + 4 images |
| Anthropic (Claude Vision-equivalent) | US West | ~250ms | ~55 img/min | 99.9% | ~$0.0013/img | ~$0.0013/img | ~32K tokens + 4 images |
Performance benchmarks
| Metric | GPT-5 Image (OpenAI) | Gemini Vision Pro (Google) | Claude 3.7 Vision (Anthropic) |
|---|---|---|---|
| Latency per Image | ~800ms | ~900ms | ~1.0s |
| Throughput | ~40 img/s | ~35 img/s | ~30 img/s |
| Max Resolution | ~4096x4096 | ~4096x4096 | ~4096x4096 |
| Price per Image | ~$0.005 | ~$0.006 | ~$0.007 |
| Supported Formats | PNG, JPG, WEBP, GIF~ | PNG, JPG, WEBP~ | PNG, JPG, WEBP~ |
| Uptime | ~99.9% | ~99.9% | ~99.9% |
30-day usage via LLM API
Architecture & Integration
One unified API. Every major model. Built-in reliability, cost control, and observability.
Dynamically route each request to the optimal model across providers based on latency, cost, and quality—without changing your integration or redeploying code.
One endpoint, every model.Enforce budgets, choose cheaper equivalents, and mix premium and commodity models automatically so you control spend without manually tuning every call.
Max performance, minimal spend.Survive provider outages and rate limits with automatic cross-vendor failover and graceful degradation policies, keeping your AI features online by default.
Stay up when others fail.Trace every request across models and providers with unified logs, metrics, and latency breakdowns so you can debug, optimize, and prove reliability in production.
See every token, everywhere.Define high-level tasks—chat, extraction, scoring, tools—and let LLM.API pick and configure the right models so you ship features, not prompt glue.
Code tasks, not prompts.Run massive offline jobs across providers with parallelized batching, retry policies, and progress tracking so you can process millions of items reliably and cheaply.
Scale from 10 to millions.Decision guide
FAQ
GPT-5 Image is an OpenAI multimodal model for image understanding and generation, accessible via the unified LLM.API gateway.
GPT-5 Image is best for tasks combining images and text, such as visual question answering, image description, UI understanding, and image editing instructions.
GPT-5 Image pricing is defined by LLM.API’s OpenAI GPT-5 Image tariff; check your LLM.API dashboard or pricing docs for up-to-date per-token and image rates.
GPT-5 Image supports a large-context text window determined by the underlying OpenAI deployment; see the LLM.API model card for the current token limit.
GPT-5 Image latency depends on input size and load, but LLM.API maintains persistent connections and routing to minimize typical response times.
GPT-5 Image supports text-plus-image inputs and text outputs, with image-related reasoning and editing instructions handled in a single multimodal endpoint.
Use the LLM.API completion or chat endpoint with the provider set to OpenAI and the model set to gpt-5-image, passing image URLs or bytes as inputs.
Compared to earlier OpenAI vision models, GPT-5 Image generally offers better reasoning, more accurate descriptions, and stronger instruction-following on complex multimodal tasks.
GPT-5 Image can misinterpret ambiguous images, hallucinate details, and should not be solely relied on for safety-critical or privacy-sensitive visual analysis.
GPT-5 Image supports batching and streaming where enabled by LLM.API, allowing higher throughput and incremental token delivery for long multimodal responses.
Compare
Recraft V4 is Recraft’s third-generation image generation model, built around professional design workflows and visual taste, with a focus on photorealism, refined composition, and high-quality raster…
Recraft V4.1 Utility is a controlled-output image generation model from Recraft, optimized for clean, predictable visuals with flat lighting and front-facing compositions. It is particularly suited…
GPT-5 Image Mini is an OpenAI model for lightweight image understanding and generation, optimized for speed and efficiency over maximum fidelity. It is designed for everyday…