- Text Generation
Devstral 2 2512 is a 123B-parameter open-source large language model from Mistral AI, optimized for agentic coding and long-context software engineering workflows. It supports a 256K/262K-token…
Powered by Black Forest Labs
FLUX.2 Klein 4B is a compact, 4‑billion‑parameter image generation and editing model from Black Forest Labs, optimized for fast, sub‑second inference on consumer GPUs. It delivers high‑quality visual outputs while unifying text‑to‑image and image‑editing capabilities in a single architecture.
Output tokens per second · Higher is better
Seconds · Lower is better
USD per 1M tokens (blended) · Lower is better
About the model
FLUX.2 Klein 4B is a 4B-parameter rectified-flow transformer model by Black Forest Labs for high-quality, low-latency image generation and editing on consumer hardware. It is mainly used for text-to-image creation in interactive applications where sub-second response and good visual fidelity are important. It is also widely used for single- and multi-reference image editing workflows, including LoRA-based personalization and fine-tuning-friendly setups. The model is part of the FLUX.2 [klein] family, a fast, compact branch of the broader FLUX.2 image-generation and editing models.
Model capabilities
Generates high-quality images from natural language prompts using a compact 4B-parameter rectified flow transformer architecture.
Edits existing images based on text instructions, enabling transformations, enhancements, and content modifications in a unified pipeline.
Combines multiple reference images with text prompts to guide style, composition, or subject while preserving visual consistency.
Optimized for sub-second image generation and editing on consumer GPUs, supporting interactive and high-volume visual workflows.
Supports fine-tuning and LoRA-based customization through its base variants, enabling domain-specific or style-specialized image models.
Use cases
Transparent pricing
LLM API offers the lowest per-image cost and best performance for FLUX.2-class 4B models.
| Provider | Region | Latency | Throughput | Uptime | Input ($/1M) | Output ($/1M) | Context |
|---|---|---|---|---|---|---|---|
| LLM API BEST | Global | ~350ms | ~120 img/min | 99.99% | $0.0006/img | $0.0000/img | 1 img up to 1024x1024 |
| Black Forest Labs (Direct) | EU West | ~550ms | ~60 img/min | ~99.9% | ~$0.0012/img | $0.0000/img | ~1 img up to 1024x1024 |
| Replicate | Global | ~700ms | ~40 img/min | ~99.5% | ~$0.0015/img | $0.0000/img | ~1 img up to 1024x1024 |
| Together AI | US East | ~600ms | ~70 img/min | ~99.9% | ~$0.0013/img | $0.0000/img | ~1 img up to 1024x1024 |
Performance benchmarks
| Metric | FLUX.2 Klein 4B | Stable Diffusion 3.5 Medium | DALL·E 3 (standard) |
|---|---|---|---|
| Latency per Image | ~900ms | ~1.1s | ~1.3s |
| Throughput | ~40 img/s | ~35 img/s | ~30 img/s |
| Max Resolution | 1536x1536 | 1536x1536 | 1792x1024 |
| Price per Image | $0.020 | $0.018 | $0.040 |
| Supported Formats | PNG, JPG | PNG, JPG | PNG, JPG |
| Uptime | 99.5% | 99.9% | 99.9% |
30-day usage via LLM API
Architecture & Integration
One unified API. Every major model. Built-in reliability, cost control, and observability.
Automatically route each request to the optimal model across providers based on latency, cost, or quality — without changing your integration.
One endpoint, every model.Control spend with smart tiering, quotas, and policy-based model selection so you always use the cheapest model that still meets requirements.
Optimize every token.Stay online when a model or provider fails with built-in health checks and seamless failover, no extra logic in your app.
Resilient by default.Trace every request across models and providers with logs, metrics, and event streams that plug into your existing monitoring stack.
See every token flow.Call high-level tasks like chat, tools, or rerank instead of model-specific APIs, so you can swap models without refactoring.
Code to tasks, not models.Process millions of inferences efficiently with batch endpoints that maximize provider throughput while handling retries and rate limits for you.
Scale inference, not ops.Decision guide
FAQ
FLUX.2 Klein 4B is a 4B-parameter image generation model from Black Forest Labs optimized for fast, efficient, high-quality image synthesis.
FLUX.2 Klein 4B supports text-to-image generation and image-to-image transformation through the LLM.API image generation endpoints.
FLUX.2 Klein 4B is best for rapid, low-cost image generation where lightweight deployment, iteration speed, and decent visual quality are priorities.
On LLM.API, FLUX.2 Klein 4B is billed per generated image or image step, with exact pricing defined in the LLM.API model catalog.
Call the LLM.API image generation endpoint with the FLUX.2 Klein 4B model identifier and your API key in the Authorization header.
Typical text-to-image requests return in a few seconds, depending on resolution, step count, and current LLM.API load.
FLUX.2 Klein 4B does not use a token-based context window; it consumes prompts as text strings and conditioning inputs for image generation.
FLUX.2 Klein 4B generally trades some visual fidelity and detail for significantly lower compute cost, faster responses, and easier deployment.
Yes, FLUX.2 Klein 4B usage is subject to LLM.API safety filters and content policies, which may block disallowed or unsafe generations.
FLUX.2 Klein 4B may struggle with very fine text rendering, complex multi-object scenes, and ultra-photorealism compared to larger image models.
Compare
Devstral 2 2512 is a 123B-parameter open-source large language model from Mistral AI, optimized for agentic coding and long-context software engineering workflows. It supports a 256K/262K-token…
GPT-4o Transcribe is an OpenAI model specialized for converting audio into accurate, time-aligned text transcripts. It is notable for handling natural speech, varied accents, and real-world…
Embed V1 0.6B is Perplexity’s 0.6‑billion‑parameter text embedding model designed for fast, low‑latency, web‑scale retrieval. It produces compact INT8 or binary embeddings optimized for dense semantic…