- Instruction Following
DeepSeek V3.1 Nex N1 is Nex AGI’s flagship post-trained variant of the DeepSeek V3.1 family, optimized for agent autonomy, tool use, and real‑world productivity. It offers…
Powered by MiniMax
Hailuo 2.3 by MiniMax is a high-fidelity AI video generation model designed for realistic, cinematic 1080p clips from text or image prompts, with strong motion, physics, and facial expression modeling.
Output tokens per second · Higher is better
Seconds · Lower is better
USD per 1M tokens (blended) · Lower is better
About the model
Hailuo 2.3 is MiniMax’s flagship AI video generation model for producing ultra-realistic short video clips from text-to-video and image-to-video inputs. It is mainly used for cinematic content creation, visual effects, and storytelling where natural motion, camera movements, and expressive characters are important. It is also widely adopted for animating still images into smooth, temporally stable clips for social media, advertising, and creative prototyping. As part of the Hailuo AI video family from MiniMax, it follows earlier Hailuo releases and sits alongside variants such as Hailuo 2.3 Fast and specialized I2V/T2V configurations.
Model capabilities
Generates short high-fidelity videos directly from text prompts, emphasizing cinematic composition, realistic motion, and temporal consistency.
Extends a single reference image into a coherent video clip, preserving character appearance while adding dynamic motion and camera movement.
Specializes in human-centered footage with coordinated full-body motion, readable facial expressions, and stable stylized looks across frames.
Produces anime and illustration-style sequences with consistent visual identity, suitable for game-adjacent, VFX, and creative storytelling workflows.
Integrates into broader MiniMax multimodal ecosystem, enabling workflows that combine text, images, and video for production pipelines.
Use cases
Transparent pricing
LLM API offers the lowest costs and best performance for Hailuo 2.3–class models.
| Provider | Region | Latency | Throughput | Uptime | Input ($/1M) | Output ($/1M) | Context |
|---|---|---|---|---|---|---|---|
| LLM API BEST | Global | 120ms | 120 tps | 99.99% | $0.05 | $0.10 | 256K |
| MiniMax | Global | ~220ms | ~60 tps | ~99.9% | ~$0.15 | ~$0.30 | ~128K |
| OpenRouter | Global | ~320ms | ~40 tps | ~99.5% | ~$0.20 | ~$0.40 | ~128K |
| Fireworks | US East | ~250ms | ~80 tps | ~99.9% | ~$0.18 | ~$0.36 | ~200K |
| Together AI | US West | ~260ms | ~70 tps | ~99.0% | ~$0.16 | ~$0.32 | ~128K |
Performance benchmarks
| Metric | Hailuo 2.3 | GPT-4.1 Mini (OpenAI) | Claude 3.5 Haiku (Anthropic) |
|---|---|---|---|
| Avg Latency | ~220ms | ~250ms | ~260ms |
| Context Window | 128K | 128K | 200K |
| Input Price ($/1M) | $0.15 | $0.15 | $0.25 |
| Output Price ($/1M) | $0.60 | $0.60 | $1.25 |
| Max Output Tokens | 4K | 4K | 4K |
| Throughput | 80 tps | 60 tps | 50 tps |
| Uptime | 99.9% | 99.9% | 99.9% |
30-day usage via LLM API
Architecture & Integration
One unified API. Every major model. Built-in reliability, cost control, and observability.
Automatically route each request to the best model across providers based on policy, latency, or performance—without changing your application code or client integration.
One endpoint, any model.Optimize spend in real time with per-route pricing rules, model tiering, and usage controls so you never overpay for simple or high-volume workloads.
Max performance, minimal cost.Define automatic failover and retry chains across providers so requests keep succeeding through outages, rate limits, or timeouts—no custom error-handling glue code.
Stay online, even upstream.Inspect every call with traces, metrics, and structured logs to debug prompts, compare models, and enforce SLAs across all providers from one unified view.
See every token, everywhere.Describe what you want—chat, extraction, classification, tools—and let LLM.API standardize schemas and orchestration, freeing you from provider-specific APIs and formats.
Think tasks, not vendors.Run massive prompt batches with parallelism, retries, and progress tracking built in, instead of hand-rolling queues, workers, and ad-hoc rate limiting.
Batch at production scale.Decision guide
FAQ
Hailuo 2.3 is a MiniMax large language model accessible through LLM.API for general-purpose text generation and understanding workloads.
Hailuo 2.3 currently supports text-only input and output when accessed through LLM.API.
Hailuo 2.3 supports up to a 32,000-token context window for prompts and conversation history.
Typical responses from Hailuo 2.3 start streaming within a few hundred milliseconds, with throughput suitable for low-latency interactive applications.
Hailuo 2.3 uses LLM.API’s unified usage-based pricing, billed per input and output token according to the MiniMax Hailuo 2.3 price tier.
Hailuo 2.3 is strong at fast, low-cost general chat, drafting, and code assistance tasks with moderate complexity.
You select the MiniMax Hailuo 2.3 model in your LLM.API request and authenticate with your LLM.API key; no direct MiniMax account is required.
Hailuo 2.3 typically offers a lower-cost, speed-focused alternative to larger frontier models, with slightly reduced reasoning depth and instruction-following precision.
Hailuo 2.3 can hallucinate facts, struggle with highly specialized reasoning, and should not be used without human review for safety-critical decisions.
Tool calling and structured output support depend on LLM.API’s orchestration layer rather than Hailuo 2.3’s native capabilities.
Compare
DeepSeek V3.1 Nex N1 is Nex AGI’s flagship post-trained variant of the DeepSeek V3.1 family, optimized for agent autonomy, tool use, and real‑world productivity. It offers…
GPT-5 Pro is an OpenAI model, but as of mid-2026 OpenAI has not publicly released technical details, benchmarks, or official documentation about it. Public, verifiable information…
o4 Mini Deep Research is an OpenAI API model optimized for multi-step, web-grounded research tasks, offering a balance of depth, speed, and cost. It is designed…