- Text Generation
Riverflow V2 Standard Preview is the standard variant of Sourceful's Riverflow 2.0 preview lineup, offering unified text-to-image and image-to-image generation focused on production-grade creative workflows.
Powered by MiniMax
MiniMax M2 is an open‑weight Mixture‑of‑Experts large language model from MiniMax, designed to deliver high coding and agentic workflow performance with low latency and cost. It uses 230B total parameters with only about 10B active per token to balance strong reasoning with efficient deployment.
Output tokens per second · Higher is better
Seconds · Lower is better
USD per 1M tokens (blended) · Lower is better
About the model
MiniMax M2 is an open‑weight, MoE-based large language model by MiniMax optimized for coding and autonomous agent workflows. It is mainly used for software development tasks such as code generation, refactoring, and debugging, as well as orchestrating multi-step agentic workflows that call tools and APIs efficiently. It also serves as a general-purpose LLM for chat, reasoning, and integration into developer tools and AI platforms. MiniMax M2 belongs to the MiniMax-M2 family of Mixture-of-Experts models and follows earlier MiniMax research lines such as the MiniMax-M1 reasoning models.
Model capabilities
Optimized for code generation, debugging, multi-file editing, and compile-run-fix loops in modern software engineering workflows.
Designed for tool use and agentic reasoning, enabling plan-act-verify loops and complex multi-step task automation.
Handles very long inputs with strong reasoning performance across benchmarks, suitable for large documents and complex problems.
Provides strong multilingual language understanding and generation, covering multiple major languages with high-quality outputs.
Exhibits outstanding optical character recognition on handwritten text, outperforming many contemporary AI models in accuracy tests.
Use cases
Transparent pricing
LLM API offers the lowest MiniMax‑class pricing and fastest response times versus other providers.
| Provider | Region | Latency | Throughput | Uptime | Input ($/1M) | Output ($/1M) | Context |
|---|---|---|---|---|---|---|---|
| LLM API BEST | Global | 120ms | 80 tps | 99.99% | $0.30 | $0.60 | 128K |
| MiniMax | Asia Pacific | ~220ms | ~40 tps | 99.9% | ~$0.40 | ~$0.80 | ~64K |
| OpenAI (closest: GPT‑4o‑mini) | Global | ~180ms | ~50 tps | 99.9% | ~$0.50 | ~$1.00 | 128K |
| Anthropic (closest: Claude 3 Haiku) | US East | ~190ms | ~45 tps | 99.9% | ~$0.55 | ~$1.10 | 200K |
| Google (closest: Gemini 1.5 Flash) | Global | ~200ms | ~45 tps | 99.9% | ~$0.45 | ~$0.90 | 1M |
Performance benchmarks
| Metric | MiniMax M2 | OpenAI GPT-4.1 Mini | Anthropic Claude 3 Haiku |
|---|---|---|---|
| Avg Latency | ~220ms | ~180ms | ~200ms |
| Context Window | ~128K | 128K | 200K |
| Input Price ($/1M) | ~$0.15 | $0.15 | $0.25 |
| Output Price ($/1M) | ~$0.60 | $0.60 | $1.25 |
| Max Output Tokens | ~4K | 4K | 4K |
| Throughput | ~45 tps | ~50 tps | ~40 tps |
| Uptime | ~99.9% | ~99.9% | ~99.9% |
30-day usage via LLM API
Architecture & Integration
One unified API. Every major model. Built-in reliability, cost control, and observability.
Dynamically route each request to the optimal model across providers based on latency, cost, and quality—no client changes or custom glue code required.
One endpoint, every modelControl spend with price-aware routing, model selection, and usage policies so you can ship AI features fast without surprise bills or manual tuning.
Lower cost, same qualityDefine automatic fallbacks to alternative models and providers when requests fail or degrade, keeping your AI features reliable even during provider outages.
Stay online, automaticallyGet unified logs, traces, and metrics across every provider, model, and endpoint so you can debug, optimize prompts, and monitor performance from a single place.
See every tokenCall high-level tasks like chat, generate, extract, or embed instead of vendor-specific APIs, giving you portable, maintainable code that outlives any single model.
Code to tasks, not vendorsProcess thousands of calls efficiently with smart batching and concurrency controls, maximizing throughput while staying within provider limits and budget.
Scale to millions of callsDecision guide
FAQ
MiniMax M2 is a large language model by MiniMax focused on efficient, general-purpose text generation and understanding for applications like chatbots and content tools.
MiniMax M2 supports a context window of up to 32K tokens via LLM.API, suitable for longer conversations and multi-document prompts.
MiniMax M2 currently supports text input and text output only when accessed via LLM.API.
MiniMax M2 typically returns first tokens within a few hundred milliseconds to a couple of seconds, depending on prompt length and load.
MiniMax M2 usage on LLM.API is billed per 1,000 input and output tokens, with exact rates shown in your LLM.API pricing dashboard.
You select the MiniMax M2 model ID in your LLM.API request and send standard Chat or Completion-style JSON with messages and parameters.
MiniMax M2 is best for cost-efficient conversational agents, drafting and editing text, and general reasoning where ultra-high-end reasoning is not mandatory.
MiniMax M2 generally offers a good balance of quality and cost, competing with mid-tier models while being cheaper than many frontier models.
MiniMax M2 can hallucinate facts, lacks real-time knowledge, and may underperform top-tier frontier models on complex reasoning or highly specialized domains.
MiniMax M2 is currently available only as a hosted base model on LLM.API, without user-managed fine-tuning.
Compare
Riverflow V2 Standard Preview is the standard variant of Sourceful's Riverflow 2.0 preview lineup, offering unified text-to-image and image-to-image generation focused on production-grade creative workflows.
Recraft V4.1 Pro is a high-aesthetics image generation model from Recraft that produces ~2K-resolution images with enhanced photorealism, smooth gradients, and strong adherence to short prompts,…
Trinity Large Thinking is Arcee AI’s open-weight, 398–400B-parameter sparse Mixture-of-Experts model focused on advanced reasoning and long-horizon agentic tasks. It is notable for activating only about…