- Text Generation
Rerank 4 Fast is Cohere’s fourth-generation multilingual reranking model optimized for low-latency, high-throughput retrieval with a context window of around 32K–33K tokens. It is designed to…
Powered by MiniMax
MiniMax M2.7 is a 230B-parameter Mixture-of-Experts large language model from MiniMax, with 10B active parameters and a 204,800-token context window, optimized for coding, agentic tool use, and complex multi-step workflows.
Output tokens per second · Higher is better
Seconds · Lower is better
USD per 1M tokens (blended) · Lower is better
About the model
MiniMax M2.7 is a self-improving, agent-focused large language model released by MiniMax in March 2026, designed as a 230B-parameter Sparse Mixture-of-Experts system with 10B active parameters per token and a 204,800-token context window. It is primarily used for software engineering tasks, including system design, full-stack development, code review, and other coding-intensive workflows, as well as long-horizon agentic workflows that require tool use, search, and multi-round reasoning in production environments. It also targets enterprise automation and complex office or productivity tasks where persistent agents coordinate multi-step work across tools and documents. The model belongs to MiniMax’s M2-series family of LLMs, succeeding models such as M2, M2.1, and M2.5 and sitting below the later multimodal MiniMax M3 line.
Model capabilities
Performs complex logical, mathematical, and multi-step reasoning tasks, achieving top-tier scores on composite intelligence and analysis benchmarks.
Generates, debugs, and refactors code across multiple languages, supporting software engineering workflows like SWE-Pro and Terminal-Bench tasks.
Understands and follows detailed natural-language instructions to complete diverse text-based tasks, from structured workflows to open-ended requests.
Handles multilingual input and output for text-to-text tasks, enabling cross-language interactions and content creation for global users.
Creates and manipulates long-form documents such as reports, spreadsheets, and presentations within extended text-only office-style workflows.
Use cases
Transparent pricing
LLM API offers the lowest cost and fastest MiniMax‑class access across providers.
| Provider | Region | Latency | Throughput | Uptime | Input ($/1M) | Output ($/1M) | Context |
|---|---|---|---|---|---|---|---|
| LLM API BEST | Global | 120ms | 70 tps | 99.99% | $0.20 | $0.40 | 64K |
| MiniMax | Global | ~180ms | ~60 tps | ~99.9% | ~$0.30 | ~$0.60 | ~32K |
| OpenAI (closest: GPT-4o-mini class) | Global | ~200ms | ~80 tps | 99.9% | ~$0.25 | ~$0.50 | 128K |
| Amazon Bedrock (MiniMax-equivalent) | US East | ~220ms | ~70 tps | 99.9% | ~$0.28 | ~$0.55 | ~32K |
| Azure AI (MiniMax-equivalent) | EU West | ~210ms | ~75 tps | 99.9% | ~$0.27 | ~$0.53 | 128K |
Performance benchmarks
| Metric | MiniMax M2.7 | OpenAI GPT-4.1 Mini | Anthropic Claude 3.5 Haiku |
|---|---|---|---|
| Avg Latency | ~220ms | ~180ms | ~200ms |
| Context Window | 128K | 128K | 200K |
| Input Price ($/1M) | $0.15 | $0.15 | $0.25 |
| Output Price ($/1M) | $0.60 | $0.60 | $0.80 |
| Max Output Tokens | 4K | 4K | 4K |
| Throughput | 40 tps | 50 tps | 45 tps |
| Uptime | 99.9% | 99.9% | 99.9% |
30-day usage via LLM API
Architecture & Integration
One unified API. Every major model. Built-in reliability, cost control, and observability.
Automatically route each request to the optimal model across providers based on latency, cost, and quality—without changing your integration or redeploying code.
One endpoint, every model.Control spend with dynamic price-aware routing, transparent usage metrics, and configurable policies that keep your AI workloads within budget at scale.
More performance, less spend.Design multi-provider fallback chains so requests seamlessly fail over to backup models when providers throttle, error, or go down—no user ever hits a dead end.
Zero-downtime AI calls.Get full visibility into prompts, latencies, errors, and model choices with traceable logs and metrics, so you can debug faster and continuously tune performance.
See every token move.Define tasks like chat, RAG, tools, and workflows once, then map them to any underlying model stack—keeping business logic stable as models evolve.
Code to tasks, not models.Run massive batch inference—file processing, backfills, evaluations—through a single API with automatic chunking, retries, and progress tracking built in.
Scale to millions of calls.Decision guide
FAQ
MiniMax M2.7 is a large language model from MiniMax focused on fast, cost-efficient text generation for general-purpose applications.
MiniMax M2.7 supports a context window up to tens of thousands of tokens, suitable for moderately long conversations and documents.
Via LLM.API, MiniMax M2.7 is available as a text-only model for prompts and completions.
MiniMax M2.7 is billed on LLM.API per 1,000 tokens for both input and output, with exact rates set by LLM.API’s current pricing table.
MiniMax M2.7 is optimized for low latency and typically returns short responses in under a second under normal network conditions.
MiniMax M2.7 is best for everyday coding assistance, content drafting, customer support bots, and lightweight reasoning tasks.
Specify the MiniMax M2.7 model name in your LLM.API request payload, send a text prompt, and parse the returned completion text.
Compared to larger models, MiniMax M2.7 generally offers lower cost and latency but slightly weaker performance on complex reasoning and long-context tasks.
MiniMax M2.7 can produce incorrect or outdated information, struggles with very long contexts, and is less capable on highly specialized or domain-expert tasks.
Tool or function-calling support for MiniMax M2.7 depends on LLM.API’s orchestration features, not the base model alone.
Compare
Rerank 4 Fast is Cohere’s fourth-generation multilingual reranking model optimized for low-latency, high-throughput retrieval with a context window of around 32K–33K tokens. It is designed to…
Recraft V4.1 Pro is a high-aesthetics image generation model from Recraft that produces ~2K-resolution images with enhanced photorealism, smooth gradients, and strong adherence to short prompts,…
Qwen3.6 Max Preview is Qwen’s flagship proprietary large language model focused on high‑end reasoning and agentic coding, offered as an early-access cloud API. It features a…