- Text Generation
Rnj-1 Instruct is an 8B-parameter, instruction-tuned open-weight model from EssentialAI, optimized for code generation, STEM reasoning, and agentic tool-using workflows with a 32K context window.
Powered by Mistral
Mistral Medium 3.5 is a 128B-parameter dense large language model from Mistral, designed as a flagship "merged" model for strong general-purpose reasoning, coding, and long-context tasks. It targets a balance of capability, latency, and cost for production AI applications.
Output tokens per second · Higher is better
Seconds · Lower is better
USD per 1M tokens (blended) · Lower is better
About the model
Mistral Medium 3.5 is a dense 128B-parameter large language model by Mistral optimized for general-purpose text understanding and generation with a 256K-token context window. It is used for software development assistance, long-running autonomous or remote coding agents, and other knowledge work requiring reliable reasoning over large contexts. It also serves as a default or backbone model in several Mistral products and third-party platforms for assistants, agents, and enterprise applications. It follows earlier Mistral Medium 3-series models and complements other Mistral families such as Mistral Large and the smaller Ministral models.
Model capabilities
Processes both text and images, performing instruction-following, logical reasoning, and complex problem solving within a unified 128B dense model.
Provides strong instruction-following, conversational responses, and system-prompt control suitable for assistants, support bots, and long-context interactions.
Generates, debugs, and refactors code, enabling sophisticated coding agents and long-running software engineering workflows with high benchmark performance.
Understands and generates text in dozens of languages, including major European and Asian languages, for global applications and content.
Performs OCR and document understanding with a custom vision encoder handling variable image sizes, layouts, and structured visual annotations.
Use cases
Transparent pricing
LLM API offers the lowest cost and highest performance for Mistral Medium 3.5–class models.
| Provider | Region | Latency | Throughput | Uptime | Input ($/1M) | Output ($/1M) | Context |
|---|---|---|---|---|---|---|---|
| LLM API BEST | Global | 90ms | 120 tps | 99.99% | $0.20 | $0.60 | 128K |
| Mistral (direct) | EU West | ~180ms | ~60 tps | 99.9% | ~$0.25 | ~$0.75 | 128K |
| Azure (Mistral-compatible) | US East | ~220ms | ~50 tps | 99.9% | ~$0.35 | ~$1.00 | 128K |
| AWS Bedrock (Mistral-like) | US West | ~210ms | ~55 tps | 99.9% | ~$0.30 | ~$0.90 | 128K |
| Replicate (Mistral-compatible) | Global | ~260ms | ~30 tps | 99.5% | ~$0.40 | ~$1.20 | ~64K |
Performance benchmarks
| Metric | Mistral Medium 3.5 | GPT-4.1 Mini | Claude 3.5 Haiku |
|---|---|---|---|
| Avg Latency | ~180ms | ~220ms | ~250ms |
| Context Window | 128K | 128K | 200K |
| Input Price ($/1M) | $0.60 | $0.15 | $0.25 |
| Output Price ($/1M) | $1.80 | $0.60 | $1.25 |
| Max Output Tokens | 4K | 4K | 4K |
| Throughput | ~70 tps | ~80 tps | ~65 tps |
| Uptime | 99.9% | 99.9% | 99.9% |
30-day usage via LLM API
Architecture & Integration
One unified API. Every major model. Built-in reliability, cost control, and observability.
Dynamically route requests across providers and model families via one endpoint, using rules or performance data to balance quality, latency, and reliability automatically.
One endpoint, every modelEnforce per-project and per-route budgets, downshift to cheaper models automatically, and compare providers so you never overspend for the same output quality.
Control spend by designDefine provider- and model-level fallbacks so requests transparently fail over on timeouts, rate limits, or outages—without changing your application code.
No single point of failureGet unified logs, traces, and metrics for every request across providers—latency, errors, tokens, and cost—so you can debug and optimize production workloads quickly.
See every token spentCall high-level tasks—chat, tools, RAG, image, embeddings—through a stable API while LLM.API handles prompt shaping, model quirks, and provider differences underneath.
Code to tasks, not modelsSubmit large batches of prompts to any provider with automatic chunking, retries, and aggregation, maximizing throughput while staying within rate and budget limits.
Ship at batch scaleDecision guide
FAQ
Mistral Medium 3.5 is a proprietary large language model by Mistral aimed at general-purpose coding, reasoning, and chat workloads with balanced cost and quality.
Mistral Medium 3.5 supports up to a 32K token context window for combined input and output via LLM.API.
Mistral Medium 3.5 usage on LLM.API is billed per input and output token; check your LLM.API pricing page for current rates.
Mistral Medium 3.5 is optimized for low-latency streaming responses, with actual speed depending on prompt size and your network conditions.
Through LLM.API, Mistral Medium 3.5 currently supports text input and text output only.
Select the Mistral provider and the Mistral Medium 3.5 model ID in your LLM.API client or HTTP requests to route calls to this model.
Mistral Medium 3.5 is best for production chatbots, code generation, data transformation, and general reasoning tasks needing a balance of capability and price.
Compared with lighter Mistral models, Mistral Medium 3.5 generally offers stronger reasoning, coding, and instruction-following at higher cost and latency.
Mistral Medium 3.5 can hallucinate incorrect facts, lacks real-time internet access, and should not be used for unsupervised high-stakes decisions.
Direct fine-tuning of Mistral Medium 3.5 is not available via LLM.API; use prompting or retrieval-augmented techniques instead.
Compare
Rnj-1 Instruct is an 8B-parameter, instruction-tuned open-weight model from EssentialAI, optimized for code generation, STEM reasoning, and agentic tool-using workflows with a 32K context window.
GPT-5.2 is an OpenAI large language model in the GPT-5 family, designed for advanced natural language understanding and generation across many tasks. It emphasizes improved reasoning,…
all-MiniLM-L12-v2 is a compact Sentence Transformers model that generates high-quality sentence embeddings for efficient semantic search and similarity tasks. It is notable for its strong performance-to-size…