- Instruction Following
Ling-2.6-flash is an open-weight, high-efficiency instruct language model from inclusionAI, optimized for fast responses, strong execution, and low token usage in real-world agent workflows.
Powered by Poolside
Laguna M.1 (free) is Poolside’s flagship agentic coding language model, offered with a free access tier via API and platforms like OpenRouter. It is a large Mixture-of-Experts model optimized for complex software engineering tasks and long-context coding workflows.
Output tokens per second · Higher is better
Seconds · Lower is better
USD per 1M tokens (blended) · Lower is better
About the model
Laguna M.1 (free) is a 225B-parameter Mixture-of-Experts language model from Poolside focused on agentic coding and complex software engineering tasks. It is mainly used for autonomous or assisted code generation, editing, and debugging within terminal agents and cloud dev environments, with long-context reasoning and tool-calling support for sophisticated coding workflows. A free tier of this proprietary model is exposed via OpenAI-compatible APIs and third-party routing platforms for experimentation and integration into existing tools. Laguna M.1 belongs to Poolside’s Laguna model family and serves as the larger companion to the open-weight Laguna XS.2 model.
Model capabilities
Specialized in complex, long-horizon software engineering tasks, autonomously editing, refactoring, and extending codebases across multiple files.
Invokes external tools and APIs from natural language instructions to run tests, interact with systems, and orchestrate workflows.
Performs strong logical and structural reasoning over code, enabling reliable bug fixing, feature implementation, and test creation.
Processes very large text contexts, maintaining coherence over long coding sessions, logs, and multi-file repositories within a single prompt.
Understands and generates text in multiple languages, supporting international codebases, comments, and documentation across diverse locales.
Use cases
Transparent pricing
LLM API offers the lowest cost and latency for Laguna M.1–class models.
| Provider | Region | Latency | Throughput | Uptime | Input ($/1M) | Output ($/1M) | Context |
|---|---|---|---|---|---|---|---|
| LLM API BEST | Global | 110ms | 85 tps | 99.99% | $0.20 | $0.20 | 128K |
| Poolside (Laguna M.1 free tier) | Global | ~220ms | ~25 tps | ~99.0% | $0.00 | $0.00 | ~32K |
| Poolside (Laguna M.1 paid) | Global | ~180ms | ~40 tps | ~99.5% | ~$0.30 | ~$0.30 | ~64K |
| OpenRouter (Laguna-equivalent model) | Global | ~260ms | ~30 tps | ~99.9% | ~$0.40 | ~$0.40 | ~128K |
| Together AI (Laguna-equivalent model) | US East | ~250ms | ~35 tps | ~99.9% | ~$0.35 | ~$0.35 | ~128K |
Performance benchmarks
| Metric | Laguna M.1 (free) | GPT-4o mini | Claude 3.5 Haiku |
|---|---|---|---|
| Avg Latency | ~800ms | ~600ms | ~700ms |
| Context Window | 128K | 128K | 200K |
| Input Price ($/1M) | ~$0.00 | ~$0.15 | ~$0.25 |
| Output Price ($/1M) | ~$0.00 | ~$0.60 | ~$1.25 |
| Max Output Tokens | 4K | 4K | 4K |
| Throughput | ~30 tps | ~50 tps | ~40 tps |
| Uptime | 99.0% | 99.9% | 99.9% |
30-day usage via LLM API
Architecture & Integration
One unified API. Every major model. Built-in reliability, cost control, and observability.
Dynamically route each request to the optimal model across providers based on latency, cost, or quality—no client changes required as your stack evolves.
One endpoint, any modelAutomatically shift traffic to cheaper equivalents, apply smart downscaling, and cap spend per workspace so you can experiment without surprise bills.
Cut cost, not coverageDefine failover policies once and let LLM.API transparently retry on alternate models or providers when timeouts, errors, or quota limits hit.
Resilience by defaultGet request tracing, latency breakdowns, cost per call, and model-level success metrics in one place to debug faster and tune routing with real data.
See every tokenDescribe tasks like chat, extraction, or classification and let LLM.API pick the right model and parameters—no more provider-specific boilerplate everywhere.
Code to tasks, not modelsStreamline large workloads with optimized batching, concurrency controls, and rate-limit aware scheduling to maximize throughput while staying within provider quotas.
Scale jobs, stay safeDecision guide
FAQ
Laguna M.1 (free) is a Poolside language model available via LLM.API, suited for general-purpose text generation and coding assistance without usage fees.
Laguna M.1 (free) is best for iterative coding, debugging, and explaining code, plus general chat and lightweight reasoning tasks.
Laguna M.1 (free) is exposed as a zero-cost tier on LLM.API, charging no per-token fees but possibly subject to fair-use limits.
Laguna M.1 (free) supports a context window of up to 16K tokens for combined prompt and completion.
Laguna M.1 (free) is optimized for low latency interactive use, typically streaming first tokens within a second under normal load.
Laguna M.1 (free) is text-only, supporting text input and text output, without native image, audio, or video understanding.
You select provider "Poolside" and model "Laguna M.1 (free)" in your LLM.API request, passing messages in the standard Chat Completions schema.
Laguna M.1 (free) targets performance comparable to strong mid-range open models while emphasizing low friction access and stable, predictable behavior.
Laguna M.1 (free) can hallucinate facts, lacks real-time browsing tools, and may underperform larger frontier models on complex reasoning or domain-expert tasks.
Laguna M.1 (free) can be used with LLM.API’s tool or function-calling abstractions when you define tools in the request and handle structured outputs.
Compare
Ling-2.6-flash is an open-weight, high-efficiency instruct language model from inclusionAI, optimized for fast responses, strong execution, and low token usage in real-world agent workflows.
GPT-5.1 is an OpenAI language model; as of mid-2026, OpenAI has not publicly released technical details or documentation about it.
GPT-5 Pro is an OpenAI model, but as of mid-2026 OpenAI has not publicly released technical details, benchmarks, or official documentation about it. Public, verifiable information…