- Instruction Following
Qwen3.5 Plus 2026-04-20 is a large-scale, proprietary multimodal language model from Qwen (Alibaba) that offers a 1M-token context window and strong reasoning and vision capabilities for…
Powered by Openrouter
Free Models Router is an OpenRouter meta-model that automatically routes requests to compatible free models, providing no-cost inference across multiple underlying LLMs. It filters candidates based on required capabilities such as text, image input, and tool use.
Output tokens per second · Higher is better
Seconds · Lower is better
USD per 1M tokens (blended) · Lower is better
About the model
Free Models Router is an OpenRouter routing model (`openrouter/free`) that selects eligible free models to handle each request instead of generating outputs itself. It is mainly used for cost-free experimentation, prototyping, and general text generation across whatever free models are currently available on OpenRouter. It also supports multimodal text-and-image inputs and can be used in applications that require capabilities like vision, reasoning, and tool use without committing to a single backend model. The model belongs to OpenRouter’s router category and is part of its family of meta-models that dynamically dispatch traffic to different hosted LLMs.
Model capabilities
Routes user requests across multiple free large language models on OpenRouter, selecting an appropriate backend model for each call.
Supports conversational interactions with natural language understanding and generation, forwarding messages to suitable underlying chat-optimized models.
Relays text translation requests to underlying models that can convert content between multiple languages with reasonable quality.
Can pass image inputs to compatible backend models to extract or utilize textual content contained within those images.
Forwards images to vision-capable models that can interpret visual content, answer questions about images, and describe visual scenes.
Use cases
Transparent pricing
LLM API offers the lowest effective token costs and best performance among Free Models Router–class APIs.
| Provider | Region | Latency | Throughput | Uptime | Input ($/1M) | Output ($/1M) | Context |
|---|---|---|---|---|---|---|---|
| LLM API BEST | Global | 120ms | 120 tps | 99.99% | $0.00 | $0.00 | 128K |
| Openrouter | Global | ~350ms | ~40 tps | ~99.9% | $0.00 | $0.00 | ~32K |
| OpenAI | Global | ~300ms | ~80 tps | 99.9% | ~$0.10 | ~$0.30 | 128K |
| Anthropic | US East | ~320ms | ~70 tps | 99.9% | ~$0.20 | ~$0.60 | 200K |
| Google AI Studio | Global | ~280ms | ~75 tps | ~99.9% | ~$0.08 | ~$0.24 | ~128K |
Performance benchmarks
| Metric | Free Models Router (Openrouter) | OpenAI gpt-4o-mini | Anthropic Claude 3 Haiku |
|---|---|---|---|
| Avg Latency | ~180ms | ~250ms | ~300ms |
| Context Window | ~128K | 128K | 200K |
| Input Price ($/1M) | ~$0.00 | $0.15 | $0.25 |
| Output Price ($/1M) | ~$0.00 | $0.60 | $1.25 |
| Max Output Tokens | ~4K | 4K | 4K |
| Throughput | ~60 tps | ~40 tps | ~35 tps |
| Uptime | ~99.5% | ~99.9% | ~99.9% |
30-day usage via LLM API
Architecture & Integration
One unified API. Every major model. Built-in reliability, cost control, and observability.
Dynamically route each request to the best model across providers based on latency, capability, and cost—without changing your integration or redeploying code.
One endpoint, every modelAutomatically balance quality and spend with configurable cost policies, price-aware routing, and transparent usage data so teams can ship faster without surprise bills.
Optimize quality per dollarDefine provider-agnostic fallback chains that auto-retry on errors, timeouts, or rate limits to keep your AI features online even when vendors degrade.
Failure-tolerant by designTrace every call across providers with logs, metrics, and structured events for prompts, latencies, and errors, making debugging and optimization straightforward in production.
Full visibility into LLMsDescribe what you want—chat, extraction, tools, RAG—and let LLM.API handle provider-specific quirks so teams can swap models without refactoring payloads.
Code to tasks, not vendorsProcess large workloads efficiently with batched requests, concurrency controls, and rate-aware scheduling, dramatically reducing per-request overhead and infrastructure complexity.
Scale workloads, not overheadDecision guide
FAQ
Free Models Router is an OpenRouter endpoint that automatically routes your request to one of several free-tier large language models.
It is best for cost-free experimentation, prototyping, and low-stakes applications where occasional quality or availability variations are acceptable.
Requests through LLM.API are billed according to LLM.API’s pricing for the Free Models Router endpoint, even though the underlying OpenRouter tier is free.
The effective context window depends on the specific underlying free model selected, so you should assume a relatively small to medium context size.
Latency can vary per request because different backing models and infrastructures may be selected, so you should not rely on consistent response times.
It primarily supports text-in, text-out interactions; image or other modalities are not guaranteed and depend on the routed underlying model.
Use the LLM.API endpoint with the model identifier corresponding to OpenRouter’s Free Models Router and authenticate with your LLM.API key.
Pinned premium models usually provide more predictable quality, latency, and features, while Free Models Router optimizes primarily for zero model-side cost.
Yes, both OpenRouter’s free tier and LLM.API’s account-level limits can restrict throughput, rate, or total tokens for this model.
You may see inconsistent model behavior, varying capabilities, and occasional capacity errors because requests are routed across multiple free models.
Compare
Qwen3.5 Plus 2026-04-20 is a large-scale, proprietary multimodal language model from Qwen (Alibaba) that offers a 1M-token context window and strong reasoning and vision capabilities for…
GPT-5.2 Chat is an OpenAI conversational language model designed for interactive dialogue and task assistance. It focuses on providing coherent, context-aware responses across a wide range…
Hailuo 2.3 by MiniMax is a high-fidelity AI video generation model designed for realistic, cinematic 1080p clips from text or image prompts, with strong motion, physics,…