- Instruction Following
Qwen3.6 35B A3B is an open-weight, multimodal Mixture-of-Experts model with 35 billion parameters (about 3 billion active per token), designed for long-context reasoning, coding, and vision-language…
Powered by Qwen
Qwen3.5 Plus 2026-04-20 is a large-scale, proprietary multimodal language model from Qwen (Alibaba) that offers a 1M-token context window and strong reasoning and vision capabilities for advanced agentic workflows.
Output tokens per second · Higher is better
Seconds · Lower is better
USD per 1M tokens (blended) · Lower is better
About the model
Qwen3.5 Plus 2026-04-20 is an updated April 2026 release of Qwen’s flagship Qwen3.5 Plus multimodal language model with a 1M-token context window. It is mainly used for complex text and code generation tasks that benefit from long-context understanding, such as processing large document collections or repositories in a single session. It is also used for multimodal reasoning over text and images in applications like visual question answering, data analysis, and tool-using AI agents. It belongs to the Qwen3.5 family of models, an evolution of earlier Qwen and Qwen2 generations from Alibaba.
Model capabilities
Engages in multi-turn dialogue, follows complex instructions, and maintains conversational context across diverse general-purpose assistant tasks.
Understands and generates source code, explains programming concepts, and helps debug or refactor code in multiple languages.
Interprets images, identifying objects, scenes, and relationships to answer questions or provide descriptions about visual content.
Translates between multiple languages, preserving meaning and tone while adapting phrasing to sound natural in the target language.
Extracts machine-readable text from images or scanned documents, enabling search, editing, and downstream processing of visual text content.
Use cases
Transparent pricing
LLM API offers the lowest token prices and best performance for Qwen3.5-class models.
| Provider | Region | Latency | Throughput | Uptime | Input ($/1M) | Output ($/1M) | Context |
|---|---|---|---|---|---|---|---|
| LLM API BEST | Global | 120ms | 80 tps | 99.99% | $0.10 | $0.30 | 256K |
| Qwen (Official) | Global | ~220ms | ~40 tps | ~99.9% | ~$0.20 | ~$0.60 | ~200K |
| Alibaba Cloud AI | APAC | ~260ms | ~35 tps | ~99.9% | ~$0.22 | ~$0.65 | ~128K |
| OpenRouter | Global | ~240ms | ~30 tps | ~99.8% | ~$0.25 | ~$0.70 | ~128K |
Performance benchmarks
| Metric | Qwen3.5 Plus 2026-04-20 | GPT-4.1 | Claude 3.5 Sonnet |
|---|---|---|---|
| Avg Latency | ~220ms | ~250ms | ~320ms |
| Context Window | 128K | 128K | 200K |
| Input Price ($/1M) | $0.40 | $5.00 | $3.00 |
| Output Price ($/1M) | $1.20 | $15.00 | $15.00 |
| Max Output Tokens | 8K | 8K | 8K |
| Throughput | 60 tps | 40 tps | 35 tps |
| Uptime | 99.9% | 99.9% | 99.9% |
30-day usage via LLM API
Architecture & Integration
One unified API. Every major model. Built-in reliability, cost control, and observability.
Automatically route each request to the optimal model across providers based on cost, latency, and quality—without changing your integration or redeploying code.
One endpoint, every modelControl spend with per-request cost policies, dynamic model selection, and real-time price visibility so you can scale usage without surprise bills or manual tuning.
Lower cost, same outputDefine provider- and model-level fallback chains so requests transparently fail over on errors, rate limits, or outages—no custom retry code needed.
Stay online by defaultInspect every call with traces, costs, latencies, and model choices in one place, making it easy to debug prompts and optimize performance in production.
See every tokenUse high-level task APIs for chat, tools, RAG, and structured outputs so you can swap models and providers without rewriting business logic.
Code to tasks, not modelsSubmit large batches with automatic chunking, concurrency control, and retries to process millions of requests efficiently while respecting rate limits across providers.
Scale jobs, not scriptsDecision guide
FAQ
Qwen3.5 Plus 2026-04-20 is a general-purpose large language model by Qwen exposed through the LLM.API unified AI gateway.
Qwen3.5 Plus 2026-04-20 is best for robust text generation, coding assistance, and instruction-following tasks where strong reasoning and reliability matter.
Qwen3.5 Plus 2026-04-20 supports a context window of up to 128K tokens via LLM.API, depending on your configured limits.
Qwen3.5 Plus 2026-04-20 is text-only through LLM.API and does not support image, audio, or video inputs.
On LLM.API, Qwen3.5 Plus 2026-04-20 is billed per token, with separate input and output token rates defined in your LLM.API pricing plan.
Typical end-to-end latency is comparable to other mid-sized hosted LLMs, but depends on prompt length, output size, and current LLM.API load.
Use the LLM.API chat or completions endpoint and set the model parameter to "Qwen3.5 Plus 2026-04-20" in your request payload.
Qwen3.5 Plus 2026-04-20 targets a balance of quality and cost, often cheaper than flagship frontier models but stronger than lightweight baselines.
It can hallucinate facts, lacks real-time knowledge beyond its training cutoff, and should not be solely relied on for safety-critical decisions.
Tool use or browsing is only available if you implement those capabilities application-side; the base model has no built-in external access.
Compare
Qwen3.6 35B A3B is an open-weight, multimodal Mixture-of-Experts model with 35 billion parameters (about 3 billion active per token), designed for long-context reasoning, coding, and vision-language…
LFM2.5-1.2B-Instruct (free) is a 1.2B-parameter, instruction-tuned hybrid language model from LiquidAI, optimized for fast, on-device inference with a ~32k token context window. It offers general-purpose conversational…
Claude Opus Latest is Anthropic’s current flagship Opus-tier large language model, designed for complex reasoning, coding, and knowledge work with strong safety and alignment features. It…