- Text Generation
Qwen3 VL 30B A3B Instruct is a 30B-parameter Mixture-of-Experts vision-language model from Qwen, offering strong multimodal understanding and generation with a 262K-token context window. It is…
Powered by Qwen
Qwen3.6 Max Preview is Qwen’s flagship proprietary large language model focused on high‑end reasoning and agentic coding, offered as an early-access cloud API. It features a very long context window and improved world knowledge and instruction following compared with earlier Qwen3.6 models.
Output tokens per second · Higher is better
Seconds · Lower is better
USD per 1M tokens (blended) · Lower is better
About the model
Qwen3.6 Max Preview is a next-generation, closed-weight flagship large language model from Qwen (Alibaba) optimized for agentic coding, long-context reasoning, and cloud deployment. It is mainly used for autonomous and tool-using coding agents, handling complex software engineering tasks and benchmark-grade code reasoning. It is also applied to general-purpose assistant use cases that need strong world knowledge, precise instruction following, and long-context document or workspace analysis. It belongs to the Qwen3.6 model family and is positioned as a higher-end successor to models such as Qwen3.6-Plus and the open-source Qwen3.6 series.
Model capabilities
Acts as a high-end conversational assistant with strong instruction following, world knowledge, and multi-turn dialogue management for complex tasks.
Excels at software development assistance, agentic coding workflows, and achieving top scores on benchmarks like SWE-bench and Terminal-Bench.
Provides native reasoning modes and structured outputs, supporting long-context chain-of-thought style problem solving and tool-using agents.
Supports many languages for prompts and responses, enabling cross-lingual reasoning and content generation across global use cases.
Can read and extract information from provided text snippets or documents to support summarization, transformation, and downstream tasks.
Use cases
Transparent pricing
LLM API offers the lowest cost and best performance for Qwen3.6 Max Preview–class models.
| Provider | Region | Latency | Throughput | Uptime | Input ($/1M) | Output ($/1M) | Context |
|---|---|---|---|---|---|---|---|
| LLM API BEST | Global | 160ms | 120 tps | 99.99% | $0.20 | $0.60 | 128K |
| Qwen | Global | ~220ms | ~70 tps | ~99.9% | ~$0.30 | ~$0.90 | 128K |
| Alibaba Cloud | AP Southeast | ~250ms | ~60 tps | ~99.9% | ~$0.35 | ~$1.00 | 128K |
| OpenRouter | Global | ~240ms | ~80 tps | ~99.9% | ~$0.32 | ~$0.96 | 128K |
| Together AI | US East | ~230ms | ~75 tps | ~99.9% | ~$0.28 | ~$0.85 | 128K |
Performance benchmarks
| Metric | Qwen3.6 Max Preview | GPT-4.1 | Claude 3.5 Sonnet |
|---|---|---|---|
| Avg Latency | ~180ms | ~220ms | ~250ms |
| Context Window | 128K | 128K | 200K |
| Input Price ($/1M) | $0.80 | $5.00 | $3.00 |
| Output Price ($/1M) | $2.40 | $15.00 | $15.00 |
| Max Output Tokens | 8K | 4K | 4K |
| Throughput | 48 tps | 40 tps | 36 tps |
| Uptime | 99.9% | 99.9% | 99.9% |
30-day usage via LLM API
Architecture & Integration
One unified API. Every major model. Built-in reliability, cost control, and observability.
Dynamically route each request to the best model across providers using rules, metadata, and performance signals—without changing your integration or redeploying code.
One endpoint, any modelBalance latency, quality, and token prices automatically with configurable policies, so you minimize spend while keeping performance and SLAs under control.
Optimize tokens, not codeDefine multi-step fallback chains across models and regions to survive outages, rate limits, and timeouts—without complex client-side error handling.
Never drop a requestGet end-to-end traces, metrics, and structured logs for every call, including provider-level breakdowns, to debug issues and tune routing strategies in minutes.
See every token hopCall high-level tasks like chat, extract, classify, or generate instead of vendor-specific APIs, and swap underlying models without rewriting business logic.
Code to tasks, not vendorsRun massive offline jobs—evaluations, backfills, reprocessing—through a single API with concurrency control, retries, and cost tracking built in.
Millions of calls, one pipelineDecision guide
FAQ
Qwen3.6 Max Preview is a large language model from Qwen focused on high-quality reasoning, coding, and general-purpose text generation.
It excels at complex reasoning, multi-step problem solving, code generation, data analysis assistance, and building advanced chat or agentic applications.
Qwen3.6 Max Preview pricing on LLM.API is usage-based per 1,000 tokens; check your LLM.API dashboard or pricing docs for current rates.
Qwen3.6 Max Preview supports a large context window suitable for long conversations and multi-file prompts; refer to LLM.API docs for the exact token limit.
Typical latency is comparable to other large frontier models, with first-token times depending on load, model size, and your selected LLM.API region.
Through LLM.API, Qwen3.6 Max Preview currently supports text input and output; check the docs to confirm any multimodal capabilities or updates.
Use the standard LLM.API chat or completions endpoint, setting the model parameter to "Qwen3.6 Max Preview" and including your messages payload.
It targets strong reasoning and coding performance with competitive quality-to-cost, making it an alternative to top-tier models from other providers.
It can still hallucinate, produce incorrect code, mishandle edge cases, or reflect training-data biases, so critical outputs should be validated.
Fine-tuning availability depends on LLM.API features at the time; check the fine-tuning section to see if this model is supported.
Compare
Qwen3 VL 30B A3B Instruct is a 30B-parameter Mixture-of-Experts vision-language model from Qwen, offering strong multimodal understanding and generation with a 262K-token context window. It is…
GPT Chat Latest is OpenAI’s most up-to-date GPT-based chat model, offering strong general-purpose reasoning, coding, and writing capabilities. It is designed for interactive conversations and assistance…
Trinity Large Thinking is Arcee AI’s open-weight, 398–400B-parameter sparse Mixture-of-Experts model focused on advanced reasoning and long-horizon agentic tasks. It is notable for activating only about…