- Instruction Following
Qwen3.6 35B A3B is an open-weight, multimodal Mixture-of-Experts model with 35 billion parameters (about 3 billion active per token), designed for long-context reasoning, coding, and vision-language…
Powered by Qwen
Qwen3.6 27B is a 27-billion-parameter large language model from Qwen, part of the Qwen3.6 series. It is designed to provide strong general-purpose reasoning and language capabilities within a relatively large, yet still deployable, model size.
Output tokens per second · Higher is better
Seconds · Lower is better
USD per 1M tokens (blended) · Lower is better
About the model
Qwen3.6 27B is a 27B-parameter large language model developed by Qwen for general-purpose AI assistance. It is mainly used for tasks such as multi-turn dialogue, content drafting, and code or data analysis support. It is also applied in building domain-specific assistants and applications that require stronger reasoning than smaller models in the same family. It belongs to the Qwen3.6 family of models, which follow earlier Qwen model generations.
Model capabilities
Engages in multi-turn, open-domain dialogue, following instructions and maintaining context for helpful, coherent conversational responses.
Analyzes and summarizes long-form text, extracting key information, making inferences, and answering detailed questions about content.
Interprets images to identify objects, scenes, and relationships, and answers questions about visual content when appropriately configured.
Translates between multiple languages, preserving meaning and tone while adapting phrasing to target-language conventions.
Reads and extracts textual content from images, enabling recognition of signs, documents, and screen text for downstream processing.
Use cases
Transparent pricing
LLM API offers the lowest cost and latency for Qwen3.6‑class 27B models.
| Provider | Region | Latency | Throughput | Uptime | Input ($/1M) | Output ($/1M) | Context |
|---|---|---|---|---|---|---|---|
| LLM API BEST | Global | 120ms | 65 tps | 99.99% | $0.30 | $0.60 | 256K |
| Qwen | Global | ~220ms | ~40 tps | ~99.9% | ~$0.40 | ~$0.80 | ~128K |
| Alibaba Cloud | APAC | ~250ms | ~35 tps | 99.9% | ~$0.45 | ~$0.90 | ~128K |
| Together AI | US East | ~210ms | ~45 tps | ~99.9% | ~$0.38 | ~$0.76 | ~128K |
| Fireworks AI | US West | ~200ms | ~42 tps | ~99.9% | ~$0.36 | ~$0.72 | ~128K |
Performance benchmarks
| Metric | Qwen3.6 27B | LLaMA 3.1 34B | Mistral Large 2 |
|---|---|---|---|
| Avg Latency | ~220ms | ~260ms | ~250ms |
| Context Window | 128K | 128K | 128K |
| Input Price ($/1M) | ~$0.60 | ~$1.00 | ~$2.00 |
| Output Price ($/1M) | ~$1.80 | ~$3.00 | ~$6.00 |
| Max Output Tokens | 8K | 8K | 8K |
| Throughput | ~45 tps | ~40 tps | ~42 tps |
| Uptime | ~99.9% | ~99.9% | ~99.9% |
30-day usage via LLM API
Architecture & Integration
One unified API. Every major model. Built-in reliability, cost control, and observability.
Automatically route each request to the best model across providers using latency, cost, and quality signals, so you keep performance high without hard-coding vendor logic.
One endpoint, every model.Enforce budgets, choose cheaper equivalents, and get transparent per-provider spend insights so you can scale usage without surprise bills or manual price tuning.
Cut spend, keep quality.Define automatic failover chains so if a provider is slow, down, or rate-limited, requests seamlessly retry against backups without code changes or user-visible errors.
Stay online, automatically.Trace every call across providers with unified logs, metrics, and latency breakdowns, making it easy to debug failures, tune routes, and prove SLAs.
See every token hop.Express work as high-level tasks—chat, retrieval, tools, scoring—while LLM.API handles prompts, parameters, and model quirks for you under a consistent schema.
Think tasks, not models.Ship millions of evaluations, datasets, or offline jobs through a single batch API with automatic chunking, retries, and progress tracking across providers.
Scale evaluations effortlessly.Decision guide
FAQ
Qwen3.6 27B is a 27-billion-parameter large language model by Qwen, focused on high-quality text generation and reasoning through LLM.API.
Qwen3.6 27B is best for complex reasoning, multi-step coding, and high-quality long-form writing where accuracy matters more than minimal latency or cost.
Qwen3.6 27B supports a context window of up to 32,768 tokens per request on LLM.API.
Qwen3.6 27B is exposed on LLM.API as a text-only model, accepting and returning UTF-8 text sequences.
Qwen3.6 27B generally has higher latency than smaller models, so it suits background tasks more than ultra-low-latency interactive workloads.
Use the standard LLM.API chat or completions endpoint and set the model parameter to the Qwen3.6 27B identifier shown in the catalog.
Compared to smaller Qwen models, Qwen3.6 27B generally offers stronger reasoning and coding performance at the cost of higher latency and price.
Qwen3.6 27B can hallucinate facts, struggle with very domain-specific knowledge, and should not be used without human review for critical decisions.
Qwen3.6 27B pricing is usage-based per input and output tokens; check your LLM.API pricing page for the latest specific rates.
Within its 32K-token context, Qwen3.6 27B maintains good coherence, but very long conversations may still cause earlier details to be forgotten.
Compare
Qwen3.6 35B A3B is an open-weight, multimodal Mixture-of-Experts model with 35 billion parameters (about 3 billion active per token), designed for long-context reasoning, coding, and vision-language…
GPT-5.1 Chat is an OpenAI conversational AI model designed for high-quality dialogue, reasoning, and assistance across many domains. It is notable for improved reliability, instruction-following, and…
GLM 5V Turbo is Z.ai’s native multimodal large language model optimized for vision-based coding and agentic workflows, able to process images, video, and text for complex…