- Instruction Following
Ring-2.6-1T is a trillion-parameter-scale open-weight "thinking" language model from inclusionAI, designed for real-world agent and coding workflows that need strong reasoning with efficient execution.
Powered by Qwen
Qwen3.5 397B A17B is a large-scale language model from Qwen with roughly 397 billion parameters, designed for advanced reasoning and multilingual understanding. It targets high-end inference scenarios where strong general capabilities and model depth are required.
Output tokens per second · Higher is better
Seconds · Lower is better
USD per 1M tokens (blended) · Lower is better
About the model
Qwen3.5 397B A17B is a 397-billion-parameter Qwen language model optimized for powerful, general-purpose AI assistance. It is used for complex text generation and understanding tasks such as drafting, analysis, and conversation. It is also applied in demanding reasoning, coding, and knowledge-intensive applications where very large models are preferred. It belongs to the Qwen (Qwen2/Qwen2.5/Qwen3.x) family of large language models developed by Qwen.
Model capabilities
Engages in multi-turn conversations, follows complex instructions, and maintains context for reasoning, coding help, and detailed explanations.
Interprets uploaded images to identify objects, text, layouts, and visual relationships, supporting description, reasoning, and grounded question answering.
Extracts and structures text from scanned documents, screenshots, and complex layouts, enabling downstream analysis, search, and transformation tasks.
Supports tool-using workflows, including calling external APIs, running code-like reasoning, and monitoring iterative steps for complex tasks.
Translates between many languages while preserving meaning, tone, and formatting, useful for cross-lingual communication and content localization.
Use cases
Transparent pricing
LLM API offers the lowest cost and highest performance for Qwen3.5‑class 397B models.
| Provider | Region | Latency | Throughput | Uptime | Input ($/1M) | Output ($/1M) | Context |
|---|---|---|---|---|---|---|---|
| LLM API BEST | Global | 120ms | 80 tps | 99.99% | $0.30 | $0.60 | 256K |
| Qwen | Global | ~220ms | ~50 tps | 99.9% | ~$0.50 | ~$1.00 | ~128K |
| Alibaba Cloud | APAC East | ~260ms | ~40 tps | 99.9% | ~$0.55 | ~$1.10 | ~128K |
| AWS Marketplace (Qwen Partner) | US East | ~250ms | ~35 tps | 99.9% | ~$0.60 | ~$1.20 | ~128K |
| Azure Marketplace (Qwen Partner) | EU West | ~240ms | ~38 tps | 99.9% | ~$0.58 | ~$1.15 | ~128K |
Performance benchmarks
| Metric | Qwen3.5 397B A17B | GPT-4.1 | Claude 3.5 Sonnet |
|---|---|---|---|
| Avg Latency | ~180ms | ~220ms | ~200ms |
| Context Window | 128K | 128K | 200K |
| Input Price ($/1M) | ~$2.00 | ~$5.00 | ~$3.00 |
| Output Price ($/1M) | ~$6.00 | ~$15.00 | ~$15.00 |
| Max Output Tokens | 8K | 4K | 8K |
| Throughput | ~70 tps | ~50 tps | ~55 tps |
| Uptime | ~99.9% | ~99.9% | ~99.9% |
30-day usage via LLM API
Architecture & Integration
One unified API. Every major model. Built-in reliability, cost control, and observability.
Automatically route each request to the optimal model across providers based on latency, price, and performance—no client changes required.
One endpoint, many modelsControl spend with per-route pricing rules, automatic downgrades, and usage caps while keeping SLAs and quality intact.
Cut costs, keep qualityDefine multi-step failover chains so requests seamlessly retry on backup models or providers when outages or timeouts occur.
Never go darkGet end-to-end traces, latency histograms, provider error rates, and payload logs in one place to debug and optimize quickly.
See every tokenUse task-level APIs (chat, tools, embeddings, rerank, image) that stay stable even as underlying models and vendors change.
Code to tasks, not vendorsSend massive batch jobs through a single endpoint with automatic sharding, rate limiting, and retries across providers.
Scale jobs, not scriptsDecision guide
FAQ
Qwen3.5 397B A17B is a large-scale Qwen language model accessible through LLM.API, optimized for complex reasoning, code, and high-quality text generation.
It excels at multi-step reasoning, advanced coding assistance, data analysis, and generating long-form, instruction-following content with strong coherence.
Pricing is usage-based per input and output token; check your LLM.API dashboard or pricing docs for the latest specific rates.
Qwen3.5 397B A17B supports a long context window suitable for extended conversations and documents; refer to LLM.API docs for the current token limit.
As a very large model it has higher latency than smaller Qwen variants, but LLM.API streams tokens progressively to improve perceived responsiveness.
Through LLM.API it supports text input and output; check the model capabilities section to confirm any additional modalities like images if enabled.
Use the standard chat or completion endpoint, specifying the model name "qwen3.5-397b-a17b" (or listed identifier) in your LLM.API request payload.
It generally offers stronger reasoning and generation quality than smaller Qwen models, at higher cost and latency per request.
It may hallucinate incorrect facts, struggle with real-time or proprietary data, and be too slow or expensive for latency-critical, high-throughput workloads.
Yes, you can run batch and backend workloads via LLM.API, but should account for its higher token cost and compute latency in your design.
Compare
Ring-2.6-1T is a trillion-parameter-scale open-weight "thinking" language model from inclusionAI, designed for real-world agent and coding workflows that need strong reasoning with efficient execution.
Reka Edge is a 7B-parameter multimodal vision-language model from RekaAI that processes text, image, and video inputs to generate text outputs, optimized for fast, efficient edge…
Gemini 2.5 Flash Lite Preview 09-2025 is a lightweight preview variant of Google’s Gemini 2.5 Flash-Lite model, optimized for fast, cost-efficient multimodal inference with long-context support.…