- Instruction Following
ERNIE 4.5 21B A3B Thinking is Baidu’s upgraded lightweight MoE language model optimized for deep reasoning, with a context window around 131K tokens and competitive pricing…
Powered by Qwen
Qwen3 VL 235B A22B Instruct is a 235B-parameter Mixture-of-Experts vision-language model from Qwen, offering open-weight, long-context (≈256K) multimodal reasoning over text, images, and video. It is instruction-tuned for chat-style interactions and agentic use, including GUI automation and tool use.
Output tokens per second · Higher is better
Seconds · Lower is better
USD per 1M tokens (blended) · Lower is better
About the model
Qwen3 VL 235B A22B Instruct is an open-weight, instruction-tuned Mixture-of-Experts vision-language model with 235B parameters (22B active) that supports text, image, and video inputs with a context window of about 256K tokens. It is mainly used for general multimodal chat and reasoning tasks such as visual question answering, document and chart understanding, and long-context analysis across mixed media. It is also applied to agentic workflows including GUI automation, visual code generation from mockups, and tool-using assistants in enterprise or research pipelines. The model belongs to the Qwen3-VL family of vision-language models developed by Qwen/Alibaba as a flagship high-capacity variant building on earlier Qwen and Qwen-VL generations.
Model capabilities
Understands and reasons over images and text jointly, enabling tasks like description, question answering, and visual-grounded instruction following.
Engages in multi-turn conversations, follows complex instructions, and performs reasoning, coding, and analysis across diverse textual domains.
Interprets screenshots, interfaces, and layouts, supporting tasks like element identification, navigation planning, and workflow explanation for applications.
Interprets technical prompts, reasons step-by-step, and can be integrated with tools or environments for advanced programmatic workflows.
Understands and generates multiple languages, enabling cross-lingual tasks such as explanation, paraphrasing, and language-aware reasoning over content.
Use cases
Transparent pricing
LLM API offers the lowest prices and best limits for Qwen3 VL–class models.
| Provider | Region | Latency | Throughput | Uptime | Input ($/1M) | Output ($/1M) | Context |
|---|---|---|---|---|---|---|---|
| LLM API BEST | Global | ~220ms | ~70 img/min | 99.99% | ~$0.40/1K tokens+image | ~$0.80/1K tokens | ~256K tokens+images |
| Qwen | Global | ~350ms | ~45 img/min | 99.9% | ~$0.90/1K tokens+image | ~$1.80/1K tokens | ~128K tokens+images |
| Alibaba Cloud | APAC East | ~420ms | ~40 img/min | 99.9% | ~$1.00/1K tokens+image | ~$2.00/1K tokens | ~128K tokens+images |
| AWS Marketplace | US East | ~380ms | ~38 img/min | 99.9% | ~$1.10/1K tokens+image | ~$2.20/1K tokens | ~128K tokens+images |
| Azure Marketplace | EU West | ~400ms | ~35 img/min | 99.9% | ~$1.20/1K tokens+image | ~$2.40/1K tokens | ~128K tokens+images |
Performance benchmarks
| Metric | Qwen3 VL 235B A22B Instruct | GPT-4.1 Vision | Claude 3.5 Sonnet Vision |
|---|---|---|---|
| Latency per Image | ~900ms | ~1.2s | ~1.0s |
| Throughput | ~40 img/s | ~35 img/s | ~30 img/s |
| Max Resolution | ~4K | ~4K | ~4K |
| Price per Image | ~$0.004 | ~$0.005 | ~$0.005 |
| Supported Formats | PNG, JPG, WebP, GIF | PNG, JPG, WebP, GIF | PNG, JPG, WebP, GIF |
| Uptime | ~99.9% | ~99.9% | ~99.9% |
30-day usage via LLM API
Architecture & Integration
One unified API. Every major model. Built-in reliability, cost control, and observability.
Dynamically route each request to the best model by latency, cost, or quality—no client changes required as providers, versions, or constraints evolve.
One endpoint, every modelControl spend with automatic price-based routing, per-project budgets, and cost insights so you can ship fast without surprise bills or manual tuning.
Optimize every tokenSurvive provider outages and rate limits with automatic cross-vendor failover, health checks, and configurable retries—all wired behind a single API.
Never drop a requestTrace every call across providers with logs, latency and error metrics, and cost breakdowns so you can debug, tune, and scale with real production data.
See every tokenDescribe the task, not the model. LLM.API picks the right tools, prompts, and providers so you keep logic clean and avoid brittle per-model code.
Code to tasks, not modelsRun massive offline jobs with provider-aware chunking, parallelization, and retries to safely process millions of items through a single unified interface.
Scale to millionsDecision guide
FAQ
Qwen3 VL 235B A22B Instruct is a large Qwen multimodal instruction-tuned model designed for high-quality vision-language and text-only reasoning tasks.
It excels at complex image understanding, detailed visual question answering, document analysis with OCR, code reasoning from screenshots, and advanced multi-step text reasoning.
Through LLM.API, it supports text input and output plus image inputs, enabling vision-language workflows and standard chat-style text generation.
The model supports a large-context window suitable for long conversations and multi-page document analysis; check LLM.API docs for the exact current token limit.
As a 235B-parameter model it has higher latency than smaller models, but LLM.API uses optimized serving to keep interactive use practical.
Pricing is usage-based per input and output token, with the exact rates published in the LLM.API pricing section for Qwen models.
Specify the model name in your LLM.API chat or completion request, pass text and optional image inputs, and handle responses like other chat models.
It generally offers stronger reasoning and visual understanding quality but at higher cost and latency than smaller Qwen3 VL variants.
It can hallucinate details, may misinterpret ambiguous images, and is not guaranteed accurate for real-time data or highly domain-specific expert advice.
By default it has no direct internet or tool access; any such capabilities must be implemented in your application around the LLM.API calls.
Compare
ERNIE 4.5 21B A3B Thinking is Baidu’s upgraded lightweight MoE language model optimized for deep reasoning, with a context window around 131K tokens and competitive pricing…
Claude Opus 4.8 (Fast) is Anthropic’s flagship Claude Opus 4.8 model running in a special fast mode that delivers significantly higher output token throughput at premium…
Ling-2.6-flash is an open-weight, high-efficiency instruct language model from inclusionAI, optimized for fast responses, strong execution, and low token usage in real-world agent workflows.