- Text Generation
Qwen3 VL 32B Instruct is a 32-billion-parameter multimodal vision-language model from Qwen, designed for high-precision understanding and reasoning over text, images, and video with a very…
Powered by Z.ai
GLM 4.6V is Z.ai’s open-source, large-scale vision-language model that supports images, video, documents, and text with a long context window and native tool use. It is notable for combining high-quality multimodal understanding with function calling and cloud- or local-friendly variants.
Output tokens per second · Higher is better
Seconds · Lower is better
USD per 1M tokens (blended) · Lower is better
About the model
GLM 4.6V is Z.ai’s 106B-parameter multimodal foundation model for visual reasoning over text, images, and video. It is mainly used for tasks like document and image understanding, code and data analysis, and agent-style workflows that rely on native function calling. It also powers applications needing long-context (around 128K–131K tokens) multimodal chat and reasoning, from research assistants to enterprise AI tools. GLM 4.6V belongs to the GLM-V family and follows earlier GLM-4.5V and GLM-4.5-Air models, alongside the smaller GLM-4.6V-Flash variant.
Model capabilities
Engages in context-aware conversations with long text and mixed media inputs using a large 128K context window.
Analyzes images, complex layouts, charts, and documents, extracting structure and semantics for downstream reasoning or generation.
Performs multi-step reasoning on text and visual inputs, supporting chain-of-thought style problem solving and complex analysis.
Reads and interprets text from screenshots, scanned documents, tables, and natural images as part of its visual understanding pipeline.
Translates between multiple languages within multimodal conversations, preserving context from accompanying images or documents.
Use cases
Transparent pricing
LLM API offers the lowest cost and latency for GLM 4.6V-class models.
| Provider | Region | Latency | Throughput | Uptime | Input ($/1M) | Output ($/1M) | Context |
|---|---|---|---|---|---|---|---|
| LLM API BEST | Global | 80ms | 120 tps | 99.99% | $0.20 | $0.40 | 200K |
| Z.ai | Global | ~220ms | ~40 tps | ~99.9% | ~$0.60 | ~$1.20 | ~128K |
| OpenAI (GPT-4.1 mini vision-equivalent) | Global | ~180ms | ~80 tps | 99.9% | ~$0.50 | ~$1.00 | 128K |
| Google (Gemini 1.5 Flash Vision-equivalent) | Global | ~190ms | ~70 tps | 99.9% | ~$0.45 | ~$0.90 | 128K |
| Anthropic (Claude 3.5 Sonnet Vision-equivalent) | Global | ~210ms | ~50 tps | 99.9% | ~$0.70 | ~$1.40 | 200K |
Performance benchmarks
| Metric | GLM 4.6V | GPT-4o | Claude 3.5 Sonnet |
|---|---|---|---|
| Latency per Image | ~220ms | ~250ms | ~260ms |
| Throughput | 45 img/s | 40 img/s | 35 img/s |
| Max Resolution | 4K | 4K | 4K |
| Price per Image | $0.003 | $0.005 | $0.004 |
| Supported Formats | PNG, JPG, WEBP | PNG, JPG, WEBP, GIF | PNG, JPG, WEBP |
| Max Output Tokens (per call) | 4K | 4K | 4K |
| Uptime | 99.9% | 99.9% | 99.9% |
30-day usage via LLM API
Architecture & Integration
One unified API. Every major model. Built-in reliability, cost control, and observability.
Define routing rules once and automatically send each request to the optimal model across providers based on latency, cost, or quality—without changing your app code.
One endpoint, any modelMix premium and budget models with granular controls, rate limits, and caps so you can aggressively optimize spend without sacrificing reliability or user experience.
Lower cost per callConfigure provider-agnostic retries and fallbacks so requests seamlessly fail over to backup models on timeouts, rate limits, or outages—no brittle error handling.
Resilience by defaultGet centralized logs, traces, and metrics for every AI call across providers, with request replay and tagging to debug issues and tune performance quickly.
See every tokenDescribe tasks like chat, extraction, or tools once and let LLM.API handle model-specific prompts, parameters, and formats behind a stable, versioned contract.
APIs, not promptsSend thousands of requests in a single batch with concurrency controls and retries, maximizing throughput while keeping provider limits and costs under control.
Scale without throttlingDecision guide
FAQ
GLM 4.6V is a multimodal Z.ai model accessible through LLM.API, designed for combined text and image understanding and generation.
GLM 4.6V is best for vision-language tasks like image captioning, visual question answering, UI understanding, and workflows mixing images with natural language.
GLM 4.6V supports text input and output plus image input, enabling rich vision-language interactions via a single API.
GLM 4.6V supports a 32K token context window for prompts and conversation history combined.
GLM 4.6V is optimized for low-latency responses, with typical first-token times under a second for short prompts, excluding network overhead.
LLM.API charges for GLM 4.6V on a pay-per-token basis for prompt and completion tokens, following the Z.ai GLM 4.6V pricing tier.
Use the unified LLM.API chat or completions endpoint and set the model parameter to the GLM 4.6V identifier provided in the dashboard.
GLM 4.6V targets strong vision-language quality with competitive cost, generally trading slightly lower raw performance for better efficiency than frontier multimodal models.
GLM 4.6V can hallucinate facts, misread small or low-resolution visual details, and should not be used for safety-critical or legal decisions.
Yes, GLM 4.6V supports server-sent events streaming on LLM.API, allowing tokens to be consumed as they are generated.
Compare
Qwen3 VL 32B Instruct is a 32-billion-parameter multimodal vision-language model from Qwen, designed for high-precision understanding and reasoning over text, images, and video with a very…
MiniMax M2.7 is a 230B-parameter Mixture-of-Experts large language model from MiniMax, with 10B active parameters and a 204,800-token context window, optimized for coding, agentic tool use,…
MiniMax M2.1 is a second-generation, open-weight Mixture-of-Experts large language model from MiniMax, optimized for real-world coding, tool use, and long-horizon agentic workflows. It is notable for…