- Text Generation
Riverflow V2 Fast is the fastest variant of Sourceful’s Riverflow 2.0 image generation and editing lineup, optimized for production deployments and latency‑critical brand creative workflows.
Powered by Z.ai
GLM 4.6 is Z.ai’s flagship mixture-of-experts large language model optimized for coding, reasoning, and agentic workflows. It is notable for its strong performance on code benchmarks and its very long ~200K token context window for complex tasks.
Output tokens per second · Higher is better
Seconds · Lower is better
USD per 1M tokens (blended) · Lower is better
About the model
GLM 4.6 is a mixture-of-experts large language model from Z.ai designed for advanced coding assistance, reasoning, and agent-style tool use. It is primarily used for software development workflows, including code generation, refactoring, and working over large repositories within integrated agents. It is also applied to general-purpose text generation and long-context tasks such as multi-step reasoning, data analysis, and orchestrated tool-calling pipelines. GLM 4.6 succeeds earlier GLM 4.x models such as GLM 4.5 in the broader GLM series developed by Zhipu AI (Z.ai).
Model capabilities
Performs complex logical reasoning and multi-step problem solving, supporting tool use for sophisticated agentic workflows and decision-making tasks.
Generates, analyzes, and debugs code across multiple languages, optimized for building coding agents and long-running software development workflows.
Processes and utilizes very long text contexts, enabling work with large documents, extended conversations, and multi-stage project instructions.
Understands and generates text in multiple languages for general-purpose chat, knowledge querying, and content creation across diverse domains.
Extracts, interprets, and restructures information from long-form text documents, supporting summarization, reformatting, and targeted information retrieval.
Use cases
Transparent pricing
LLM API offers the lowest cost and highest performance for GLM 4.6–class models.
| Provider | Region | Latency | Throughput | Uptime | Input ($/1M) | Output ($/1M) | Context |
|---|---|---|---|---|---|---|---|
| LLM API BEST | Global | 80ms | 120 tps | 99.99% | $0.15 | $0.45 | 256K |
| Z.ai | Global | ~180ms | ~40 tps | ~99.9% | ~$0.60 | ~$1.80 | ~128K |
| OpenAI (closest: GPT-4.1) | Global | ~220ms | ~30 tps | 99.9% | ~$2.50 | ~$10.00 | 128K |
| Anthropic (closest: Claude 3.5 Sonnet) | US & EU | ~210ms | ~25 tps | 99.9% | ~$3.00 | ~$15.00 | 200K |
| Google (closest: Gemini 1.5 Pro) | Global | ~240ms | ~20 tps | ~99.9% | ~$2.00 | ~$8.00 | 1M |
Performance benchmarks
| Metric | GLM 4.6 (Z.ai) | GPT-4.1 (OpenAI) | Claude 3.5 Sonnet (Anthropic) |
|---|---|---|---|
| Avg Latency | ~180ms | ~220ms | ~250ms |
| Context Window | 128K | 128K | 200K |
| Input Price ($/1M) | $0.70 | $5.00 | $3.00 |
| Output Price ($/1M) | $2.10 | $15.00 | $15.00 |
| Max Output Tokens | 4K | 4K | 4K |
| Throughput | 70 tps | 50 tps | 45 tps |
| Uptime | 99.9% | 99.9% | 99.9% |
30-day usage via LLM API
Architecture & Integration
One unified API. Every major model. Built-in reliability, cost control, and observability.
Automatically route each request to the optimal model by provider, latency, and capability so you ship faster without hard-coding vendor logic.
One endpoint, any modelBalance quality and price with per-request cost controls, policies, and mix-and-match providers so you never overspend on routine workloads.
More performance, less spendDefine multi-provider fallback chains so timeouts, rate limits, or provider outages fail over automatically—keeping your AI features online.
Designed for failure modesGet unified logs, traces, metrics, and per-provider analytics across all AI calls so you can debug prompts, tune latency, and track usage in one place.
One pane of glassDefine reusable tasks like chat, extraction, search, or tools once, then swap models underneath without touching application code.
Program tasks, not vendorsSend large batches of requests through a single pipeline with built-in retries, concurrency controls, and aggregation for massive throughput and lower effective cost.
Scale to millions of callsDecision guide
FAQ
GLM 4.6 is a large language model from Z.ai focused on fast, general-purpose text generation and reasoning through the LLM.API gateway.
GLM 4.6 is best for chatbots, code assistance, document summarization, and general reasoning where balanced quality and speed are important.
Through LLM.API, GLM 4.6 currently supports text input and text output only.
GLM 4.6 supports up to a 128K token context window for prompts plus generated output combined.
GLM 4.6 is optimized for low initial latency and high token throughput, making it suitable for interactive applications and batched backend workloads.
LLM.API exposes GLM 4.6 with token-based pricing; see the LLM.API pricing page for current per‑million input and output token rates.
Specify the provider as "Z.ai" and the model name "GLM 4.6" in your LLM.API completion or chat endpoint request payload.
Compared to similar general-purpose models, GLM 4.6 targets a balance of competitive reasoning quality, longer context, and cost efficiency.
If enabled by LLM.API, GLM 4.6 can consume structured function schemas and produce arguments for tool invocation like other compatible models.
GLM 4.6 can hallucinate facts, lacks real-time knowledge, and should not be used without human review for safety-critical or compliance-sensitive decisions.
Compare
Riverflow V2 Fast is the fastest variant of Sourceful’s Riverflow 2.0 image generation and editing lineup, optimized for production deployments and latency‑critical brand creative workflows.
Nemotron 3 Nano 30B A3B is a 30-billion-parameter NVIDIA language model variant optimized for compact deployment with efficient inference. It targets on-device or resource-constrained environments while…
Anthropic Claude Haiku (Latest) is a lightweight, fast Claude family model optimized for low-latency, cost‑efficient tasks while maintaining strong language understanding. It is notable for offering…