- Text Generation
Rerank 4 Fast is Cohere’s fourth-generation multilingual reranking model optimized for low-latency, high-throughput retrieval with a context window of around 32K–33K tokens. It is designed to…
Powered by Z.ai
GLM 4.7 is Z.ai’s flagship large language model, optimized for strong coding performance and stable multi-step reasoning. It is notable for its very large context window and open-source availability under Apache 2.0.
Output tokens per second · Higher is better
Seconds · Lower is better
USD per 1M tokens (blended) · Lower is better
About the model
GLM 4.7 is a frontier-scale large language model from Z.ai, designed as a general-purpose assistant with particular strengths in software development and complex reasoning tasks. It is widely used for code generation, debugging, and agent-style tools that require reliable multi-step execution, as well as for advanced chat, analysis, and content creation workloads across long contexts. It belongs to Z.ai’s GLM model family as a successor to earlier GLM-4.x releases and is provided as an open-source Apache 2.0 MoE-based model.
Model capabilities
Engages in multi-turn, natural conversations, following instructions and maintaining context over long text-only interactions with high coherence.
Generates, edits, and explains source code, supporting complex programming workflows and real-world development environments with strong reliability.
Produces well-formed JSON and other structured formats, supporting function calling and tool invocation for agentic applications.
Handles complex, long-horizon tasks with stable multi-step reasoning, suitable for agents and tool-using workflows.
Understands and generates text in multiple languages, enabling cross-lingual tasks, content creation, and language transformation.
Use cases
Transparent pricing
Save up to 70% on GLM‑class models with LLM API’s optimized pricing.
| Provider | Region | Latency | Throughput | Uptime | Input ($/1M) | Output ($/1M) | Context |
|---|---|---|---|---|---|---|---|
| LLM API BEST | Global | ~150ms | ~120 tps | 99.99% | $0.20 | $0.30 | 256K |
| Z.ai | Global | ~220ms | ~80 tps | ~99.9% | ~$0.35 | ~$0.60 | ~128K |
| OpenAI (closest: GPT-4.1-mini / o3-mini class) | Global | ~250ms | ~90 tps | 99.9% | ~$0.50 | ~$1.50 | 128K |
| Anthropic (closest: Claude 3.5 Sonnet) | US & EU | ~260ms | ~70 tps | ~99.9% | ~$3.00 | ~$15.00 | 200K |
| Azure AI (closest: GPT-4.1 via Azure) | US East / EU West | ~280ms | ~85 tps | 99.9% | ~$2.80 | ~$11.20 | 128K |
Performance benchmarks
| Metric | GLM 4.7 (Z.ai) | GPT-4.1 Mini (OpenAI) | Claude 3 Haiku (Anthropic) |
|---|---|---|---|
| Avg Latency | ~180ms | ~220ms | ~250ms |
| Context Window | 128K | 128K | 200K |
| Input Price ($/1M) | $0.20 | $0.15 | $0.25 |
| Output Price ($/1M) | $0.60 | $0.60 | $0.80 |
| Max Output Tokens | 8K | 4K | 8K |
| Throughput | ~80 tps | ~100 tps | ~70 tps |
| Uptime | 99.9% | 99.9% | 99.9% |
30-day usage via LLM API
Architecture & Integration
One unified API. Every major model. Built-in reliability, cost control, and observability.
Intelligently route each request across models and providers based on latency, cost, and quality. One API, dynamic routing rules, no client rewrites.
One endpoint, any modelEnforce per-project and per-model budgets, caps, and policies at the gateway. Optimize spend automatically without touching your application logic.
Control cost per callAutomatically fail over to backup models or providers on errors, timeouts, or rate limits. Keep your AI features online without custom retry logic.
No single point of failureTrace every request across models with metrics, logs, and timelines. Debug latency, failures, and behavior from a single observability layer.
See every token hopDefine higher-level tasks—like classification or extraction—then plug in any compatible model. Swap models without rewriting business logic.
Code to tasks, not modelsSend massive batches through one optimized pipeline with concurrency, backoff, and retries built in. Maximize throughput while respecting provider limits.
Millions of calls, one pipeDecision guide
FAQ
GLM 4.7 is a large language model from Z.ai focused on strong general-purpose reasoning and code generation, accessible via the LLM.API gateway.
GLM 4.7 is best for multi-step reasoning, code assistance, and building chat-style assistants that require stable, predictable behavior.
GLM 4.7 usage is billed per token through LLM.API, with exact input and output token rates defined in your LLM.API pricing plan.
GLM 4.7 supports a large context window suitable for multi-turn chats and long documents; check LLM.API docs for the exact token limit.
Typical latencies are comparable to other mid-to-large LLMs, with streaming responses available to reduce perceived delay for end users.
Through LLM.API, GLM 4.7 supports text input and output; additional modalities depend on the capabilities LLM.API exposes for this model.
Use the LLM.API chat or completion endpoint with the model identifier for GLM 4.7, including your API key and standard request parameters.
GLM 4.7 targets a balance of quality, speed, and cost comparable to mainstream general-purpose LLMs in its size and capability class.
GLM 4.7 can hallucinate facts, lacks real-time knowledge or browsing, and should not be used without human review for high-stakes decisions.
Function calling and tool-use support depend on LLM.API’s integration for GLM 4.7; check the model features table in the LLM.API documentation.
Compare
Rerank 4 Fast is Cohere’s fourth-generation multilingual reranking model optimized for low-latency, high-throughput retrieval with a context window of around 32K–33K tokens. It is designed to…
GPT-5.4 Image 2 is an OpenAI multimodal model that can understand and generate both text and images. It is notable for combining advanced language capabilities with…
Veo 3.1 is Google’s latest high-fidelity video generation model that creates short, cinematic clips from text or image prompts with native audio. It focuses on strong…