- Text Generation
Gemma 4 26B A4B is a 26-billion-parameter multimodal Mixture-of-Experts model from Google’s Gemma 4 family, optimized for high-throughput reasoning with long context windows. It supports text…
Powered by inclusionAI
Ling-2.6-1T is inclusionAI’s trillion-parameter flagship instruction model optimized for fast, efficient execution in real-world agentic, coding, and complex reasoning workflows.
Output tokens per second · Higher is better
Seconds · Lower is better
USD per 1M tokens (blended) · Lower is better
About the model
Ling-2.6-1T is a 1-trillion-parameter flagship language model from inclusionAI designed as a high-efficiency instant/instruct model for complex real-world tasks. It is mainly used for advanced coding, large-scale agent workflows, and long-context applications that require both strong reasoning and high throughput. It is also used for everyday language tasks such as writing, summarization, and explanation where low latency and tool use/structured outputs are important. Ling-2.6-1T belongs to the Ling 2.6 family of open-weight models, alongside variants like Ling-2.6-Flash and the reasoning-focused Ring-2.6-1T.
Model capabilities
Engages in multi-turn, context-aware chat, answering questions, following instructions, and maintaining coherent dialogue across various topics.
Translates text between multiple languages, preserving meaning and tone for general-purpose content and everyday communication.
Understands and summarizes written content, extracting key points, intent, and sentiment from diverse text sources.
Analyzes images to recognize objects, people, and scenes, generating concise descriptions of visual content.
Extracts machine-readable text from scanned documents and photos of text, enabling downstream search, editing, and analysis.
Use cases
Transparent pricing
LLM API offers the lowest cost and highest performance for Ling-2.6-1T–class models.
| Provider | Region | Latency | Throughput | Uptime | Input ($/1M) | Output ($/1M) | Context |
|---|---|---|---|---|---|---|---|
| LLM API BEST | Global | 80ms | 120 tps | 99.99% | $0.30 | $0.60 | 256K |
| inclusionAI | US East | ~140ms | ~70 tps | ~99.9% | ~$0.40 | ~$0.80 | ~128K |
| OpenAI | Global | ~150ms | ~80 tps | 99.9% | ~$0.50 | ~$1.00 | 128K |
| Anthropic | US West | ~160ms | ~60 tps | ~99.9% | ~$0.55 | ~$1.10 | 200K |
| Google Cloud AI | Global | ~170ms | ~65 tps | 99.9% | ~$0.45 | ~$0.90 | 128K |
Performance benchmarks
| Metric | Ling-2.6-1T (inclusionAI) | GPT-4.1 (OpenAI) | Claude 3.5 Sonnet (Anthropic) |
|---|---|---|---|
| Avg Latency | ~180ms | ~220ms | ~210ms |
| Context Window | 128K | 128K | 200K |
| Input Price ($/1M tokens) | $1.20 | $5.00 | $3.00 |
| Output Price ($/1M tokens) | $3.60 | $15.00 | $15.00 |
| Max Output Tokens | 8K | 4K | 4K |
| Throughput | 60 tps | 4K | 4K |
| Throughput | ~80 tps | ~60 tps | ~50 tps |
| Uptime | 99.9% | 99.9% | 99.9% |
30-day usage via LLM API
Architecture & Integration
One unified API. Every major model. Built-in reliability, cost control, and observability.
Automatically send each request to the best-fit model across providers based on latency, cost, or quality—without changing your integration or redeploying code.
One API, any model.Optimize spend with policy-based routing, budget guards, and granular usage controls so you can experiment freely without surprise bills or vendor lock-in.
Max control, minimal spend.Define automatic failover and degradation paths when a provider is down, slow, or rate-limited so your production workloads stay online and predictable.
Fail gracefully, not silently.Get unified logs, traces, metrics, and structured payloads across all providers to debug prompts, compare models, and tune performance from one place.
See every token, everywhere.Define high-level tasks like chat, embeddings, tools, or RAG once, then swap underlying models and vendors without touching application logic.
Code to tasks, not models.Run large-scale batch workloads with queueing, concurrency control, and automatic retries so you can process millions of tasks reliably and cost-efficiently.
From prototype to millions.Decision guide
FAQ
Ling-2.6-1T is a large language model from inclusionAI focused on high-quality text generation and reasoning, accessible through the LLM.API unified gateway.
Ling-2.6-1T is best for complex reasoning, multi-step data processing, and robust code and text generation across a wide range of developer use cases.
Ling-2.6-1T supports a context window of up to 32,000 tokens for combined input and output through LLM.API.
Ling-2.6-1T currently supports text-in, text-out interactions only when accessed through LLM.API.
Ling-2.6-1T uses a pay-per-token billing model on LLM.API, with separate input and output token rates defined in your LLM.API pricing plan.
Typical end-to-end latencies for Ling-2.6-1T are usually in the low-seconds range, depending on prompt size and concurrent load.
You specify the model name "inclusionai/ling-2.6-1T" in your LLM.API completion or chat request, plus your API key and usual parameters.
Ling-2.6-1T aims to balance strong reasoning and generation quality with more predictable costs than many similarly sized frontier models.
Ling-2.6-1T can hallucinate facts, reflect training-data biases, and should not be relied on for safety-critical or legally binding decisions.
Yes, Ling-2.6-1T supports token streaming on LLM.API when you enable the streaming option in your request parameters.
Compare
Gemma 4 26B A4B is a 26-billion-parameter multimodal Mixture-of-Experts model from Google’s Gemma 4 family, optimized for high-throughput reasoning with long context windows. It supports text…
Claude Sonnet 4.6 is Anthropic’s most capable Sonnet‑tier large language model, offering Opus‑class performance in coding, computer use, and long‑context reasoning with a 1 million token…
Claude Opus 4.5 is Anthropic’s frontier large language model optimized for advanced reasoning, coding, and long-context, agentic workflows. It is positioned as a flagship, high-intelligence model…