- Text Generation
Kimi K2.6 (free) is MoonshotAI’s open-source, multimodal Mixture-of-Experts model optimized for long-horizon coding, autonomous agents, and large-context reasoning. The free variant provides access to these capabilities…
Powered by NVIDIA
Nemotron 3 Super is NVIDIA’s open-weight, 120B-parameter hybrid Mamba-Transformer Mixture-of-Experts language model optimized for high-throughput agentic reasoning workloads. It is notable for combining LatentMoE experts, long-context support, and efficient NVFP4 training to deliver competitive accuracy with substantially higher inference efficiency than comparable open models.
Output tokens per second · Higher is better
Seconds · Lower is better
USD per 1M tokens (blended) · Lower is better
About the model
Nemotron 3 Super is a 120B-parameter (12B active) open Mixture-of-Experts hybrid Mamba-Attention large language model from NVIDIA, designed for efficient, high-quality agentic reasoning. It is primarily used for building autonomous AI agents that perform multi-step reasoning, tool use, and long-running workflows in domains like software engineering, data analysis, and complex enterprise automation. It is also used as a foundation text model for high-throughput, long-context applications such as large document understanding and large-scale code generation on NVIDIA GPU infrastructure. It is part of the Nemotron 3 family of open models, alongside smaller Nano and larger Ultra variants that share common training data, recipes, and architecture principles.
Model capabilities
Supports agent-style workflows, enabling planning, tool use, and multi-step decision-making for complex autonomous and semi-autonomous AI agents.
Acts as a large language model optimized for natural, multi-turn dialogue with strong instruction following and contextual understanding.
Processes and reasons over very long text contexts, supporting extended documents, workflows, and multi-document inputs in a single session.
Hybrid Mamba-Transformer Mixture-of-Experts architecture with multi-token prediction enables high-throughput, low-latency text generation at scale.
Handles multiple languages, enabling understanding and generation across diverse linguistic inputs for global applications and datasets.
Use cases
Transparent pricing
LLM API delivers the lowest cost and latency for Nemotron-class models versus major providers.
| Provider | Region | Latency | Throughput | Uptime | Input ($/1M) | Output ($/1M) | Context |
|---|---|---|---|---|---|---|---|
| LLM API BEST | Global | 80ms | 120 tps | 99.99% | $0.20 | $0.20 | 256K |
| NVIDIA | US West | ~160ms | ~60 tps | 99.9% | ~$0.60 | ~$0.60 | 128K |
| AWS Bedrock | US East | ~180ms | ~45 tps | 99.9% | ~$0.70 | ~$0.70 | 64K |
| Azure AI | EU West | ~190ms | ~40 tps | 99.9% | ~$0.75 | ~$0.75 | 128K |
| Google Cloud | Global | ~170ms | ~50 tps | 99.9% | ~$0.80 | ~$0.80 | 128K |
Performance benchmarks
| Metric | Nemotron 3 Super (NVIDIA) | GPT-4 Turbo (OpenAI) | Claude 3 Sonnet (Anthropic) |
|---|---|---|---|
| Avg Latency | ~350ms | ~400ms | ~450ms |
| Context Window | ~128K | 128K | 200K |
| Input Price ($/1M) | ~$0.60 | ~$0.50 | ~$0.60 |
| Output Price ($/1M) | ~$2.40 | ~$1.50 | ~$1.80 |
| Max Output Tokens | ~4K | 4K | 4K |
| Throughput | ~60 tps | ~50 tps | ~40 tps |
| Uptime | ~99.9% | ~99.9% | ~99.9% |
30-day usage via LLM API
Architecture & Integration
One unified API. Every major model. Built-in reliability, cost control, and observability.
Automatically route each request to the best-performing model across providers based on latency, cost, and quality—without changing your integration.
One endpoint, optimal modelControl spend with transparent per-token accounting, guardrails, and smart selection of cheaper equivalent models while preserving quality for critical workloads.
Optimize every tokenKeep production flows resilient with built-in provider failover and model-level retries, so transient outages never break your user experience.
Resilient by defaultInspect latency, errors, tokens, and prompts across providers in one place, enabling fast debugging, regression detection, and performance tuning.
See every requestExpress complex AI workflows as high-level tasks—grounding, tools, classification, generation—while LLM.API handles prompt shaping, execution, and model differences.
Ship workflows, not glueProcess millions of requests efficiently with batch APIs that parallelize across providers, maximize throughput, and minimize cost for bulk inference jobs.
Scale jobs to millionsDecision guide
FAQ
Nemotron 3 Super is an NVIDIA large language model focused on high‑quality text generation and reasoning, accessible through the LLM.API unified gateway.
Nemotron 3 Super is best for code generation, data analysis assistance, structured tool-calling workflows, and general-purpose chatbots needing strong reasoning and instruction-following.
Nemotron 3 Super supports a context window of up to 8,192 tokens for combined input and output via LLM.API.
Nemotron 3 Super currently supports text-in, text-out interactions only; image, audio, and video inputs are not supported.
Nemotron 3 Super typically returns first tokens within a few hundred milliseconds and can stream responses for lower perceived latency.
Nemotron 3 Super is billed per 1,000 tokens, with separate rates for input and output tokens as defined in your LLM.API pricing plan.
You select provider "NVIDIA" and model "Nemotron 3 Super" in the LLM.API request payload, then send standard chat or completion-style requests.
Nemotron 3 Super targets stronger reasoning and coding performance than smaller Nemotron variants, at higher compute cost but improved quality.
Nemotron 3 Super can hallucinate facts, lacks real-time internet access, and should not be solely relied on for high-stakes or legally binding decisions.
Direct fine-tuning is not exposed via LLM.API; instead, you should use techniques like system prompts, few-shot examples, and retrieval-augmented generation.
Compare
Kimi K2.6 (free) is MoonshotAI’s open-source, multimodal Mixture-of-Experts model optimized for long-horizon coding, autonomous agents, and large-context reasoning. The free variant provides access to these capabilities…
MiniMax M2.7 is a 230B-parameter Mixture-of-Experts large language model from MiniMax, with 10B active parameters and a 204,800-token context window, optimized for coding, agentic tool use,…
GPT-5.4 is an OpenAI language model, but as of now OpenAI has not publicly released technical details or documentation about this specific version, so only its…