- Text Generation
GPT-4o Transcribe is an OpenAI model specialized for converting audio into accurate, time-aligned text transcripts. It is notable for handling natural speech, varied accents, and real-world…
Powered by NVIDIA
Nemotron 3 Super (free) is NVIDIA’s open‑weights, high‑throughput 120B-parameter hybrid mixture‑of‑experts language model, optimized for complex agentic AI and multi‑agent reasoning workloads. It is notable for combining a hybrid Mamba‑Transformer architecture, LatentMoE sparsity, and a 1M‑token context window to deliver efficient long‑horizon reasoning.
Output tokens per second · Higher is better
Seconds · Lower is better
USD per 1M tokens (blended) · Lower is better
About the model
Nemotron 3 Super is an open, 120B-parameter hybrid Mamba-Transformer mixture-of-experts model from NVIDIA designed for high-accuracy, efficient agentic reasoning. It is mainly used to power multi-agent and enterprise AI workflows that require long-context reasoning, planning, and orchestration across many tools or services. It is also well-suited for code, math, and complex multistep generation tasks where high throughput and long sequences are important. It belongs to the Nemotron 3 family of open models (Nano, Super, Ultra), succeeding earlier Nemotron generations.
Model capabilities
Supports multi‑agent, tool-using AI workflows, coordinating complex tasks with high throughput and long-horizon reasoning across agents.
Handles sequences up to around one million tokens, enabling analysis of large documents, codebases, and extended conversations without losing context.
Generates and understands text in multiple languages, including English and Japanese, for global applications and cross-lingual workflows.
Engages in open-domain dialogue, following instructions, answering questions, and assisting with writing or brainstorming in natural language.
Trained on diverse web, code, and technical data, enabling structured outputs, explanations, and reasoning over text-based information sources.
Use cases
Transparent pricing
LLM API offers the lowest cost and highest performance for Nemotron-class models.
| Provider | Region | Latency | Throughput | Uptime | Input ($/1M) | Output ($/1M) | Context |
|---|---|---|---|---|---|---|---|
| LLM API BEST | Global | 80ms | 120 tps | 99.99% | $0.10 | $0.10 | 128K tokens |
| NVIDIA | Global | ~180ms | ~60 tps | ~99.9% | $0.00 | $0.00 | ~128K tokens |
| AWS Bedrock | US East | ~220ms | ~45 tps | ~99.9% | ~$0.25 | ~$0.25 | ~128K tokens |
| Google Cloud | Global | ~210ms | ~50 tps | ~99.9% | ~$0.24 | ~$0.24 | ~128K tokens |
| Azure AI | EU West | ~230ms | ~40 tps | ~99.9% | ~$0.26 | ~$0.26 | ~128K tokens |
Performance benchmarks
| Metric | Nemotron 3 Super (free) | Llama 3 8B Instruct (free) | Mistral 7B Instruct (free) |
|---|---|---|---|
| Avg Latency | ~800ms | ~900ms | ~850ms |
| Context Window | 8K | 8K | 8K |
| Input Price ($/1M) | $0.00 | $0.00 | $0.00 |
| Output Price ($/1M) | $0.00 | $0.00 | $0.00 |
| Max Output Tokens | 2K | 2K | 2K |
| Throughput | ~30 tps | ~25 tps | ~25 tps |
| Uptime | 99.9% | 99.9% | 99.9% |
30-day usage via LLM API
Architecture & Integration
One unified API. Every major model. Built-in reliability, cost control, and observability.
Intelligently route each request across providers and models based on latency, capability, or custom rules. One API, always the best path for your workload.
Smart traffic, single endpointAutomatically balance performance and price with configurable policies. Use premium models when it matters, fall back to cheaper ones when it doesn’t.
Optimize spend by defaultDefine multi-step failover chains across providers so requests keep flowing through outages, rate limits, or model errors—without touching your application code.
Stay online under stressGet full visibility into requests, tokens, latency, errors, and providers with structured logs and traces. Debug faster and tune workloads with real data.
See every token spentDescribe tasks like chat, tools, reranking, or extraction once and run them on any model. Ship features without rewriting prompts per provider.
Code to tasks, not modelsSubmit massive batch jobs through a single API with queuing, retries, and cost controls built-in. Process millions of inputs without custom infrastructure.
Scale jobs, not opsDecision guide
FAQ
Nemotron 3 Super (free) is an NVIDIA large language model accessible via LLM.API, tuned for general-purpose text generation and assistant-style conversations.
It is best for fast, low-cost chat-style interactions, drafting content, and lightweight reasoning where cost and accessibility matter more than cutting-edge intelligence.
The free tier incurs no direct per-token charges to you, but may be subject to rate limits and usage caps enforced by LLM.API.
Nemotron 3 Super (free) supports a context window of up to 8K tokens, including both prompt and response tokens.
Latency is typically low for short prompts, but can increase under heavy shared-load conditions because the free tier runs on pooled infrastructure.
Nemotron 3 Super (free) supports text-in, text-out interactions only and does not natively process images, audio, or video.
You select the model by its identifier in the LLM.API completion or chat endpoint, passing your prompt and standard configuration parameters like temperature.
Compared to larger or paid frontier models, it is generally cheaper and more accessible but weaker on complex reasoning, coding, and long-context tasks.
It may hallucinate facts, struggle with very long or deeply technical tasks, and lacks multimodal capabilities and fine-grained enterprise controls.
You can, but should account for potential rate limits, variable performance, and weaker reliability than dedicated, paid production-grade NVIDIA deployments.
Compare
GPT-4o Transcribe is an OpenAI model specialized for converting audio into accurate, time-aligned text transcripts. It is notable for handling natural speech, varied accents, and real-world…
Qwen3.5-9B is a 9‑billion‑parameter multimodal language model from Qwen that supports long-context reasoning over text and images. It is designed to offer strong reasoning, coding, and…
GLM 4.7 is Z.ai’s flagship large language model, optimized for strong coding performance and stable multi-step reasoning. It is notable for its very large context window…