- Instruction Following
GLM 4.7 Flash is a 30B-class Mixture-of-Experts language model from Z.ai, optimized for speed and efficiency while maintaining strong performance on coding and agentic reasoning tasks.
Powered by NVIDIA
Nemotron 3 Nano Omni (free) is NVIDIA’s open multimodal large language model that unifies understanding of video, audio, images, documents, GUIs, and text in a single MoE architecture. It is optimized to act as a high-throughput, low-latency perception and reasoning sub-agent for agentic AI workflows.
Output tokens per second · Higher is better
Seconds · Lower is better
USD per 1M tokens (blended) · Lower is better
About the model
Nemotron 3 Nano Omni (free) is an open-weight, ~30B-parameter hybrid mixture-of-experts multimodal model from NVIDIA that processes video, audio, images, documents, charts, GUIs, and text with around 3B active parameters per token. It is mainly used to power agentic AI systems that need unified perception and reasoning over long-context multimodal inputs such as document intelligence, video understanding, and audio or screen-based Q&A. It also supports enterprise workflows like summarization, transcription, and multimodal question answering with up to 9x higher throughput than comparable open omni models at similar interactivity levels. It belongs to NVIDIA’s Nemotron 3 family and succeeds earlier Nemotron Nano multimodal models such as Nemotron Nano V2 VL within the broader Nemotron multimodal series.
Model capabilities
Engages in multi-turn text conversations, answering questions, following instructions, and maintaining context across user interactions.
Helps with programming tasks by explaining code, suggesting snippets, and assisting with debugging for common languages and frameworks.
Translates between multiple natural languages, preserving core meaning and providing reasonably fluent outputs for everyday text.
Analyzes input images to identify objects and describe visible content, enabling basic visual understanding in context.
Reads text content from images or screenshots, enabling basic optical character recognition for further processing or understanding.
Use cases
Transparent pricing
LLM API offers the lowest cost and highest performance for Nemotron 3 Nano-class models.
| Provider | Region | Latency | Throughput | Uptime | Input ($/1M) | Output ($/1M) | Context |
|---|---|---|---|---|---|---|---|
| LLM API BEST | Global | 80ms | 120 tps | 99.99% | $0.02 | $0.02 | 128K |
| NVIDIA | US West | ~140ms | ~45 tps | ~99.9% | $0.00 | $0.00 | ~32K |
| AWS Bedrock | US East | ~160ms | ~40 tps | 99.9% | ~$0.08 | ~$0.08 | ~32K |
| Azure AI | EU West | ~170ms | ~35 tps | 99.9% | ~$0.09 | ~$0.09 | ~32K |
| Google Cloud | Global | ~150ms | ~50 tps | ~99.9% | ~$0.07 | ~$0.07 | ~64K |
Performance benchmarks
| Metric | Nemotron 3 Nano Omni (free) | GPT-4o mini (OpenAI) | Gemini 1.5 Flash (Google) |
|---|---|---|---|
| Avg Latency | ~180ms | ~220ms | ~250ms |
| Context Window | 128K | 128K | 1M |
| Input Price ($/1M) | $0.00 | $0.15 | $0.08 |
| Output Price ($/1M) | $0.00 | $0.60 | $0.30 |
| Max Output Tokens | 4K | 4K | 8K |
| Throughput | ~80 tps | ~60 tps | ~70 tps |
| Uptime | 99.5% | 99.9% | 99.9% |
30-day usage via LLM API
Architecture & Integration
One unified API. Every major model. Built-in reliability, cost control, and observability.
Dynamically route each request to the optimal model across providers based on latency, cost, and performance—without changing your integration or redeploying code.
One endpoint, best modelAutomatically blend premium and budget models using your rules and budgets, so you cut AI spend without manually rewriting prompts or switching providers.
Control spend, not qualityRecover gracefully from provider outages, timeouts, or rate limits with configurable fallback rules that keep your AI features online and your SLAs intact.
Stay online by defaultTrace every request across models and providers with logs, metrics, and replayable sessions so you can debug regressions and optimize prompts in production.
See every token, everywhereDescribe tasks like chat, extraction, or scoring once and let LLM.API choose the right model and parameters, simplifying complex workflows into a clean API.
Think tasks, not modelsProcess massive workloads efficiently with parallelized, rate-limit-aware batching that maximizes throughput while staying within provider quotas and cost targets.
Scale jobs, not painDecision guide
FAQ
Nemotron 3 Nano Omni (free) is an NVIDIA language model accessible via LLM.API, optimized for lightweight, general-purpose text generation and assistance.
Nemotron 3 Nano Omni (free) is best for fast, low-cost text generation, code assistance, and lightweight reasoning where ultra-low latency matters more than raw capability.
Nemotron 3 Nano Omni (free) is offered with zero per-token charges on LLM.API, subject to platform-level free-tier quotas and rate limits.
Nemotron 3 Nano Omni (free) supports a 4K-token context window, suitable for short conversations, prompts, and small documents.
Nemotron 3 Nano Omni (free) is optimized for very low latency and high throughput, making it well-suited for real-time and interactive applications.
Nemotron 3 Nano Omni (free) is a text-only model, accepting text prompts and returning text completions without native image or audio support.
You call the unified LLM.API completion or chat endpoint, specifying the NVIDIA provider and Nemotron 3 Nano Omni (free) as the model identifier.
Nemotron 3 Nano Omni (free) is smaller and cheaper, trading off complex reasoning and long-context performance for lower latency and resource usage.
Nemotron 3 Nano Omni (free) may hallucinate, struggle with very long or complex tasks, and is not suitable for mission-critical or highly factual applications.
Yes, it is well-suited to batch and high-volume workloads, but throughput is governed by LLM.API’s global quotas and rate limits for free models.
Compare
GLM 4.7 Flash is a 30B-class Mixture-of-Experts language model from Z.ai, optimized for speed and efficiency while maintaining strong performance on coding and agentic reasoning tasks.
Qwen3.5 Plus 2026-04-20 is a large-scale, proprietary multimodal language model from Qwen (Alibaba) that offers a 1M-token context window and strong reasoning and vision capabilities for…
Step 3.5 Flash is StepFun’s sparse Mixture-of-Experts language model that delivers frontier-level reasoning and agentic capabilities while remaining highly efficient and fast for production use.