- Text Generation
GPT-5.4 Mini is an OpenAI language model variant optimized for lightweight, general-purpose assistant tasks. It is designed to balance capability with efficiency for everyday conversational and…
Powered by NVIDIA
Llama Nemotron Embed VL 1B V2 (free) is NVIDIA’s 1B-parameter multimodal embedding model optimized for question-answering retrieval over text and visual document data. It produces dense vector embeddings from text, images, or combined image–text inputs for high-quality semantic search and RAG systems.
Output tokens per second · Higher is better
Seconds · Lower is better
USD per 1M tokens (blended) · Lower is better
About the model
Llama Nemotron Embed VL 1B V2 (free) is a combined language–vision embedding model from NVIDIA designed for multimodal question-answering retrieval over text and document images. It is mainly used to embed large corpora of documents (including pages with text, tables, charts, and infographics) into dense vectors for semantic retrieval, enterprise search, and knowledge indexing. It is also used to power RAG pipelines that retrieve relevant visual or textual context given a text query, supporting text, image, and text+image to embedding modalities with a large context window. It belongs to NVIDIA’s Nemotron RAG collection and Llama Nemotron embedding family, and is offered as a free variant via providers like OpenRouter and Remova.
Model capabilities
Generates dense vector embeddings from text, images, or combined image-text document pages for retrieval over multimodal corpora.
Embeds textual queries and passages so semantically related documents can be efficiently retrieved using vector similarity search.
Encodes page images containing text, tables, charts, and infographics to enable semantic search over scanned or PDF documents.
Optimized to embed user questions and relevant pages so answer-containing documents are ranked highly in retrieval pipelines.
Provides multilingual text embeddings, enabling cross-language retrieval where queries and documents may be written in different languages.
Use cases
Transparent pricing
LLM API offers the lowest cost and highest performance for Llama Nemotron–class vision-language embeddings.
| Provider | Region | Latency | Throughput | Uptime | Input ($/1M) | Output ($/1M) | Context |
|---|---|---|---|---|---|---|---|
| LLM API BEST | Global | 50ms | 120 img/s | 99.99% | $0.00 | $0.00 | 4096 tokens |
| NVIDIA | US West | ~120ms | ~40 img/s | ~99.9% | $0.00 | $0.00 | ~4096 tokens |
| AWS Bedrock | US East | ~160ms | ~30 img/s | 99.9% | ~$0.60 / 1M tokens | ~$0.60 / 1M tokens | ~4096 tokens |
| Azure AI | EU West | ~170ms | ~25 img/s | 99.9% | ~$0.70 / 1M tokens | ~$0.70 / 1M tokens | ~4096 tokens |
| Replicate | Global | ~200ms | ~20 img/s | ~99.5% | ~$1.20 / 1M tokens | ~$1.20 / 1M tokens | ~4096 tokens |
Performance benchmarks
| Metric | Llama Nemotron Embed VL 1B V2 (free) | OpenAI text-embedding-3-small | Cohere Embed v3 English |
|---|---|---|---|
| Dimensions | 1024 | 1536 | 1024 |
| Max Input Tokens | ~8K | 8192 | ~8K |
| Price per 1M Tokens | $0.00 | $0.02 | $0.10 |
| Throughput | ~5K tok/s | ~10K tok/s | ~7K tok/s |
| Avg Latency | ~120ms | ~100ms | ~140ms |
| Uptime | ~99.5% | ~99.9% | ~99.9% |
30-day usage via LLM API
Architecture & Integration
One unified API. Every major model. Built-in reliability, cost control, and observability.
Dynamically route each request to the best model across providers based on latency, cost, and quality—no client changes, just smarter infrastructure.
One endpoint, every modelOptimize spend by mixing premium and budget models behind a single API, with pricing controls and per-route policies baked into your architecture.
Cut costs, keep qualityAutomatic failover to backup models and regions when a provider degrades, keeping your AI features reliable without extra retry logic in your code.
Stay online under failureTrace every call across providers with logs, metrics, and structured events so you can debug latency, failures, and quality from one place.
See every token hopDescribe what you want—chat, tools, RAG, workflows—once, and let LLM.API map tasks to the right models and parameters automatically.
Think tasks, not modelsSubmit massive batches across providers with built-in queuing, parallelization, and retry semantics, instead of building and tuning your own job runner.
Millions of calls, one APIDecision guide
FAQ
Llama Nemotron Embed VL 1B V2 (free) is an NVIDIA vision-language embedding model that generates joint vector representations for text and images.
It is best for semantic search, multimodal retrieval, clustering, and recommendation systems that require aligned embeddings of text and visual content.
The Llama Nemotron Embed VL 1B V2 (free) tier is available at zero API usage cost on LLM.API, subject to platform-wide rate limits.
It supports multimodal input, allowing you to encode text-only, image-only, or combined image-plus-text into a single embedding space.
Llama Nemotron Embed VL 1B V2 (free) supports text inputs up to 8,192 tokens per request via LLM.API.
As a compact 1B-parameter model, it is optimized for low latency embedding generation, typically returning results in tens of milliseconds per request.
Specify the model name "nvidia/llama-nemotron-embed-vl-1b-v2-free" in your LLM.API request along with your text and image payloads.
Compared to larger multimodal embedders, it generally offers lower latency and cost with slightly lower embedding quality on complex, fine-grained tasks.
No, it is an embedding model designed solely to produce vector representations, not to generate or continue natural language text.
It may struggle with very long documents, highly specialized domains, or detailed image reasoning compared to larger, domain-tuned multimodal models.
Compare
GPT-5.4 Mini is an OpenAI language model variant optimized for lightweight, general-purpose assistant tasks. It is designed to balance capability with efficiency for everyday conversational and…
Claude Sonnet 4.5 is an Anthropic large language model optimized for software development, computer use, and agentic workflows, offering strong performance on coding and reasoning tasks…
Ling-2.6-1T is inclusionAI’s trillion-parameter flagship instruction model optimized for fast, efficient execution in real-world agentic, coding, and complex reasoning workflows.