- Text Generation
Qwen3 ASR Flash is Qwen’s high-accuracy, multilingual automatic speech recognition (ASR) service optimized for real-time transcription of short audio. It is built on the Qwen3-Omni foundation…
Powered by Intfloat
Multilingual-E5-Large by Intfloat is a large multilingual text-embedding model that maps text from 90+ languages into a shared dense vector space for semantic similarity and retrieval tasks.
Output tokens per second · Higher is better
Seconds · Lower is better
USD per 1M tokens (blended) · Lower is better
About the model
Multilingual-E5-Large is a sentence- and document-level text embedding model that encodes multilingual inputs into 1024-dimensional vectors optimized for semantic similarity and retrieval across over 90 languages. It is mainly used for applications such as semantic search, multilingual similarity search, and cross-lingual information retrieval in RAG and search systems. It is also applied to clustering, classification, and other downstream NLP tasks that rely on high-quality multilingual embeddings. The model is part of the E5 family of embedding models from Intfloat, which includes small and base multilingual variants as well as English-only E5 models.
Model capabilities
Generates dense text embeddings for over 90 languages, enabling unified semantic representations across diverse multilingual content and applications.
Encodes sentences and documents so similar meanings are close in vector space, supporting clustering, deduplication, and semantic grouping tasks.
Optimized for retrieval tasks where user queries and passages are embedded and compared to power high-quality search and RAG pipelines.
Supports searching documents in one language using queries in another by mapping all texts into a shared multilingual embedding space.
Provides versatile text feature vectors usable in downstream models for tasks like classification, ranking, recommendation, and anomaly detection.
Use cases
Transparent pricing
LLM API offers the lowest embedding prices and best performance for Multilingual-E5-Large–class models.
| Provider | Region | Latency | Throughput | Uptime | Input ($/1M) | Output ($/1M) | Context |
|---|---|---|---|---|---|---|---|
| LLM API BEST | Global | ~120ms | ~8k tps | 99.99% | ~$0.03 per 1M tokens | $0.00 | ~64K tokens |
| Intfloat (Direct / HF Inference) | Global | ~220ms | ~3k tps | ~99.5% | ~$0.08 per 1M tokens | $0.00 | ~32K tokens |
| OpenAI (text-embedding-3-large) | Global | ~180ms | ~5k tps | ~99.9% | ~$0.13 per 1M tokens | $0.00 | 8K tokens |
| Azure OpenAI (Embeddings) | US East | ~200ms | ~4k tps | 99.9% | ~$0.15 per 1M tokens | $0.00 | 8K tokens |
| Together AI (Similar Embedding Model) | US West | ~210ms | ~3.5k tps | ~99.5% | ~$0.10 per 1M tokens | $0.00 | ~32K tokens |
Performance benchmarks
| Metric | Multilingual-E5-Large | text-embedding-3-large (OpenAI) | gte-large (Alibaba-NLP) |
|---|---|---|---|
| Dimensions | 1024 | 3072 | 1024 |
| Max Input Tokens | 4K | 8K | 2K |
| Price per 1M Tokens | $0.15 | $0.13 | $0.05 |
| Avg Latency | ~120ms | ~140ms | ~110ms |
| Throughput | 900 tps | 850 tps | 950 tps |
| Languages Supported | 100+ | 90+ | 80+ |
| Uptime | 99.5% | 99.9% | 99.0% |
30-day usage via LLM API
Architecture & Integration
One unified API. Every major model. Built-in reliability, cost control, and observability.
Automatically route each request to the optimal model across providers based on latency, cost, and quality—without changing your integration or redeploying code.
One endpoint, every modelControl spend with per-route budgets, model-level pricing rules, and smart downgrades that keep quality high while preventing surprise bills in production.
Optimize quality per dollarDefine provider-agnostic failover chains so requests transparently retry on backup models when providers throttle, fail, or degrade—no custom error handling needed.
Stay online, even upstreamGet full visibility into every call with traces, metrics, logs, and payload sampling to debug latency, errors, and cost across all providers in one place.
See every token, everywhereDescribe tasks—chat, RAG, tools, classification—once and let LLM.API choose and tune the right models and parameters per use case and environment.
Code to tasks, not modelsShip large jobs efficiently with batched and asynchronous execution, automatic chunking, and concurrency controls that maximize throughput while respecting provider limits.
Scale jobs, not headachesDecision guide
FAQ
Multilingual-E5-Large is an Intfloat text-embedding model optimized for multilingual semantic search, clustering, and retrieval across many languages.
Multilingual-E5-Large is a purely text-based model that converts text inputs into dense vector embeddings.
You call the LLM.API embeddings endpoint, specifying provider 'Intfloat' and model 'Multilingual-E5-Large' in your request parameters.
It is best for multilingual semantic search, dense retrieval, reranking pipelines, and deduplication where cross-language similarity detection is important.
Pricing is usage-based per embedding token on LLM.API, and you should check the LLM.API pricing page for current Multilingual-E5-Large rates.
Multilingual-E5-Large supports reasonably long text inputs, but you should chunk very long documents before embedding for best performance and latency.
Latency is typically low and dominated by text length and batch size, making it suitable for real-time or near real-time search applications.
It generally offers better performance on non-English and cross-lingual tasks, while strong English-only models may outperform it on purely English benchmarks.
Yes, you can send multiple texts in one embeddings request to Multilingual-E5-Large to reduce overhead and improve throughput.
It cannot generate text, may underperform on languages outside its training distribution, and can encode training-time biases into the resulting embeddings.
Compare
Qwen3 ASR Flash is Qwen’s high-accuracy, multilingual automatic speech recognition (ASR) service optimized for real-time transcription of short audio. It is built on the Qwen3-Omni foundation…
Seedance 1.5 Pro is ByteDance’s flagship native joint audio‑video generation model, focused on high‑quality, lip‑synced video with synchronized sound. It is notable for producing short, production‑ready…
LFM2.5-1.2B-Thinking (free) is LiquidAI’s 1.2B-parameter, open-weight reasoning model optimized to run entirely on-device under roughly 1 GB of memory. It focuses on chain-of-thought style “thinking” before…