- Text Generation
Text Embedding 3 Large is OpenAI’s high‑capacity embedding model optimized for semantic search, retrieval, and clustering tasks. It provides high‑quality vector representations of text with strong…
Powered by Alibaba
Wan 2.7 is Alibaba’s latest open-source multimodal visual generation model for high-quality video and image creation, offering text-to-video, image-to-video, text-to-image, and editing in a single architecture.
Output tokens per second · Higher is better
Seconds · Lower is better
USD per 1M tokens (blended) · Lower is better
About the model
Wan 2.7 is an AI visual generation model from Alibaba that unifies video and image generation and editing in one system. It is mainly used for generating cinematic short videos from text or image prompts and for image-to-video transformations in creative, marketing, and storytelling workflows. It is also used for text-to-image creation, reference-guided image generation, and instruction-based image or video editing for design and content production teams. Wan 2.7 is part of Alibaba’s Wan video model family developed within the broader Qwen ecosystem, succeeding earlier Wan 2.x releases.
Model capabilities
Generates high-quality video clips directly from detailed text prompts, supporting controllable camera movement, scenes, and lighting for creators.
Animates still images into coherent motion videos, preserving subject appearance and layout while adding realistic movement and transitions.
Edits and extends existing video using reference frames and instructions, enabling consistent subjects, motion control, and frame-level refinements.
Acts as a multimodal visual model handling text-to-video, image-to-video, text-to-image, and image editing within a single architecture.
Interprets user intent before rendering with a dedicated thinking phase, improving creative consistency, controllability, and reducing failed generations.
Use cases
Transparent pricing
LLM API offers the lowest cost and highest performance for Wan 2.7–class models.
| Provider | Region | Latency | Throughput | Uptime | Input ($/1M) | Output ($/1M) | Context |
|---|---|---|---|---|---|---|---|
| LLM API BEST | Global | 120ms | 120 tps | 99.99% | $0.15 | $0.45 | 128K tokens |
| Alibaba Cloud | APAC East | ~260ms | ~60 tps | ~99.95% | ~$0.40 | ~$1.20 | ~32K tokens |
| OpenAI | Global | ~180ms | ~90 tps | 99.9% | ~$0.50 | ~$1.50 | ~128K tokens |
| Azure AI | US East | ~200ms | ~80 tps | 99.9% | ~$0.55 | ~$1.60 | ~128K tokens |
| Anthropic | US West | ~190ms | ~70 tps | ~99.9% | ~$0.60 | ~$1.80 | ~200K tokens |
Performance benchmarks
| Metric | Wan 2.7 | Qwen2.5-72B-Instruct | Llama 3.1 70B |
|---|---|---|---|
| Avg Latency | ~220ms | ~250ms | ~260ms |
| Context Window | 128K | 128K | 128K |
| Input Price ($/1M) | $0.60 | $0.80 | $1.00 |
| Output Price ($/1M) | $2.40 | $3.20 | $4.00 |
| Max Output Tokens | 4K | 4K | 4K |
| Throughput | 48 tps | 40 tps | 42 tps |
| Uptime | 99.9% | 99.9% | 99.9% |
30-day usage via LLM API
Architecture & Integration
One unified API. Every major model. Built-in reliability, cost control, and observability.
Dynamically route each request to the optimal model across providers based on performance, cost, and availability—without changing your code or client integration.
One endpoint, any modelAutomatically balance premium and budget models with configurable cost ceilings, so you keep latency low and quality high while tightly controlling spend.
Optimize quality per dollarDefine provider and model failover chains so requests transparently retry on alternates, insulating your app from regional outages, rate limits, or model regressions.
Stay online, even upstreamGet unified traces, latency, error, and token metrics across all providers with request-level logs for fast debugging, tuning, and capacity planning.
See every token, everywhereCall high-level tasks—chat, generation, tools, and more—instead of vendor-specific APIs, so you can swap models without rewriting application logic.
Program to tasks, not modelsSubmit large batches of prompts in a single call with automatic chunking, concurrency control, and retries to maximize throughput and minimize overhead.
Scale workloads, not codeDecision guide
FAQ
Wan 2.7 is an Alibaba large language model accessible via LLM.API, targeting general-purpose text generation and understanding tasks.
Wan 2.7 is best for cost-efficient chatbots, content generation, and general NLP tasks where balanced quality and efficiency matter.
Wan 2.7 supports a context window of up to 8,192 tokens via LLM.API.
Wan 2.7 is optimized for low latency on LLM.API, typically returning first tokens within a few hundred milliseconds under normal load.
Wan 2.7 is a text-only model on LLM.API, supporting text inputs and text outputs.
Wan 2.7 uses LLM.API’s unified token-based billing, with separate input and output token rates shown in your LLM.API pricing dashboard.
You select provider 'Alibaba' and model 'Wan 2.7' in the LLM.API request payload, keeping the standard chat or completion schema unchanged.
Wan 2.7 generally trades slightly lower peak quality than top-tier frontier models for better cost efficiency and predictable performance.
Yes, Wan 2.7 supports token streaming via LLM.API by enabling the standard 'stream' flag in your request.
Wan 2.7 may struggle with highly specialized domain knowledge, strict mathematical reasoning, and tasks requiring very long-context retention beyond its context window.
Compare
Text Embedding 3 Large is OpenAI’s high‑capacity embedding model optimized for semantic search, retrieval, and clustering tasks. It provides high‑quality vector representations of text with strong…
Relace Search is a text-only large language model from Relace optimized for agentic multi-step search over large codebases, using parallel file-inspection tools to return highly relevant…
Nemotron 3 Super is NVIDIA’s open-weight, 120B-parameter hybrid Mamba-Transformer Mixture-of-Experts language model optimized for high-throughput agentic reasoning workloads. It is notable for combining LatentMoE experts, long-context…