- Text Generation
Riverflow V2 Fast Preview is Sourceful’s fastest preview variant in the Riverflow V2 lineup, offering high-throughput text-to-image and image-to-image generation with an 8K token context window.
Powered by ByteDance
Seedance 2.0 is ByteDance’s next-generation multimodal AI video generation model that natively combines audio and video to create highly realistic clips from simple prompts. It is notable for its quad-modal inputs (text, image, audio, video) and strong consistency across complex, multi-shot scenes.
Output tokens per second · Higher is better
Seconds · Lower is better
USD per 1M tokens (blended) · Lower is better
About the model
Seedance 2.0 is a unified multimodal audio-video generation model developed by ByteDance for high-fidelity, realistic video creation. It is mainly used for text-to-video and story-driven clip generation where creators script cinematic sequences with detailed control over motion, camera, and scene transitions. It is also used in consumer and professional tools like CapCut and Dreamina-style services to turn scripts, reference images, or rough edits into polished short-form content and trailers. Seedance 2.0 follows earlier Seedance 1.0/1.5 video models within ByteDance’s broader SEED family that also includes the Doubao language models and Seedream image generators.
Model capabilities
Generates coherent, cinematic video clips directly from text prompts, supporting multi-shot narratives with consistent characters, scenes, and camera movements.
Accepts combined text, image, audio, and video inputs in a unified model to guide structure, style, and motion of generated videos.
Jointly generates synchronized soundtracks, effects, and dialogue with videos, enabling frame-accurate lip-sync and environment-aware sound design.
Supports shot-by-shot control and reference-based editing, allowing users to refine pacing, composition, and continuity without manual post-production.
Handles prompts and control inputs in multiple languages, enabling creators worldwide to direct and customize video generation workflows.
Use cases
Transparent pricing
LLM API offers the lowest cost and highest performance for Seedance 2.0–class models.
| Provider | Region | Latency | Throughput | Uptime | Input ($/1M) | Output ($/1M) | Context |
|---|---|---|---|---|---|---|---|
| LLM API BEST | Global | ~150ms | ~120 tps | 99.99% | $0.10 | $0.30 | 128K tokens |
| ByteDance | Asia Pacific | ~220ms | ~80 tps | ~99.9% | ~$0.18 | ~$0.50 | ~64K tokens |
| OpenAI-compatible Gateway | Global | ~260ms | ~70 tps | ~99.9% | ~$0.22 | ~$0.60 | ~64K tokens |
| Cloud Aggregator X | US East | ~240ms | ~65 tps | ~99.5% | ~$0.25 | ~$0.70 | ~32K tokens |
Performance benchmarks
| Metric | Seedance 2.0 | GPT-4o | Claude 3.5 Sonnet |
|---|---|---|---|
| Avg Latency | ~220ms | ~280ms | ~180ms |
| Context Window | 128K | 200K | 128K |
| Input Price ($/1M) | $0.40 | $1.00 | $0.15 |
| Output Price ($/1M) | $0.60 | $2.00 | $0.60 |
| Max Output Tokens | 8K | 8K | 4K |
| Throughput | 40 tps | 30 tps | 60 tps |
| Uptime | 99.9% | 99.9% | 99.9% |
30-day usage via LLM API
Architecture & Integration
One unified API. Every major model. Built-in reliability, cost control, and observability.
Automatically route each request to the optimal model across providers based on latency, quality, and cost, without changing your integration or redeploying code.
One endpoint, any modelOptimize spend with dynamic provider selection, granular usage controls, and model-tier policies so you can keep quality high while staying within strict budgets.
Maximum value per tokenDefine automatic failover chains across providers and models, ensuring your production workloads keep running even during outages, rate limits, or degraded service.
No more single pointsGet full visibility into latency, errors, token usage, and provider performance with request-level traces and metrics that plug into your existing monitoring stack.
Debug across providersCall high-level tasks like chat, tools, or RAG through a unified schema, while LLM.API handles provider quirks, formats, and evolving capabilities under the hood.
Code to tasks, not APIsProcess large workloads with efficient batching, concurrency controls, and job-level status APIs, letting you scale evaluations, backfills, and bulk inference safely.
Scale jobs, not stressDecision guide
FAQ
Seedance 2.0 is a large language model by ByteDance focused on fast, low-cost text generation for general-purpose applications.
Seedance 2.0 is best for high-volume chatbots, content generation, and lightweight reasoning where throughput and cost efficiency matter more than frontier-level capabilities.
Seedance 2.0 supports a 16K token context window, suitable for long conversations, multi-step tools workflows, and moderately long documents.
Through LLM.API, Seedance 2.0 currently supports text input and text output only, without native image, audio, or video support.
Typical end-to-end latency ranges from 300ms to a few seconds per request on LLM.API, depending on prompt length and concurrency.
LLM.API exposes Seedance 2.0 with per-token pricing, charging separately for input tokens and output tokens, plus any provider-specific minimums.
Use the LLM.API chat or completion endpoint with the provider set to "bytedance" and the model name set to "Seedance 2.0".
Seedance 2.0 typically trades off peak reasoning quality for higher throughput and lower cost than many flagship frontier models.
Yes, Seedance 2.0 can be integrated with tools on LLM.API using the standard function-calling or tool-calling schema supported by the gateway.
Seedance 2.0 may struggle with complex long-horizon reasoning, precise mathematical proofs, and tasks requiring up-to-date proprietary or domain-specific knowledge.
Compare
Riverflow V2 Fast Preview is Sourceful’s fastest preview variant in the Riverflow V2 lineup, offering high-throughput text-to-image and image-to-image generation with an 8K token context window.
MiMo-V2-Flash is an open-source Mixture-of-Experts language model from Xiaomi optimized for fast, long-context reasoning and coding. It combines a 309B-parameter MoE architecture with only 15B active…
MiniMax M2.5 is a frontier-class, agent-native large language model from MiniMax that combines a Mixture-of-Experts architecture with long-context, cost-efficient inference for real-world productivity tasks.