- Text Generation
Lyria 3 Clip Preview is Google's preview music-generation model optimized for creating short, 30‑second musical clips, loops, and previews from text or image prompts.
Powered by Sourceful
Riverflow V2 Fast is the fastest variant of Sourceful’s Riverflow 2.0 image generation and editing lineup, optimized for production deployments and latency‑critical brand creative workflows.
Output tokens per second · Higher is better
Seconds · Lower is better
USD per 1M tokens (blended) · Lower is better
About the model
Riverflow V2 Fast is a production-grade image generation and editing model from Sourceful, tuned for rapid, low-latency performance. It is mainly used for marketing and creative applications where teams need fast iteration on packaging, campaign visuals, and other brand assets at scale. It also serves latency‑sensitive deployments such as interactive design tools and high-throughput content pipelines. Riverflow V2 Fast belongs to the Riverflow 2.0 family of visual AI models, following earlier Riverflow 1 and Riverflow V2 preview releases.
Model capabilities
Supports low-latency conversational interactions with an 8K token context window, suitable for production chat and assistant experiences.
Creates images from text prompts, optimized for latency-critical workflows using Sourceful’s Riverflow 2.0 text-to-image architecture.
Performs image-to-image transformations, enabling complex multi-step edits and enhancements guided by an integrated reasoning model.
Designed for production deployments with high throughput, making it suitable for large-scale, continuously running applications and services.
Handles prompts in multiple languages for image generation and editing, enabling localized creative workflows across global user bases.
Use cases
Transparent pricing
LLM API offers the lowest costs and fastest performance for Riverflow V2 Fast–class models.
| Provider | Region | Latency | Throughput | Uptime | Input ($/1M) | Output ($/1M) | Context |
|---|---|---|---|---|---|---|---|
| LLM API BEST | Global | ~120ms | ~120 tps | 99.99% | ~$0.08 | ~$0.16 | ~256K |
| Sourceful | Global | ~180ms | ~80 tps | ~99.9% | ~$0.12 | ~$0.24 | ~128K |
| OpenAI (comparable fast model) | Global | ~220ms | ~70 tps | ~99.9% | ~$0.15 | ~$0.60 | ~128K |
| Anthropic (comparable fast model) | US East | ~250ms | ~60 tps | ~99.9% | ~$0.20 | ~$0.80 | ~200K |
| Fireworks.ai (comparable fast model) | US West | ~200ms | ~90 tps | ~99.9% | ~$0.10 | ~$0.30 | ~128K |
Performance benchmarks
| Metric | Riverflow V2 Fast (Sourceful) | OpenAI gpt-4.1-mini | Anthropic Claude 3 Haiku |
|---|---|---|---|
| Avg Latency | ~180ms | ~220ms | ~250ms |
| Context Window | 128K | 128K | 200K |
| Input Price ($/1M tokens) | $0.10 | $0.15 | $0.25 |
| Output Price ($/1M tokens) | $0.25 | $0.60 | $1.25 |
| Max Output Tokens | 4K | 4K | 4K |
| Throughput | 80 tps | 60 tps | 55 tps |
| Uptime | 99.9% | 99.9% | 99.9% |
30-day usage via LLM API
Architecture & Integration
One unified API. Every major model. Built-in reliability, cost control, and observability.
Dynamically route each request to the optimal model across providers based on cost, latency, or performance—no client changes, just smarter traffic decisions.
One endpoint, every modelControl spend with per-request cost policies, price-aware routing, and usage caps so you can scale AI features without surprise bills or manual tuning.
Reduce cost, keep qualityHandle provider outages and rate limits automatically with policy-based failover to backup models, preserving uptime and user experience without custom recovery code.
Resilient by defaultGet full visibility into latency, errors, tokens, and provider performance with request-level tracing and structured logs that plug into your existing monitoring stack.
Trace every tokenDefine reusable tasks—chat, RAG, tools, moderation—once, then run them on any model or provider with consistent inputs, outputs, and guardrails.
Ship features, not callsRun massive batch inference jobs efficiently with parallelized execution, retry semantics, and cost controls tuned for large datasets and background processing.
Scale jobs, not codeDecision guide
FAQ
Riverflow V2 Fast is a Sourceful language model optimized for fast, low-cost text generation accessed through the unified LLM.API gateway.
Riverflow V2 Fast is best for high-volume tasks like chatbots, lightweight reasoning, and content generation where speed and cost-efficiency matter most.
Riverflow V2 Fast supports a context window of up to 8,192 tokens via LLM.API.
Riverflow V2 Fast is designed for low-latency responses, typically suitable for interactive applications and real-time chat workloads.
Riverflow V2 Fast currently supports text-only input and output via LLM.API.
Riverflow V2 Fast uses a pay-as-you-go per-token pricing model on LLM.API, with separate rates for input and output tokens.
You select the Sourceful Riverflow V2 Fast model name in your LLM.API request and send standard chat or completion-style payloads.
Riverflow V2 Fast trades some reasoning depth and accuracy for significantly better throughput, latency, and cost efficiency than larger flagship models.
Riverflow V2 Fast may struggle with very complex reasoning, long multi-step instructions, and highly specialized domain knowledge compared to larger models.
Riverflow V2 Fast can handle moderately long documents within its context window but may require chunking for extensive documents or multi-document workflows.
Compare
Lyria 3 Clip Preview is Google's preview music-generation model optimized for creating short, 30‑second musical clips, loops, and previews from text or image prompts.
Riverflow V2 Max Preview is Sourceful’s most powerful Riverflow V2 preview model, a unified text-to-image and image-to-image generator. It is designed to exceed the performance of…
Aion-2.0 is a text-only large language model from AionLabs, fine-tuned from DeepSeek V3.2 and optimized for immersive roleplaying and storytelling. It offers a 131K-token context window…