- Text Generation
Claude Sonnet 4.5 is an Anthropic large language model optimized for software development, computer use, and agentic workflows, offering strong performance on coding and reasoning tasks…
Powered by Google
Lyria 3 Clip Preview is Google's preview music-generation model optimized for creating short, 30‑second musical clips, loops, and previews from text or image prompts.
Output tokens per second · Higher is better
Seconds · Lower is better
USD per 1M tokens (blended) · Lower is better
About the model
Lyria 3 Clip Preview is a Google music-generation model that produces high-quality short audio clips from text and image inputs. It is mainly used to generate 30-second music snippets, loops, and previews for creative, media, and sound design workflows. It also supports features like vocal or instrumental modes, user- or model-generated lyrics, and controls such as BPM and intensity to shape the resulting clip. It belongs to the Lyria 3 family of music-generation models, alongside Lyria 3 Pro, and is offered as a preview model via the Gemini API and related Google Cloud platforms.
Model capabilities
Generates 30-second high-quality stereo music clips from detailed text prompts, including structure, style, mood, and instrumentation guidance.
Creates musical clips and previews conditioned on input images, translating visual themes and scenes into coherent audio compositions.
Supports vocal generation, lyric generation, and user-provided lyrics to produce clips with synchronized singing and musical phrasing.
Offers BPM and intensity controls plus instrumental mode, enabling tailored rhythmic feel, energy, and arrangement for generated clips.
Applies input filtering, output recitation filtering, and vocal similarity filtering, alongside audio watermarking for safer music outputs.
Use cases
Transparent pricing
LLM API offers the lowest cost and latency for Lyria 3 Clip–class models.
| Provider | Region | Latency | Throughput | Uptime | Input ($/1M) | Output ($/1M) | Context |
|---|---|---|---|---|---|---|---|
| Global | ~550ms | ~8 img/s | 99.9% | ~$1.20/1K images | $0.00 | ~10 min video or 20 images | |
| Vertex AI (Google Cloud) | US East | ~480ms | ~10 img/s | 99.9% | ~$1.30/1K images | $0.00 | ~10 min video or 20 images |
| Replicate | US West | ~750ms | ~5 img/s | 99.5% | ~$1.80/1K images | $0.00 | ~8 min video or 16 images |
| LLM API BEST | Global | 180ms | 20 img/s | 99.99% | $0.80/1K images | $0.00 | 12 min video or 32 images |
Performance benchmarks
| Metric | Lyria 3 Clip Preview (Google) | CLIP ViT-L/14 (OpenAI) | SigLIP Large Patch16-384 (Google) |
|---|---|---|---|
| Avg Latency | ~800ms | ~700ms | ~900ms |
| Context Window | 128K | 128K | 200K |
| Input Price ($/1M tokens) | ~$0.20 | $0.15 | $0.25 |
| Output Price ($/1M tokens) | ~$0.60 | $0.60 | $1.25 |
| Max Output Tokens | 8K | 16K | 4K |
| Throughput | ~80 tps | ~100 tps | ~60 tps |
| Uptime | ~99.9% | ~99.9% | ~99.9% |
30-day usage via LLM API
Architecture & Integration
One unified API. Every major model. Built-in reliability, cost control, and observability.
Dynamically route each request to the optimal model across providers based on latency, cost, and performance—without changing your code or integrations.
One endpoint, every modelOptimize spend by automatically choosing cheaper equivalents, downgrading when quality allows, and enforcing per-project budgets with centralized cost controls and analytics.
Max performance, minimal spendEliminate single-vendor downtime with automatic failover to backup models and providers, using configurable rules, health checks, and graceful degradation strategies.
Stay online, even when APIs failTrace every call across providers with logs, metrics, and structured events to debug prompts, track latency, and understand model behavior in production.
See every token, every hopDescribe tasks—chat, generation, extraction, tools—once and let LLM.API select and configure the right models, prompts, and parameters for each use case.
Think in tasks, not modelsRun large-scale batch inference with automatic chunking, concurrency control, retries, and progress tracking designed for data pipelines and offline processing.
Ship millions of calls safelyDecision guide
FAQ
Lyria 3 Clip Preview is a Google model for generating short video clips from text prompts, accessible through the unified LLM.API gateway.
It is best for quickly prototyping and previewing short video concepts, storyboards, and motion ideas directly from text descriptions.
Lyria 3 Clip Preview consumes text prompts and outputs short video clips, without support for audio or image-only inputs in this preview tier.
Pricing is usage-based per generated clip or generated video-second, with exact rates defined in your LLM.API pricing dashboard.
The model supports moderately long text prompts, typically a few paragraphs, but you should avoid excessively long scripts or scene-by-scene screenplays.
Latency depends on clip length and load, but you should expect generation to take several seconds up to a couple of minutes per request.
Use the standard LLM.API generation endpoint, specifying the Google provider and the "Lyria 3 Clip Preview" model name in your request payload.
It emphasizes quick preview-quality clips for ideation rather than long-duration, high-fidelity cinematic video compared to heavier video generation models.
Limitations include short clip duration, preview-level visual quality, occasional motion artifacts, and imperfect adherence to very detailed or complex scene instructions.
It is better suited for concepting and iteration; production workflows typically require post-processing or higher-fidelity models for final output.
Compare
Claude Sonnet 4.5 is an Anthropic large language model optimized for software development, computer use, and agentic workflows, offering strong performance on coding and reasoning tasks…
Devstral 2 2512 is a 123B-parameter open-source large language model from Mistral AI, optimized for agentic coding and long-context software engineering workflows. It supports a 256K/262K-token…
Grok Imagine Image Quality is an image-focused evaluation or enhancement component from xAI’s Grok ecosystem, aimed at assessing or improving the visual fidelity of generated images.…