- Text Generation
Anthropic Claude Sonnet Latest refers to the most recent mid-tier Claude Sonnet language model from Anthropic, designed to balance strong intelligence with speed and cost-efficiency. It…
Powered by Sesame
CSM 1B is a 1‑billion‑parameter conversational speech model from Sesame that turns text (and optionally audio context) into natural‑sounding English speech. It is notable for its Llama-based architecture and RVQ/Mimi audio code generation that enables high‑quality, low‑latency voice output.
Output tokens per second · Higher is better
Seconds · Lower is better
USD per 1M tokens (blended) · Lower is better
About the model
CSM 1B is a 1‑billion‑parameter conversational speech generation model from Sesame that converts text into English speech. It is mainly used for dialogue-oriented applications such as voice assistants, interactive agents, and chat-style voice interfaces that need realistic, contextual speech output. It is also applied in text-to-speech pipelines for content creation, accessibility tools, and other products that require controllable, high‑fidelity synthetic voices. CSM 1B belongs to Sesame’s CSM (Conversational Speech Model) family, which uses a Llama backbone with a specialized audio decoder that generates RVQ/Mimi audio codes.
Model capabilities
Generates high-quality, natural-sounding speech audio directly from text using a Llama-based backbone and specialized audio decoder.
Maintains contextual awareness across turns, adjusting tone, pauses, and inflection to match dialogue flow and emotional nuance.
Processes both text and audio inputs, using prior audio context to guide consistent speech patterns and expressive delivery.
Supports efficient, low-latency speech synthesis suitable for real-time or interactive applications and local deployment scenarios.
Can be fine-tuned on additional languages, enabling customized voices and language support beyond the base English-focused model.
Use cases
Transparent pricing
LLM API offers the lowest cost and highest performance for CSM 1B-class models.
| Provider | Region | Latency | Throughput | Uptime | Input ($/1M) | Output ($/1M) | Context |
|---|---|---|---|---|---|---|---|
| LLM API BEST | Global | ~120ms | ~120 tps | 99.99% | $0.08 | $0.08 | 64K tokens |
| Sesame | Global | ~220ms | ~60 tps | ~99.9% | ~$0.15 | ~$0.15 | ~32K tokens |
| OpenRouter | Global | ~250ms | ~55 tps | ~99.9% | ~$0.18 | ~$0.18 | ~32K tokens |
| Replicate | US East | ~280ms | ~45 tps | ~99.5% | ~$0.20 | ~$0.20 | ~16K tokens |
Performance benchmarks
| Metric | CSM 1B (Sesame) | Llama 3.2 1B (Meta) | Phi-3 Mini 1.1B (Microsoft) |
|---|---|---|---|
| Avg Latency | ~220ms | ~250ms | ~260ms |
| Context Window | ~32K | 8K | 8K |
| Input Price ($/1M tokens) | ~$0.05 | ~$0.10 | ~$0.05 |
| Output Price ($/1M tokens) | ~$0.15 | ~$0.30 | ~$0.20 |
| Max Output Tokens | ~4K | 2K | 2K |
| Throughput | ~60 tps | ~40 tps | ~35 tps |
| Uptime | ~99.9% | ~99.9% | ~99.9% |
30-day usage via LLM API
Architecture & Integration
One unified API. Every major model. Built-in reliability, cost control, and observability.
Dynamically route each request to the best-fit model across providers based on latency, accuracy, or cost—without changing your integration or redeploying code.
One endpoint, every modelOptimize spend with per-route price caps, automatic model downgrades, and detailed usage insights so you can keep quality high while staying within budget.
Max quality, min costAvoid downtime by failing over to alternative models or providers on errors, rate limits, or outages—configured once, enforced globally in real time.
Resilience by defaultTrace every call across providers with request logs, latency breakdowns, errors, and cost metrics so you can debug faster and tune performance with confidence.
See every tokenUse high-level tasks—chat, tools, RAG, agents—instead of provider-specific APIs, so you can upgrade models or vendors without rewriting your application logic.
Code to tasks, not APIsRun large-scale jobs—evaluations, backfills, fine-tuning prep—through a single batch interface with concurrency, retries, and throttling handled by the platform.
Ship bulk workloads fastDecision guide
FAQ
CSM 1B is a 1-billion-parameter language model from Sesame, accessible through LLM.API for lightweight, cost-efficient text generation and understanding.
CSM 1B is best for lightweight chatbots, autocomplete, short-form content generation, and simple classification or extraction tasks where low cost matters.
CSM 1B supports a 4K-token context window, making it suitable for short conversations, prompts, and small documents.
CSM 1B is a text-only model, accepting text prompts and returning text completions without image, audio, or video support.
Use the LLM.API completions or chat endpoint with the model parameter set to "Sesame/CSM-1B" and include your usual authorization header.
CSM 1B is cheaper and faster but generally less capable on complex reasoning, long-context, and nuanced instruction-following than larger Sesame models.
CSM 1B is optimized for low latency, typically returning first tokens faster than larger models at similar throughput settings.
CSM 1B may struggle with long multi-step reasoning, very domain-specific technical tasks, and maintaining consistency over extended dialogs.
Yes, you can enable streaming in LLM.API requests to receive CSM 1B tokens incrementally as they are generated.
CSM 1B is priced as a budget-friendly tier on LLM.API, with lower per-token costs than larger Sesame and frontier models.
Compare
Anthropic Claude Sonnet Latest refers to the most recent mid-tier Claude Sonnet language model from Anthropic, designed to balance strong intelligence with speed and cost-efficiency. It…
Cydonia 24B V4.1 is a 24-billion-parameter, open-source text language model by TheDrummer, fine-tuned from Mistral Small 3.2 and optimized for uncensored creative writing with a 131K-token…
Riverflow V2 Pro is Sourceful’s most powerful Riverflow 2.0 model, focused on high-quality, controllable image generation and perfect text rendering.