- Text Generation
Aion-2.0 is a text-only large language model from AionLabs, fine-tuned from DeepSeek V3.2 and optimized for immersive roleplaying and storytelling. It offers a 131K-token context window…
Powered by TheDrummer
Cydonia 24B V4.1 is a 24-billion-parameter, open-source text language model by TheDrummer, fine-tuned from Mistral Small 3.2 and optimized for uncensored creative writing with a 131K-token context window. It is notable for combining strong long-context handling with budget-friendly pricing for high-volume use.
Output tokens per second · Higher is better
Seconds · Lower is better
USD per 1M tokens (blended) · Lower is better
About the model
Cydonia 24B V4.1 is an open‑source, text‑to‑text language model by TheDrummer built on Mistral Small 3.2 24B with a ~131K token context window. It is primarily used for uncensored creative writing, roleplay, and narrative-heavy chat where mood, nuance, and consistent characterization matter over long conversations. It is also applied as a general-purpose assistant model in enterprise and hobbyist settings, offering relatively low per-token costs for large-context workloads. Cydonia 24B V4.1 continues TheDrummer’s Cydonia series, improving on earlier variants such as Cydonia-22B and Cydonia-24B-v2.x in focus, coherence, and writing quality.
Model capabilities
Engages in multi-turn dialogue, answering questions, following instructions, and maintaining context within general conversational and assistant tasks.
Reads and writes code or technical text, explaining behavior, debugging issues, and providing structured suggestions within its training scope.
Processes image inputs to identify objects and scenes and provide descriptive text responses within its supported visual understanding abilities.
Extracts readable text from images or screenshots and converts it into machine-readable form for further processing or analysis.
Translates written text between multiple languages, preserving meaning and tone as closely as possible within its training limitations.
Use cases
Transparent pricing
LLM API offers the lowest cost and highest performance for Cydonia 24B–class models.
| Provider | Region | Latency | Throughput | Uptime | Input ($/1M) | Output ($/1M) | Context |
|---|---|---|---|---|---|---|---|
| LLM API BEST | Global | ~120ms | ~65 tps | ~99.99% | ~$0.25 | ~$0.25 | ~256K |
| TheDrummer | Global | ~220ms | ~30 tps | ~99.5% | ~$0.80 | ~$0.80 | ~64K |
| Together AI | US East | ~210ms | ~35 tps | ~99.9% | ~$0.70 | ~$0.70 | ~128K |
| RunPod | US West | ~260ms | ~25 tps | ~99.0% | ~$0.90 | ~$0.90 | ~32K |
| Banana | Global | ~240ms | ~28 tps | ~99.5% | ~$0.85 | ~$0.85 | ~64K |
Performance benchmarks
| Metric | Cydonia 24B V4.1 | Llama 3 70B Instruct | Qwen2 72B |
|---|---|---|---|
| Avg Latency | ~220ms | ~260ms | ~280ms |
| Context Window | 128K | 8K | 32K |
| Input Price ($/1M) | $0.40 | $0.60 | $0.45 |
| Output Price ($/1M) | $0.60 | $0.90 | $0.70 |
| Max Output Tokens | 4K | 4K | 4K |
| Throughput | 65 tps | 50 tps | 55 tps |
| Uptime | 99.5% | 99.9% | 99.5% |
30-day usage via LLM API
Architecture & Integration
One unified API. Every major model. Built-in reliability, cost control, and observability.
Dynamically route each request across providers and models based on latency, cost, and quality—without changing your integration.
One endpoint, any modelAutomatically pick the most cost-effective model for each task and track spend per project, environment, and feature in one place.
Optimize tokens, not codeDefine fallback chains across providers so requests transparently recover from outages, rate limits, and model regressions.
Keep responses flowingGet end-to-end traces, latency and error metrics, and model-level analytics to debug prompts and production traffic in real time.
See every token hopDescribe tasks like chat, tool use, search, or generation once, then plug in any model or provider behind the same interface.
Ship tasks, not glue codeFan out thousands of requests per call with built-in retries, rate management, and structured result aggregation.
Scale from 10 to 10M callsDecision guide
FAQ
Cydonia 24B V4.1 is a 24-billion-parameter language model by TheDrummer focused on fast, general-purpose code and text generation via LLM.API.
It is best for code completion, technical writing, tool-using agents, and structured data generation where latency and cost matter.
Cydonia 24B V4.1 supports a context window of up to 32,000 tokens per request.
Cydonia 24B V4.1 is a text-only model that accepts and outputs UTF-8 text.
Pricing is usage-based per 1,000 tokens, with separate rates for input and output tokens defined in your LLM.API account.
Typical end-to-end latency is in the low hundreds of milliseconds for short prompts, depending on load and request size.
Specify the model name "TheDrummer/cydonia-24b-v4.1" in your LLM.API completion or chat endpoint requests with your API key.
It targets a balance of stronger coding ability and lower latency than many open 20–30B models at similar price points.
Yes, you can use LLM.API's tool or function-calling conventions with this model for agent-style workflows.
It may hallucinate facts, lacks real-time knowledge, and is not guaranteed safe for high-stakes decisions without human review.
Compare
Aion-2.0 is a text-only large language model from AionLabs, fine-tuned from DeepSeek V3.2 and optimized for immersive roleplaying and storytelling. It offers a 131K-token context window…
Qwen3 VL 30B A3B Instruct is a 30B-parameter Mixture-of-Experts vision-language model from Qwen, offering strong multimodal understanding and generation with a 262K-token context window. It is…
Qwen3 ASR Flash is Qwen’s high-accuracy, multilingual automatic speech recognition (ASR) service optimized for real-time transcription of short audio. It is built on the Qwen3-Omni foundation…