- Text Generation
Olmo 3 32B Think is a 32-billion-parameter open-weight reasoning model from the Allen Institute for AI, optimized for deep chain-of-thought reasoning and complex instruction following. It…
Powered by Writer
Palmyra X5 is Writer's most advanced enterprise large language model, featuring an extremely long context window and adaptive reasoning for complex business workflows. It is purpose-built for building and scaling AI agents across the enterprise with strong performance on long-form, text-heavy tasks.
Output tokens per second · Higher is better
Seconds · Lower is better
USD per 1M tokens (blended) · Lower is better
About the model
Palmyra X5 is Writer’s flagship enterprise large language model designed for adaptive reasoning over very long text inputs. It is used for enterprise content generation and long-document analysis, such as processing extensive reports, knowledge bases, and regulatory or research materials, and for powering AI agents that automate complex business workflows across domains like finance, healthcare, and software. It belongs to Writer’s Palmyra family of foundation models and succeeds earlier generations such as Palmyra X4.
Model capabilities
Performs deep, multi-step reasoning over complex business tasks, enabling reliable enterprise agents and sophisticated decision-support workflows.
Processes and grounds responses in very long inputs, supporting analysis of large document sets and extensive enterprise knowledge bases.
Calls external tools and composes multi-step AI agents, orchestrating workflows such as retrieval, APIs, and database interactions.
Understands and generates text in over 30 languages, enabling global enterprise deployments and cross-lingual workflows.
Accepts images as inputs to inform responses, allowing multimodal enterprise workflows that combine visual data with text.
Use cases
Transparent pricing
LLM API offers the lowest token prices and latency for Palmyra X5–class models.
| Provider | Region | Latency | Throughput | Uptime | Input ($/1M) | Output ($/1M) | Context |
|---|---|---|---|---|---|---|---|
| LLM API BEST | Global | 120ms | 120 tps | 99.99% | $0.40 | $0.80 | 128K |
| Writer | US | ~220ms | ~60 tps | 99.9% | ~$0.60 | ~$1.20 | 32K |
| OpenAI (closest: GPT-4.1-mini) | Global | ~250ms | ~80 tps | 99.9% | ~$0.50 | ~$1.00 | 128K |
| Anthropic (closest: Claude 3.5 Haiku) | US East | ~260ms | ~70 tps | 99.9% | ~$0.55 | ~$1.10 | 200K |
| Google Cloud (closest: Gemini 1.5 Pro) | Global | ~280ms | ~65 tps | 99.9% | ~$0.70 | ~$1.40 | 1M |
Performance benchmarks
| Metric | Palmyra X5 (Writer) | GPT-4.1 Mini (OpenAI) | Claude 3.5 Sonnet (Anthropic) |
|---|---|---|---|
| Avg Latency | ~220ms | ~180ms | ~250ms |
| Context Window | 128K | 128K | 200K |
| Input Price ($/1M) | $0.80 | $0.15 | $3.00 |
| Output Price ($/1M) | $2.40 | $0.60 | $15.00 |
| Max Output Tokens | 8K | 4K | 8K |
| Throughput | 40 tps | 60 tps | 35 tps |
| Uptime | 99.9% | 99.9% | 99.9% |
30-day usage via LLM API
Architecture & Integration
One unified API. Every major model. Built-in reliability, cost control, and observability.
Automatically route each request to the best model across providers based on cost, latency, and quality—no client changes required when your stack evolves.
One endpoint, every modelEnforce budgets, compare provider pricing, and transparently shift traffic to cheaper equivalents while preserving quality so you never overspend on inference again.
Cut spend, keep qualityDefine automatic fallbacks across models and providers so timeouts, rate limits, or outages degrade gracefully instead of taking your product offline.
Never fail on 500sTrace every request across providers with metrics, logs, and latency breakdowns so you can debug incidents and tune model routing in minutes, not days.
See every tokenDescribe tasks like chat, extraction, or classification once and let LLM.API pick the right models and prompts, simplifying integration and future migrations.
Code to tasks, not modelsSend massive batches through a single API with concurrency controls and provider-optimized chunking to cut latency and costs for large-scale workloads.
Ship thousands at onceDecision guide
FAQ
Palmyra X5 is a large language model from Writer focused on enterprise-grade text generation, editing, and knowledge-intensive tasks.
Palmyra X5 is best for long-form content generation, marketing copy, product documentation, and domain-specific enterprise workflows requiring consistent style and tone.
Through LLM.API, Palmyra X5 is accessible as a text-only model for prompts and completions.
Palmyra X5 supports a context window of up to 32K tokens via LLM.API.
Palmyra X5 pricing is usage-based per input and output token, with exact rates defined in LLM.API’s pricing documentation.
On LLM.API, Palmyra X5 is optimized for low-latency interactive use, with typical responses in the sub-second to few-second range depending on prompt size.
Specify the model name "writer/palmyra-x5" in your LLM.API request along with your API key and standard completion parameters.
Palmyra X5 emphasizes enterprise safety, controllability, and writing quality, making it competitive with other mid-to-large models for business content generation.
If enabled by LLM.API, Palmyra X5 can be used with the platform’s standardized tool-calling interface similar to other supported models.
Palmyra X5 can hallucinate facts, may be less suitable for code-heavy workloads, and should not be used without human review for critical decisions.
Compare
Olmo 3 32B Think is a 32-billion-parameter open-weight reasoning model from the Allen Institute for AI, optimized for deep chain-of-thought reasoning and complex instruction following. It…
Trinity Large Thinking is Arcee AI’s open-weight, 398–400B-parameter sparse Mixture-of-Experts model focused on advanced reasoning and long-horizon agentic tasks. It is notable for activating only about…
Laguna XS.2 (free) by Poolside is a compact, open‑weight agentic coding model optimized for fast, affordable software engineering workflows, available at no cost via selected providers.…