- Instruction Following
Qwen3.5 397B A17B is a large-scale language model from Qwen with roughly 397 billion parameters, designed for advanced reasoning and multilingual understanding. It targets high-end inference…
Powered by Upstage
Solar Pro 3 is Upstage’s Mixture-of-Experts large language model with 102B total parameters (12B active), a 128K-token context window, and strong extended reasoning and tool-use capabilities.
Output tokens per second · Higher is better
Seconds · Lower is better
USD per 1M tokens (blended) · Lower is better
About the model
Solar Pro 3 is a proprietary Mixture-of-Experts language model from Upstage optimized for efficient, high-quality text generation and reasoning. It is used for complex multi-step reasoning, agentic workflows, and long-context tasks such as document analysis and large-codebase assistance. It also serves enterprise applications that need reliable tool use, structured outputs, and multilingual support focused on Korean with additional English and Japanese coverage. Solar Pro 3 follows earlier Solar-series models such as Solar Pro 2, offering increased parameter scale and improved reasoning performance within the same general model family.
Model capabilities
Generates and edits high-quality text responses across domains, suitable for content creation, SEO workflows, and structured writing tasks.
Processes and reasons over long inputs with a context window up to 128K tokens, supporting document-heavy and retrieval-oriented applications.
Supports tool use and function calling, enabling agentic workflows that interact with external systems and APIs programmatically.
Produces well-structured JSON and schema-conformant outputs, useful for automation pipelines and programmatic integration with downstream systems.
Handles multiple languages with strong performance in Korean and solid English and Japanese support for multilingual applications.
Use cases
Transparent pricing
LLM API offers the lowest cost and latency for Solar Pro 3–class models.
| Provider | Region | Latency | Throughput | Uptime | Input ($/1M) | Output ($/1M) | Context |
|---|---|---|---|---|---|---|---|
| LLM API BEST | Global | 120ms | 80 tps | 99.99% | $0.20 | $0.60 | 256K |
| Upstage | Global | ~250ms | ~40 tps | ~99.9% | ~$0.25 | ~$0.75 | ~128K |
| OpenRouter | Global | ~320ms | ~35 tps | ~99.9% | ~$0.30 | ~$0.90 | ~128K |
| Together AI | US East | ~280ms | ~45 tps | ~99.9% | ~$0.28 | ~$0.85 | ~128K |
| Fireworks AI | US West | ~260ms | ~50 tps | ~99.95% | ~$0.26 | ~$0.80 | ~128K |
Performance benchmarks
| Metric | Solar Pro 3 (Upstage) | GPT-4.1 (OpenAI) | Claude 3.5 Sonnet (Anthropic) |
|---|---|---|---|
| Avg Latency | ~180ms | ~220ms | ~250ms |
| Context Window | 200K | 128K | 200K |
| Input Price ($/1M) | $0.80 | $5.00 | $3.00 |
| Output Price ($/1M) | $4.00 | $15.00 | $15.00 |
| Max Output Tokens | 4K | 4K | 4K |
| Throughput | ~80 tps | ~60 tps | ~50 tps |
| Uptime | 99.9% | 99.9% | 99.9% |
30-day usage via LLM API
Architecture & Integration
One unified API. Every major model. Built-in reliability, cost control, and observability.
Dynamically route each request to the optimal model across providers using latency, cost, and quality signals—without changing your integration or redeploying code.
One endpoint, every model.Define budgets, price ceilings, and routing rules so LLM.API automatically picks the cheapest viable model while preserving output quality and performance SLAs.
Optimize spend by default.Guard against provider outages and rate limits with configurable failover logic that instantly retries on backup models, maintaining uptime without custom error-handling glue.
Resiliency built in.Get per-request traces, latency and cost breakdowns, and structured logs across all providers from a single dashboard and API, ready for alerting and analytics.
One pane of glass.Describe tasks at a higher level—chat, extraction, tools—and let LLM.API select prompts, parameters, and models, so you ship features instead of tuning configs.
Think tasks, not prompts.Send massive batches of jobs through a single API call with concurrency controls, retries, and progress tracking, ideal for backfills, evaluations, and bulk processing.
Scale jobs, not code.Decision guide
FAQ
Solar Pro 3 is a large language model by Upstage optimized for high-quality reasoning, coding, and general-purpose chat via the LLM.API gateway.
Solar Pro 3 is best for complex reasoning, code generation and debugging, multi-step tool use, and production-grade chatbots needing strong instruction following.
Solar Pro 3 supports a long context window suitable for large documents and multi-step conversations; check the LLM.API model card for the exact token limit.
Typical end-to-end latency is on the order of seconds for short prompts, with streaming responses and scalable throughput handled by LLM.API infrastructure.
Solar Pro 3 is a text-only model that accepts text prompts and returns text completions or chat responses.
You can select the upstage/solar-pro-3 model name in the LLM.API completion or chat endpoint, passing your prompt and any temperature or max_tokens parameters.
Solar Pro 3 uses pay-as-you-go, per-token billing; see the LLM.API pricing page for current input and output token rates.
Solar Pro 3 is positioned as a high-quality, cost-efficient general model competitive with other top-tier reasoning and coding LLMs in its price bracket.
Solar Pro 3 can hallucinate, lacks real-time knowledge or browsing, and may underperform on highly specialized domain tasks without careful prompting or grounding.
Fine-tuning support depends on LLM.API capabilities at the time; check the model page for whether custom fine-tunes or adapters are available for Solar Pro 3.
Compare
Qwen3.5 397B A17B is a large-scale language model from Qwen with roughly 397 billion parameters, designed for advanced reasoning and multilingual understanding. It targets high-end inference…
Free Models Router is an OpenRouter meta-model that automatically routes requests to compatible free models, providing no-cost inference across multiple underlying LLMs. It filters candidates based…
GPT-5.3 Chat is an OpenAI conversational large language model designed for general-purpose dialogue and task assistance, with improved reasoning and instruction-following over prior GPT chat models.