- Instruction Following
GPT-5.5 Pro is an OpenAI model name that has been mentioned publicly but has not been formally documented or specified by OpenAI as of now. Reliable…
Powered by StepFun
Step 3.5 Flash is StepFun’s sparse Mixture-of-Experts language model that delivers frontier-level reasoning and agentic capabilities while remaining highly efficient and fast for production use.
Output tokens per second · Higher is better
Seconds · Lower is better
USD per 1M tokens (blended) · Lower is better
About the model
Step 3.5 Flash is a sparse Mixture-of-Experts large language model from StepFun designed to combine frontier-level reasoning with high-throughput, low-latency inference. It is mainly used for complex reasoning tasks, code generation, and agentic workflows that benefit from its long context window and efficient token usage. The model is also applied to natural language processing, data analysis, and long-document or codebase understanding where cost and speed are critical. It belongs to StepFun’s Step family of models and builds on the Step 3.x architecture and research line.
Model capabilities
Handles general-purpose dialogue, explanations, brainstorming, and question answering, optimized for fast, low-latency text generation and responses.
Generates well-formed JSON and structured text suitable for programmatic consumption, including configuration data, responses, and tool outputs.
Translates between multiple natural languages, leveraging its large context and reasoning capabilities to preserve meaning and style.
Writes and edits code, explains snippets, and assists with debugging across common programming languages, tuned for agentic coding workflows.
Performs reasoning and synthesis over very long texts using its 256K-token context window for documents, logs, or multi-step analyses.
Use cases
Transparent pricing
LLM API offers the lowest cost and fastest access to Step 3.5 Flash–class models.
| Provider | Region | Latency | Throughput | Uptime | Input ($/1M) | Output ($/1M) | Context |
|---|---|---|---|---|---|---|---|
| LLM API BEST | Global | 110ms | 80 tps | 99.99% | $0.08 | $0.24 | 256K |
| StepFun | Global | ~180ms | ~40 tps | ~99.9% | ~$0.10 | ~$0.30 | ~128K |
| OpenAI-compatible gateway | US East | ~220ms | ~35 tps | ~99.9% | ~$0.12 | ~$0.36 | ~128K |
| Cloud Hyperscaler A | EU West | ~250ms | ~30 tps | ~99.95% | ~$0.14 | ~$0.40 | ~128K |
Performance benchmarks
| Metric | Step 3.5 Flash (StepFun) | GPT-4.1 mini (OpenAI) | Claude 3 Haiku (Anthropic) |
|---|---|---|---|
| Avg Latency | ~180ms | ~200ms | ~220ms |
| Context Window | 128K | 128K | 200K |
| Input Price ($/1M) | $0.10 | $0.15 | $0.25 |
| Output Price ($/1M) | $0.40 | $0.60 | $0.80 |
| Max Output Tokens | 4K | 4K | 4K |
| Throughput | ~120 tps | ~100 tps | ~90 tps |
| Uptime | 99.9% | 99.9% | 99.9% |
30-day usage via LLM API
Architecture & Integration
One unified API. Every major model. Built-in reliability, cost control, and observability.
Automatically route each request to the optimal model based on latency, cost, and quality—without changing your integration or redeploying services.
One endpoint, every model.Control spend with dynamic model selection, price ceilings, and transparent usage metrics so you can scale AI features without runaway cloud bills.
Max performance, minimal spend.Define automatic failover chains across providers so timeouts, rate limits, or outages don’t break your workloads or user experience.
Stay online, even upstream.Trace every call across providers with logs, metrics, and latency breakdowns to debug faster and optimize prompt, model, and routing decisions.
See every token’s journey.Describe high-level tasks once and let LLM.API handle tool calls, multi-step flows, and provider choices for consistent results across environments.
Think tasks, not models.Send thousands of operations in a single request with controlled concurrency, retries, and deduplication to power large-scale inference pipelines efficiently.
Ship at batch scale.Decision guide
FAQ
Step 3.5 Flash is a fast, cost-efficient StepFun language model for general-purpose text generation and reasoning, accessible through the LLM.API gateway.
Step 3.5 Flash is best for high-throughput tasks like chatbots, agents, data processing, and lightweight reasoning where low latency and low cost matter.
Step 3.5 Flash supports a context window of up to 32K tokens, including both prompt and completion tokens.
Step 3.5 Flash is optimized for low latency, typically returning first tokens within a few hundred milliseconds depending on prompt size and load.
Step 3.5 Flash is a text-only model that accepts textual prompts and returns textual completions.
Specify the provider as "StepFun" and the model name "step-3.5-flash" in your LLM.API request, sending standard chat or completion payloads.
On LLM.API, Step 3.5 Flash is billed per 1,000 tokens of prompt and completion; check the LLM.API pricing page for exact rates.
Compared to larger StepFun models, Step 3.5 Flash is cheaper and faster but generally less capable on complex reasoning and intricate long-context tasks.
Yes, Step 3.5 Flash can stream tokens incrementally when you enable streaming mode in your LLM.API request.
Step 3.5 Flash may struggle with very long multi-step reasoning, domain-expert tasks, strict factual accuracy, and does not access external tools or the internet.
Compare
GPT-5.5 Pro is an OpenAI model name that has been mentioned publicly but has not been formally documented or specified by OpenAI as of now. Reliable…
Qwen3.5-Flash is a hosted, production-oriented large language model from Qwen, optimized for fast, efficient text and vision-language generation. It corresponds to the Qwen3.5-35B-A3B model and offers…
Gemma 4 31B (free) is a large language model from Google’s Gemma 4 family, offered in a 31-billion-parameter configuration with free access in some platforms. It…