- Text Generation
Sora 2 Pro is an OpenAI model name that has been mentioned publicly, but as of now OpenAI has not released authoritative technical details or documentation…
Powered by Google
Gemini 3.1 Pro Preview is a preview large language model from Google’s Gemini family, offering advanced reasoning and multimodal capabilities for early experimentation and feedback. As a preview model, its behavior and performance may change as Google continues development before general availability.
Output tokens per second · Higher is better
Seconds · Lower is better
USD per 1M tokens (blended) · Lower is better
About the model
Gemini 3.1 Pro Preview is an experimental version of Google’s Gemini 3.1 Pro large language model made available for limited testing. It is primarily used by developers and researchers to explore its capabilities in tasks such as code assistance, structured reasoning, and information retrieval. It is also used to prototype multimodal and agentic applications so Google can refine quality, safety, and performance. It belongs to the Gemini model family, following earlier Gemini 1.x and 2.x generations and the non-preview Gemini Pro variants.
Model capabilities
Performs complex logical reasoning and problem solving, excelling on benchmarks like ARC-AGI-2 and SWE-Bench for difficult tasks.
Understands text, code, images, audio, video, and PDFs within a very long context window for rich cross-modal analysis.
Processes and synthesizes information from large documents and datasets, supporting enterprise knowledge tasks and technical analysis.
Supports code understanding and generation, autonomous software engineering tasks, and tool-assisted code execution workflows.
Handles multiple languages for reading and generation, enabling cross-language understanding and globally-deployed conversational applications.
Use cases
Transparent pricing
LLM API offers the lowest effective cost and latency for Gemini 3.1 Pro–class models.
| Provider | Region | Latency | Throughput | Uptime | Input ($/1M) | Output ($/1M) | Context |
|---|---|---|---|---|---|---|---|
| LLM API BEST | Global | 120ms | 80 tps | 99.99% | $0.30 | $0.60 | 128K |
| Global | ~220ms | ~40 tps | 99.9% | ~$0.50 | ~$1.50 | 128K | |
| Vertex AI (Google Cloud) | US East | ~260ms | ~35 tps | 99.9% | ~$0.55 | ~$1.60 | 128K |
| Fireworks AI | US West | ~200ms | ~50 tps | 99.9% | ~$0.45 | ~$1.40 | 64K |
Performance benchmarks
| Metric | Gemini 3.1 Pro Preview | GPT-4.1 | Claude 3.5 Sonnet |
|---|---|---|---|
| Avg Latency | ~250ms | ~300ms | ~280ms |
| Context Window | 128K | 128K | 200K |
| Input Price ($/1M) | $0.40 | $5.00 | $3.00 |
| Output Price ($/1M) | $1.20 | $15.00 | $15.00 |
| Max Output Tokens | 8K | 4K | 4K |
| Throughput | ~50 tps | ~40 tps | ~35 tps |
| Uptime | 99.9% | 99.9% | 99.9% |
30-day usage via LLM API
Architecture & Integration
One unified API. Every major model. Built-in reliability, cost control, and observability.
Dynamically route each request to the optimal model across providers based on latency, cost, or quality—without changing your integration or redeploying.
One endpoint, every modelControl spend with per-route cost caps, dynamic model downgrades, and usage insights so you ship rich AI features without surprise bills.
Cut spend, keep qualityIf a provider throttles or fails, LLM.API seamlessly retries on backup models, keeping your AI workflows online without custom failover logic.
Resilience by defaultTrace every request across models and providers with rich logs, metrics, and timelines to debug prompts, tune routing, and prove reliability in production.
See every tokenCall higher-level tasks—chat, RAG, tools, moderation—instead of raw models, letting LLM.API manage prompts, memory, and orchestration under a stable interface.
Code to tasks, not modelsProcess thousands of prompts in parallel with batch operations, reducing overhead, smoothing rate limits, and maximizing throughput for large-scale workloads.
Scale from day oneDecision guide
FAQ
Gemini 3.1 Pro Preview is a Google frontier language model optimized for high‑quality reasoning, coding, and general-purpose chat use cases.
Through LLM.API, Gemini 3.1 Pro Preview currently supports text input and output, with image and other modalities exposed as the provider enables them.
Gemini 3.1 Pro Preview is billed on a pay-as-you-go per-token basis, with separate input and output token rates defined by LLM.API.
Gemini 3.1 Pro Preview supports a large context window suitable for multi-thousand token prompts and long conversations, as configured by LLM.API.
Typical latency is comparable to other large frontier models, with first-token times dependent on prompt size and current LLM.API and Google load.
It excels at multi-step reasoning, complex code generation, data analysis, and high-quality natural language generation across many domains.
You select the model name "google/gemini-3.1-pro-preview" (or similar identifier) in LLM.API and send standard chat or completion-style requests.
Gemini 3.1 Pro Preview targets similar advanced reasoning and coding capabilities, but performance, cost, and latency vary by task and provider configuration.
Yes, when enabled in your LLM.API request, Gemini 3.1 Pro Preview can return tokens incrementally for lower perceived latency.
It can hallucinate, may contain training-data biases, and should not be relied on for authoritative legal, medical, or safety-critical decisions.
Compare
Sora 2 Pro is an OpenAI model name that has been mentioned publicly, but as of now OpenAI has not released authoritative technical details or documentation…
GLM 5.1 is Z.ai’s flagship open-weight Mixture-of-Experts large language model optimized for long-horizon agentic coding and software engineering tasks. It is notable for its very large…
Qwen3.6 Plus is Alibaba’s flagship Qwen 3.6 series multimodal reasoning model that offers a very large context window and strong agentic capabilities for complex tasks. It…