- Text Generation
Gemini 3.1 Pro Preview is a preview large language model from Google’s Gemini family, offering advanced reasoning and multimodal capabilities for early experimentation and feedback. As…
Powered by Google
Gemini 3.1 Pro Preview Custom Tools is a preview large language model from Google’s Gemini 3.1 Pro line that supports integration with user-defined tools and APIs. It is notable for enabling tailored, tool-augmented workflows while providing the advanced reasoning and multimodal capabilities of the Gemini 3.1 Pro family.
Output tokens per second · Higher is better
Seconds · Lower is better
USD per 1M tokens (blended) · Lower is better
About the model
Gemini 3.1 Pro Preview Custom Tools is a Google Gemini 3.1 Pro–series model variant that allows developers to connect custom tools and APIs for augmented reasoning and action-taking. It is mainly used to build applications where the model can call external services, trigger workflows, or fetch live data through developer-defined tools. It also supports use cases that require more controlled, domain-specific behavior by combining Gemini’s core capabilities with bespoke tool integrations. It belongs to the Gemini 3.1 Pro family of models, which succeeds earlier Gemini Pro generations.
Model capabilities
Engages in multi-turn dialogue, answering questions, following instructions, and adapting responses to user context and prior conversation.
Accepts image inputs to identify objects, infer context, and answer questions about visual content when enabled by the provider.
Translates text between multiple languages, supporting cross-lingual understanding and communication in both short prompts and longer documents.
Orchestrates custom tools and APIs, interpreting user requests and invoking external functions to retrieve data or perform actions.
Reads and extracts textual content from documents or screenshots, enabling question answering and information retrieval over provided materials.
Use cases
Transparent pricing
LLM API offers the lowest prices and best performance for Gemini 3.1 Pro–class models.
| Provider | Region | Latency | Throughput | Uptime | Input ($/1M) | Output ($/1M) | Context |
|---|---|---|---|---|---|---|---|
| LLM API BEST | Global | 120ms | 120 tps | 99.99% | $0.20 | $0.60 | 200K |
| Global | ~350ms | ~60 tps | 99.9% | ~$0.30 | ~$0.90 | 128K | |
| Vertex AI (Google Cloud) | US Central | ~380ms | ~55 tps | 99.9% | ~$0.32 | ~$0.96 | 128K |
| Anthropic | US East | ~300ms | ~70 tps | 99.9% | ~$0.40 | ~$1.20 | 200K |
| OpenAI | Global | ~250ms | ~80 tps | 99.9% | ~$0.35 | ~$1.05 | 128K |
Performance benchmarks
| Metric | Gemini 3.1 Pro Preview Custom Tools | GPT-4.1 | Claude 3.5 Sonnet |
|---|---|---|---|
| Avg Latency | ~220ms | ~250ms | ~230ms |
| Context Window | 128K | 128K | 200K |
| Input Price ($/1M) | $0.80 | $5.00 | $3.00 |
| Output Price ($/1M) | $2.40 | $15.00 | $15.00 |
| Max Output Tokens | 4K | 4K | 4K |
| Throughput | 60 tps | 50 tps | 45 tps |
| Uptime | 99.9% | 99.9% | 99.9% |
30-day usage via LLM API
Architecture & Integration
One unified API. Every major model. Built-in reliability, cost control, and observability.
Automatically direct each request to the best model across providers based on performance, latency, and cost—without changing your integration or redeploying code.
One endpoint, every modelOptimize spend by mixing premium and budget models per request, with guardrails and policies that keep your AI bill predictable at scale.
Lower cost, same outputDesign automatic provider and model failover so your workloads survive rate limits, outages, and timeouts—without complex custom retry logic.
No single point of failureTrace every request across providers with unified logs, metrics, and latency analysis so you can debug faster and continuously tune model performance.
See every token flowDescribe what you want—chat, extraction, classification, tools—and let LLM.API pick the right model, parameters, and prompts for the job.
Think tasks, not modelsProcess millions of requests efficiently with queued, retried, and parallelized batch jobs that respect provider limits and maximize throughput automatically.
Scale from 10 to millionsDecision guide
FAQ
Gemini 3.1 Pro Preview Custom Tools is a Google multimodal large language model with support for custom tool calling and integration via LLM.API.
It is best for complex reasoning, multi-step tool-using workflows, and mixed text-plus-structured data applications where external tools or APIs must be orchestrated.
LLM.API meters usage per token for input and output, with specific Gemini 3.1 Pro Preview Custom Tools rates shown in the LLM.API pricing section.
Gemini 3.1 Pro Preview Custom Tools supports a multi-thousand token context window; check the model card on LLM.API for the exact current limit.
Typical latency is similar to other large frontier models, and depends on prompt size, output length, and any custom tool calls executed.
Through LLM.API it supports text input and output, and can interact with external tools; additional modalities depend on LLM.API’s enabled interfaces.
Specify the model name "google/gemini-3.1-pro-preview-custom-tools" in your LLM.API request, include your API key, and send standard chat completion payloads.
Compared to standard Gemini Pro variants, it emphasizes robust tool-calling behavior and orchestration over pure generative throughput.
It can still hallucinate, miscall tools, and may not reflect real-time information; you must validate critical outputs and handle tool errors gracefully.
If LLM.API enables streaming for this model, you can request streamed responses using the standard streaming flags in the API.
Compare
Gemini 3.1 Pro Preview is a preview large language model from Google’s Gemini family, offering advanced reasoning and multimodal capabilities for early experimentation and feedback. As…
Google Gemini Pro Latest is the most recent Pro-tier model in Google’s Gemini family of large multimodal models, optimized for complex reasoning and agentic tasks across…
Granite 4.0 Micro is a 3B-parameter dense language model from IBM’s Granite 4.0 family, optimized for low-latency, cost-efficient workloads and local or edge deployment.