- Text Classification
gpt-oss-safeguard-20b is an OpenAI model name that appears to reference a 20-billion-parameter, safety-focused open-source-style GPT variant, but OpenAI has not publicly released authoritative technical details about…
Powered by Perceptron
Perceptron Mk1 is Perceptron's highest‑quality proprietary vision‑language model focused on video understanding and embodied visual reasoning. It is notable for combining multimodal inputs (text, images, video) with long‑context reasoning and structured visual outputs for production use.
Output tokens per second · Higher is better
Seconds · Lower is better
USD per 1M tokens (blended) · Lower is better
About the model
Perceptron Mk1 is a closed-source vision-language model from Perceptron designed for image and video understanding, OCR, object detection, document parsing, and embodied visual reasoning. It is mainly used for tasks like video question answering, summarization, event and temporal segment detection, and detailed scene analysis across robotics, autonomous systems, surveillance, and AR applications. It also supports structured outputs such as spatial annotations (points, boxes, polygons) and temporal clips for production APIs that need precise localization and grounding in complex visual data. Perceptron Mk1 is the first and highest-quality model in the proprietary Perceptron Mk family of multimodal vision and reasoning models.
Model capabilities
Engages in multi-turn text conversations, following instructions, answering questions, and adapting responses to user context and intent.
Interprets visual content in images, identifying objects, scenes, relationships, and other salient details to support downstream reasoning tasks.
Translates written text between multiple languages while preserving meaning, tone, and basic formatting across diverse domains and styles.
Processes user interface or webpage-like content, enabling reasoning about layouts, elements, and on-screen information for digital tasks.
Extracts machine-readable text from images, screenshots, or scanned documents to enable search, editing, or further automated processing.
Use cases
Transparent pricing
LLM API offers the lowest cost and best performance for Perceptron‑class models.
| Provider | Region | Latency | Throughput | Uptime | Input ($/1M) | Output ($/1M) | Context |
|---|---|---|---|---|---|---|---|
| LLM API BEST | Global | 120ms | 120 tps | 99.99% | $0.20 | $0.20 | 64K tokens |
| Perceptron | US East | ~220ms | ~60 tps | ~99.9% | ~$0.45 | ~$0.45 | ~32K tokens |
| OpenAI | Global | ~250ms | ~80 tps | ~99.9% | ~$0.50 | ~$0.50 | ~32K tokens |
| Anthropic | US West | ~260ms | ~70 tps | ~99.9% | ~$0.55 | ~$0.55 | ~200K tokens |
| Azure AI | EU West | ~240ms | ~75 tps | ~99.95% | ~$0.52 | ~$0.52 | ~32K tokens |
Performance benchmarks
| Metric | Perceptron Mk1 | OpenAI GPT-4.1 Mini | Anthropic Claude 3 Haiku |
|---|---|---|---|
| Avg Latency | ~180ms | ~200ms | ~220ms |
| Context Window | 128K | 128K | 200K |
| Input Price ($/1M) | $0.15 | $0.15 | $0.25 |
| Output Price ($/1M) | $0.60 | $0.60 | $0.80 |
| Max Output Tokens | 8K | 4K | 8K |
| Throughput | ~120 tps | ~100 tps | ~90 tps |
| Uptime | 99.9% | 99.9% | 99.9% |
30-day usage via LLM API
Architecture & Integration
One unified API. Every major model. Built-in reliability, cost control, and observability.
Dynamically route each request to the optimal model across providers using policies, constraints, and real-time signals—no client rewrites or vendor-specific logic needed.
One endpoint, any modelControl spend with price-based routing, per-project limits, and smart downgrades—keeping quality high while preventing surprise bills in production workloads.
Optimize quality per dollarDefine graceful failover chains across models and providers so timeouts, rate limits, or outages degrade smoothly instead of breaking user experiences.
Never ship single points of failureTrace every request across providers with unified logs, metrics, and latency breakdowns to debug issues fast and tune prompts and routing policies confidently.
See every token, everywhereDescribe tasks like chat, embeddings, tools, or reranking once and let LLM.API translate them into each provider’s API, schema, and capabilities.
Standard tasks, many backendsShip large workloads as batches with concurrency control, retries, and partial-failure handling built in—cutting overhead and maximizing throughput across providers.
Batch at scale, safelyDecision guide
FAQ
Perceptron Mk1 is a large language model from Perceptron focused on fast, low-cost text generation for general software and product development use cases.
Perceptron Mk1 is best for code generation, refactoring, technical writing, and structured data transformations where low latency and high throughput matter.
Perceptron Mk1 uses LLM.API’s unified per-token billing; check your LLM.API pricing dashboard for current input and output token rates.
Perceptron Mk1 supports a context window defined by the LLM.API integration; refer to the model card for the latest maximum token limit.
Perceptron Mk1 is optimized for low latency, typically returning first tokens quickly enough for interactive applications when used via LLM.API.
Perceptron Mk1 is a text-only model, accepting text prompts and returning text completions via the LLM.API interface.
Specify the Perceptron Mk1 model name in your LLM.API completion or chat endpoint request, using the same authentication and parameters as other models.
Perceptron Mk1 typically trades some reasoning depth for higher throughput and lower cost compared with larger, more capable general-purpose models.
Perceptron Mk1 can struggle with complex multi-step reasoning, domain-expert answers, and may occasionally produce incorrect or hallucinated information.
Yes, when enabled in LLM.API, Perceptron Mk1 supports streaming token responses to improve perceived latency for end users.
Compare
gpt-oss-safeguard-20b is an OpenAI model name that appears to reference a 20-billion-parameter, safety-focused open-source-style GPT variant, but OpenAI has not publicly released authoritative technical details about…