perception-0-fast
Efficient visual reasoning at scale. Ideal for simple queries and visual tasks with high throughput.
perception-0-thinking
Our most powerful vision model for complex and multi-step reasoning problems on video.
SOTA on LVBench
perception-0-fast
Intelligent visual understanding at scale. This model offers state-of-the-art VLM performance, augmented for speed and efficiency. Available for images, videos, and collections of either. Our system indexes your inputs once and enables reliable and faster querying at scale.
Best for: Visual Q&A, image classification, video summaries, large collection queries, low-latency or synchronous applications.
perception-0-thinking
The worlds most powerful vision model for real world tasks. A self-improving visual reasoning engine that can solve any visual task, as well or better than a human in any domain, through: [1] powerful visual analysis primitives & tools, and [2] a system that autonomously builds a fully formed expert ‘model’ for any domain. Reach out to[email protected] for API access.
Best for: High recall and precision, long context, multi-step temporal/spatial reasoning, counting, domains with complex knowledge required.