AI Models with Vision

Vision support is now table stakes for document parsing, screenshot understanding, and multimodal agents. Every model below accepts image input. Where they differ is in how much that capability costs and how large a context they pair it with, which matters a lot when you are feeding in high-resolution pages rather than single thumbnails.

Showing top 25 of 465 matching models · last synced 7 hours ago

Top pick right now: Aya Vision 8B from Cohere. It pairs image input with the lowest output price in this set (N/A/Mtok) and a 16K context window.

# Model Provider Input $/Mtok Output $/Mtok Context Capabilities
8
Auto Router
OpenRouter N/A N/A 2,000,000 vision, tools, reasoning, structured output
19
Free Models Router
OpenRouter Free Free 200,000 vision, tools, reasoning, structured output
20
Seedream 4.5
OpenRouter Free Free 4,096 vision, open
21
FLUX.2 Max
OpenRouter Free Free 46,864 vision
22
FLUX.2 Flex
OpenRouter Free Free 67,344 vision
23
FLUX.2 Pro
OpenRouter Free Free 46,864 vision
24
FLUX.2 Klein 4B
OpenRouter Free Free 40,960 vision, open
25
Nemotron 3 Nano Omni (free)
OpenRouter Free Free 256,000 vision, tools, reasoning, open

Browse other categories

Data synced daily from models.dev. Always verify pricing with the provider before deploying to production.