AI Models with Vision

Vision support is now table stakes for document parsing, screenshot understanding, and multimodal agents. Every model below accepts image input. Where they differ is in how much that capability costs and how large a context they pair it with, which matters a lot when you are feeding in high-resolution pages rather than single thumbnails.

Showing top 25 of 393 matching models · last synced 1 day ago

Top pick right now: Aya Vision 8B from Cohere. It pairs image input with the lowest output price in this set (N/A/Mtok) and a 16K context window.

# Model Provider Input $/Mtok Output $/Mtok Context Capabilities
8
Auto Router
OpenRouter N/A N/A 2,000,000 vision, tools, reasoning, structured output
18
Free Models Router
OpenRouter Free Free 200,000 vision, tools, reasoning, structured output
19
Seedream 4.5
OpenRouter Free Free 4,096 vision, open
20
FLUX.2 Max
OpenRouter Free Free 46,864 vision
21
FLUX.2 Flex
OpenRouter Free Free 67,344 vision
22
FLUX.2 Pro
OpenRouter Free Free 46,864 vision
23
FLUX.2 Klein 4B
OpenRouter Free Free 40,960 vision, open
24
Nemotron 3 Nano Omni (free)
OpenRouter Free Free 256,000 vision, tools, reasoning, open
25
Nemotron Nano 12B 2 VL (free)
OpenRouter Free Free 128,000 vision, tools, reasoning, open

Browse other categories

Data synced daily from models.dev. Always verify pricing with the provider before deploying to production.