Multimodal
Models that take in at least two of image, audio and video alongside text.
- models
- 72
- providers
- 15
Prices and specs checked 2026-09-20
- Google
Gemini 3.8 Flash
Gemini FlashLatest Gemini Flash. Reads text, images, audio and video; 1M context.
MultimodalText + Images + Video + Files + Audio → Text- Audio & video
- Agentic work
- Long documents
- Pricing
- $0.75 / $3.75 per 1M tokens, in / out
- Context window
- 1M context
Ready on endueUse in agent - Google
Gemini 3.5 Flash
Gemini FlashNear-Pro coding and reasoning at Flash cost, with parallel tool use.
MultimodalText + Images + Video + Files + Audio → Text- Coding
- Audio & video
- Agentic work
- Pricing
- $1.5 / $9 per 1M tokens, in / out
- Context window
- 1M context
Ready on endueUse in agent - Google
Gemini 3.6 Flash
Gemini FlashFlash model for coding and web/app development with fewer stray edits.
MultimodalText + Images + Video + Files + Audio → Text- Coding
- Audio & video
- Pricing
- $0.75 / $3.75 per 1M tokens, in / out
- Context window
- 1M context
Ready on endueUse in agent - Google
Gemini 3.5 Flash Lite
Gemini Flash-LiteCheap Gemini for sub-agents that run focused tasks.
MultimodalText + Images + Video + Files + Audio → Text- Low cost
- Audio & video
- Fast replies
- Pricing
- $0.3 / $2.5 per 1M tokens, in / out
- Context window
- 1M context
Ready on endueUse in agent - Google
Gemini 2.5 Pro
Gemini ProGemini with step-by-step thinking for math, science and code.
MultimodalText + Images + Files + Audio + Video → Text- Reasoning
- Audio & video
- Long documents
- Pricing
- $1.25 / $10 per 1M tokens, in / out
- Context window
- 1M context
Ready on endueUse in agent - Google
Gemini 3.1 Flash Lite
Gemini Flash-LiteLow-latency Gemini for high volume. Reads audio, video and PDF.
MultimodalText + Images + Video + Files + Audio → Text- Low cost
- Fast replies
- Audio & video
- Pricing
- $0.25 / $1.5 per 1M tokens, in / out
- Context window
- 1M context
Ready on endueUse in agent - Google
Gemini 2.5 Flash Lite
Gemini Flash-LiteAmong the cheapest multimodal models here, 1M context.
MultimodalText + Images + Files + Audio + Video → Text- Low cost
- Fast replies
- Audio & video
- Pricing
- $0.1 / $0.4 per 1M tokens, in / out
- Context window
- 1M context
Ready on endueUse in agent - Google
Gemma 4 31B
Gemma 4Open-weight dense model with image and video input and function calling.
MultimodalImages + Text + Video → Text- Open weights
- Low cost
- Images
- Pricing
- $0.09 / $0.34 per 1M tokens, in / out
- Context window
- 262K context
Ready on endueUse in agent - Google
Gemma 4 26B
Gemma 4Open-weight MoE close to Gemma 4 31B quality at lower compute.
MultimodalImages + Text + Video → Text- Open weights
- Low cost
- Fast replies
- Pricing
- $0.09 / $0.3 per 1M tokens, in / out
- Context window
- 262K context
Ready on endueUse in agent - Moonshot AI
Kimi K3
KimiLarge open-weight multimodal model for coding and long agent runs.
MultimodalText + Images + Video → Text- Coding
- Agentic work
- Open weights
- Pricing
- $1.7 / $8.5 per 1M tokens, in / out
- Context window
- 1M context
Ready on endueUse in agent - Alibaba Qwen
Qwen3.8 Max
Qwen3.8Largest Qwen, with text, image and video input.
MultimodalText + Images + Video → Text- Reasoning
- Coding
- Audio & video
- Pricing
- $2 / $6 per 1M tokens, in / out
- Context window
- 1M context
Ready on endueUse in agent - Alibaba Qwen
Qwen3.8 Flash
Qwen3.8Cheap multimodal Qwen for charts, documents and long videos.
MultimodalText + Images + Video → Text- Low cost
- Audio & video
- Images
- Pricing
- $0.15 / $0.47 per 1M tokens, in / out
- Context window
- 1M context
Ready on endueUse in agent - Alibaba Qwen
Qwen3.6 Flash
Qwen3.6Fast Qwen with image and video input and a 1M context.
MultimodalText + Images + Video → Text- Fast replies
- Low cost
- Audio & video
- Pricing
- $0.19 / $1.13 per 1M tokens, in / out
- Context window
- 1M context
Ready on endueUse in agent - Z.ai
GLM-5V Turbo
GLM-5VGLM that reads images and video, for vision-based coding.
MultimodalImages + Text + Video → Text- Images
- Coding
- Audio & video
- Pricing
- $1.2 / $4 per 1M tokens, in / out
- Context window
- 203K context
Ready on endueUse in agent - MiniMax
MiniMax M3
MiniMaxMultimodal model with a 1M context for long agent work.
MultimodalText + Images + Video → Text- Audio & video
- Agentic work
- Low cost
- Pricing
- $0.3 / $1.2 per 1M tokens, in / out
- Context window
- 1M context
Ready on endueUse in agent - Meta
Muse Spark 1.3
Muse SparkKeeps track of long tasks. Reads text, images, audio, video and PDF.
MultimodalText + Images + Video + Files + Audio → Text- Agentic work
- Audio & video
- Long documents
- Pricing
- $1.25 / $4.25 per 1M tokens, in / out
- Context window
- 1M context
Ready on endueUse in agent - Meta
Muse Spark 1.1
Muse SparkMultimodal reasoning for agent tasks, 1M context.
MultimodalText + Images + Video + Files + Audio → Text- Agentic work
- Audio & video
- Pricing
- $1.25 / $4.25 per 1M tokens, in / out
- Context window
- 1M context
Ready on endueUse in agent - StepFun
Step 3.7 Flash
StepEfficient MoE with native image and video understanding.
MultimodalText + Images + Video → Text- Audio & video
- Low cost
- Images
- Pricing
- $0.2 / $1.15 per 1M tokens, in / out
- Context window
- 262K context
Ready on endueUse in agent
No model matches these filters.