Image
Models that can read images you send. endue does not offer image generation models yet.
- models
- 72
- providers
- 15
Prices and specs checked 2026-09-20
- xAI
Grok 4.6
Grok 4xAI’s top Grok for coding, STEM and knowledge work.
TextText + Images + Files → Text- Coding
- Reasoning
- Long documents
- Pricing
- $2 / $6 per 1M tokens, in / out
- Context window
- 500K context
Ready on endueUse in agent - xAI
Grok 4.3
Grok 4Reasoning model that follows instructions closely, with a 1M context. endue’s default.
TextText + Images + Files → Text- Agentic work
- Reasoning
- Long documents
- Pricing
- $1.25 / $2.5 per 1M tokens, in / out
- Context window
- 1M context
Ready on endueUse in agent - xAI
Grok 4.20
Grok 4Fast tool calling with a 2M context for very long inputs.
TextText + Images + Files → Text- Agentic work
- Long documents
- Fast replies
- Pricing
- $1.25 / $2.5 per 1M tokens, in / out
- Context window
- 2M context
Ready on endueUse in agent - xAI
Grok 4.5
Grok 4Balanced Grok for coding and knowledge work, 500K context.
TextText + Images + Files → Text- Coding
- Reasoning
- Pricing
- $2 / $6 per 1M tokens, in / out
- Context window
- 500K context
Ready on endueUse in agent - xAI
Grok Build 0.1
Grok BuildFast coding model tuned for interactive software engineering.
TextText + Images + Files → Text- Coding
- Fast replies
- Low cost
- Pricing
- $1 / $2 per 1M tokens, in / out
- Context window
- 256K context
Ready on endueUse in agent - OpenAI
GPT-6 Astra
GPT-6OpenAI’s flagship for long, demanding work: analysis, research, large codebases.
TextFiles + Images + Text → Text- Reasoning
- Coding
- Long documents
- Pricing
- $10 / $50 per 1M tokens, in / out
- Context window
- 1.1M context
Ready on endueUse in agent - OpenAI
GPT-6 Astra Pro
GPT-6GPT-6 Astra with deeper reasoning per answer. Same price, slower replies.
TextFiles + Images + Text → Text- Reasoning
- Web research
- Pricing
- $10 / $50 per 1M tokens, in / out
- Context window
- 1.1M context
Ready on endueUse in agent - OpenAI
GPT-5.6 Sol
GPT-5.6Top GPT-5.6 tier, strong at multi-step and command-line coding.
TextFiles + Images + Text → Text- Coding
- Agentic work
- Reasoning
- Pricing
- $2 / $10 per 1M tokens, in / out
- Context window
- 1.1M context
Ready on endueUse in agent - OpenAI
GPT-5.6 Terra
GPT-5.6Middle GPT-5.6 tier for everyday coding and agent tasks.
TextFiles + Images + Text → Text- Coding
- Agentic work
- Pricing
- $2 / $12 per 1M tokens, in / out
- Context window
- 1.1M context
Ready on endueUse in agent - OpenAI
GPT-5.6 Luna
GPT-5.6Cheap, fast GPT-5.6 tier for chat, classification and high volume.
TextFiles + Images + Text → Text- Fast replies
- Low cost
- Long documents
- Pricing
- $0.2 / $1.2 per 1M tokens, in / out
- Context window
- 1.1M context
Ready on endueUse in agent - OpenAI
GPT-5.5
GPT-5Reliable reasoning for professional workloads, 1M context.
TextFiles + Images + Text → Text- Reasoning
- Long documents
- Pricing
- $5 / $30 per 1M tokens, in / out
- Context window
- 1.1M context
Ready on endueUse in agent - OpenAI
GPT-5.4
GPT-5Codex and GPT merged into one model, with a 1M context.
TextText + Images + Files → Text- Coding
- Reasoning
- Long documents
- Pricing
- $2.5 / $15 per 1M tokens, in / out
- Context window
- 1.1M context
Ready on endueUse in agent - OpenAI
GPT-5.4 Mini
GPT-5Smaller GPT-5.4 for high-throughput reasoning and coding.
TextFiles + Images + Text → Text- Fast replies
- Coding
- Pricing
- $0.75 / $4.5 per 1M tokens, in / out
- Context window
- 400K context
Ready on endueUse in agent - OpenAI
GPT-5.4 Nano
GPT-5Lightest GPT-5.4 for latency-sensitive, high-volume tasks.
TextFiles + Images + Text → Text- Fast replies
- Low cost
- Pricing
- $0.2 / $1.25 per 1M tokens, in / out
- Context window
- 400K context
Ready on endueUse in agent - OpenAI
GPT-5.3 Codex
GPT-5 CodexOpenAI’s agentic coding model for long software tasks.
TextText + Images + Files → Text- Coding
- Agentic work
- Pricing
- $1.75 / $14 per 1M tokens, in / out
- Context window
- 400K context
Ready on endueUse in agent - OpenAI
GPT-5
GPT-5Step-by-step reasoning with careful instruction following.
TextText + Images + Files → Text- Reasoning
- Coding
- Pricing
- $1.25 / $10 per 1M tokens, in / out
- Context window
- 400K context
Ready on endueUse in agent - OpenAI
GPT-4o mini
GPT-4oSmall non-reasoning model for quick, inexpensive replies.
TextText + Images + Files → Text- Fast replies
- Low cost
- Pricing
- $0.15 / $0.6 per 1M tokens, in / out
- Context window
- 128K context
Ready on endueUse in agent - Anthropic
Claude Opus 5
Claude OpusAnthropic’s flagship for code review, bug finding and long agent runs.
TextText + Images + Files → Text- Coding
- Agentic work
- Reasoning
- Pricing
- $5 / $25 per 1M tokens, in / out
- Context window
- 1M context
Ready on endueUse in agent - Anthropic
Claude Fable 5.1
Claude FableFable 5 improved for long refactors, front-end work and knowledge tasks.
TextText + Images + Files → Text- Coding
- Agentic work
- Long documents
- Pricing
- $10 / $50 per 1M tokens, in / out
- Context window
- 1M context
Ready on endueUse in agent - Anthropic
Claude Fable 5
Claude FableBuilt for autonomous knowledge work and coding.
TextText + Images + Files → Text- Agentic work
- Coding
- Pricing
- $10 / $50 per 1M tokens, in / out
- Context window
- 1M context
Ready on endueUse in agent - Anthropic
Claude Opus 4.8
Claude OpusLatest Opus 4 release, 1M context.
TextText + Images + Files → Text- Coding
- Reasoning
- Long documents
- Pricing
- $5 / $25 per 1M tokens, in / out
- Context window
- 1M context
Ready on endueUse in agent - Anthropic
Claude Sonnet 5
Claude SonnetStrong coding and agent work at a lower price than Opus.
TextText + Images + Files → Text- Coding
- Agentic work
- Pricing
- $2 / $10 per 1M tokens, in / out
- Context window
- 1M context
Ready on endueUse in agent - Anthropic
Claude Sonnet 4.5
Claude SonnetSonnet tuned for real-world agents and coding workflows.
TextText + Images + Files → Text- Coding
- Agentic work
- Pricing
- $3 / $15 per 1M tokens, in / out
- Context window
- 1M context
Ready on endueUse in agent - Anthropic
Claude Haiku 4.5
Claude HaikuFastest Claude, near Sonnet 4 quality at lower cost.
TextText + Images + Files → Text- Fast replies
- Low cost
- Pricing
- $1 / $5 per 1M tokens, in / out
- Context window
- 200K context
Ready on endueUse in agent - Anthropic
Claude Opus 4.7
Claude OpusOpus for long-running, asynchronous agents.
TextText + Images + Files → Text- Agentic work
- Coding
- Pricing
- $5 / $25 per 1M tokens, in / out
- Context window
- 1M context
Ready on endueUse in agent - Anthropic
Claude Opus 4.6
Claude OpusOpus for agents that run whole workflows, not single prompts.
TextText + Images + Files → Text- Agentic work
- Coding
- Pricing
- $5 / $25 per 1M tokens, in / out
- Context window
- 1M context
Ready on endueUse in agent - Anthropic
Claude Sonnet 4.6
Claude SonnetIterative development and navigating large codebases.
TextText + Images + Files → Text- Coding
- Agentic work
- Pricing
- $3 / $15 per 1M tokens, in / out
- Context window
- 1M context
Ready on endueUse in agent - Anthropic
Claude Opus 4.5
Claude OpusOpus for complex software engineering and computer use, 200K context.
TextFiles + Images + Text → Text- Coding
- Agentic work
- Pricing
- $5 / $25 per 1M tokens, in / out
- Context window
- 200K context
Ready on endueUse in agent - Anthropic
Claude Opus 4.1
Claude OpusEarlier Opus. The most expensive Claude here; newer Opus models cost less.
TextImages + Text + Files → Text- Coding
- Reasoning
- Pricing
- $15 / $75 per 1M tokens, in / out
- Context window
- 200K context
Ready on endueUse in agent - Google
Gemini 3.8 Flash
Gemini FlashLatest Gemini Flash. Reads text, images, audio and video; 1M context.
MultimodalText + Images + Video + Files + Audio → Text- Audio & video
- Agentic work
- Long documents
- Pricing
- $0.75 / $3.75 per 1M tokens, in / out
- Context window
- 1M context
Ready on endueUse in agent - Google
Gemini 3.5 Flash
Gemini FlashNear-Pro coding and reasoning at Flash cost, with parallel tool use.
MultimodalText + Images + Video + Files + Audio → Text- Coding
- Audio & video
- Agentic work
- Pricing
- $1.5 / $9 per 1M tokens, in / out
- Context window
- 1M context
Ready on endueUse in agent - Google
Gemini 3.6 Flash
Gemini FlashFlash model for coding and web/app development with fewer stray edits.
MultimodalText + Images + Video + Files + Audio → Text- Coding
- Audio & video
- Pricing
- $0.75 / $3.75 per 1M tokens, in / out
- Context window
- 1M context
Ready on endueUse in agent - Google
Gemini 3.5 Flash Lite
Gemini Flash-LiteCheap Gemini for sub-agents that run focused tasks.
MultimodalText + Images + Video + Files + Audio → Text- Low cost
- Audio & video
- Fast replies
- Pricing
- $0.3 / $2.5 per 1M tokens, in / out
- Context window
- 1M context
Ready on endueUse in agent - Google
Gemini 2.5 Pro
Gemini ProGemini with step-by-step thinking for math, science and code.
MultimodalText + Images + Files + Audio + Video → Text- Reasoning
- Audio & video
- Long documents
- Pricing
- $1.25 / $10 per 1M tokens, in / out
- Context window
- 1M context
Ready on endueUse in agent - Google
Gemini 3.1 Flash Lite
Gemini Flash-LiteLow-latency Gemini for high volume. Reads audio, video and PDF.
MultimodalText + Images + Video + Files + Audio → Text- Low cost
- Fast replies
- Audio & video
- Pricing
- $0.25 / $1.5 per 1M tokens, in / out
- Context window
- 1M context
Ready on endueUse in agent - Google
Gemini 2.5 Flash Lite
Gemini Flash-LiteAmong the cheapest multimodal models here, 1M context.
MultimodalText + Images + Files + Audio + Video → Text- Low cost
- Fast replies
- Audio & video
- Pricing
- $0.1 / $0.4 per 1M tokens, in / out
- Context window
- 1M context
Ready on endueUse in agent - Google
Gemma 4 31B
Gemma 4Open-weight dense model with image and video input and function calling.
MultimodalImages + Text + Video → Text- Open weights
- Low cost
- Images
- Pricing
- $0.09 / $0.34 per 1M tokens, in / out
- Context window
- 262K context
Ready on endueUse in agent - Google
Gemma 4 26B
Gemma 4Open-weight MoE close to Gemma 4 31B quality at lower compute.
MultimodalImages + Text + Video → Text- Open weights
- Low cost
- Fast replies
- Pricing
- $0.09 / $0.3 per 1M tokens, in / out
- Context window
- 262K context
Ready on endueUse in agent - Moonshot AI
Kimi K3
KimiLarge open-weight multimodal model for coding and long agent runs.
MultimodalText + Images + Video → Text- Coding
- Agentic work
- Open weights
- Pricing
- $1.7 / $8.5 per 1M tokens, in / out
- Context window
- 1M context
Ready on endueUse in agent - Moonshot AI
Kimi K2.7 Code
KimiCoding-focused Kimi for end-to-end programming over long contexts.
TextText + Images → Text- Coding
- Agentic work
- Pricing
- $0.71 / $3.21 per 1M tokens, in / out
- Context window
- 262K context
Ready on endueUse in agent - Moonshot AI
Kimi K2.6
KimiLong coding tasks, UI generation from code and multi-agent setups.
TextText + Images → Text- Coding
- Agentic work
- Images
- Pricing
- $0.95 / $4 per 1M tokens, in / out
- Context window
- 262K context
Ready on endueUse in agent - DeepSeek
DeepSeek V4.1 Flash
DeepSeek V4Low-cost DeepSeek with image input and a 1M context.
TextText + Images → Text- Low cost
- Long documents
- Images
- Pricing
- $0.15 / $0.6 per 1M tokens, in / out
- Context window
- 1M context
Ready on endueUse in agent - Alibaba Qwen
Qwen3.8 Max
Qwen3.8Largest Qwen, with text, image and video input.
MultimodalText + Images + Video → Text- Reasoning
- Coding
- Audio & video
- Pricing
- $2 / $6 per 1M tokens, in / out
- Context window
- 1M context
Ready on endueUse in agent - Alibaba Qwen
Qwen3.8 Flash
Qwen3.8Cheap multimodal Qwen for charts, documents and long videos.
MultimodalText + Images + Video → Text- Low cost
- Audio & video
- Images
- Pricing
- $0.15 / $0.47 per 1M tokens, in / out
- Context window
- 1M context
Ready on endueUse in agent - Alibaba Qwen
Qwen3.7 Plus
Qwen3.7Cost-effective Qwen3.7 with image input.
TextText + Images → Text- Low cost
- Images
- Pricing
- $0.32 / $1.28 per 1M tokens, in / out
- Context window
- 1M context
Ready on endueUse in agent - Alibaba Qwen
Qwen3.6 Flash
Qwen3.6Fast Qwen with image and video input and a 1M context.
MultimodalText + Images + Video → Text- Fast replies
- Low cost
- Audio & video
- Pricing
- $0.19 / $1.13 per 1M tokens, in / out
- Context window
- 1M context
Ready on endueUse in agent - Z.ai
GLM-5V Turbo
GLM-5VGLM that reads images and video, for vision-based coding.
MultimodalImages + Text + Video → Text- Images
- Coding
- Audio & video
- Pricing
- $1.2 / $4 per 1M tokens, in / out
- Context window
- 203K context
Ready on endueUse in agent - MiniMax
MiniMax M3
MiniMaxMultimodal model with a 1M context for long agent work.
MultimodalText + Images + Video → Text- Audio & video
- Agentic work
- Low cost
- Pricing
- $0.3 / $1.2 per 1M tokens, in / out
- Context window
- 1M context
Ready on endueUse in agent - Mistral AI
Mistral Medium 3.5
Mistral MediumDense 128B model for agent workflows and coding.
TextText + Images + Files → Text- Agentic work
- Coding
- Pricing
- $1.5 / $7.5 per 1M tokens, in / out
- Context window
- 262K context
Ready on endueUse in agent - Mistral AI
Mistral Small 2603
Mistral SmallMistral Small 4: reasoning and image input in one cheap model.
TextText + Images → Text- Low cost
- Reasoning
- Images
- Pricing
- $0.15 / $0.6 per 1M tokens, in / out
- Context window
- 262K context
Ready on endueUse in agent - Mistral AI
Ministral 14B
MinistralEfficient 14B with image input at a flat low price.
TextText + Images → Text- Low cost
- Images
- Fast replies
- Pricing
- $0.2 / $0.2 per 1M tokens, in / out
- Context window
- 262K context
Ready on endueUse in agent - Meta
Muse Spark 1.3
Muse SparkKeeps track of long tasks. Reads text, images, audio, video and PDF.
MultimodalText + Images + Video + Files + Audio → Text- Agentic work
- Audio & video
- Long documents
- Pricing
- $1.25 / $4.25 per 1M tokens, in / out
- Context window
- 1M context
Ready on endueUse in agent - Meta
Muse Spark 1.1
Muse SparkMultimodal reasoning for agent tasks, 1M context.
MultimodalText + Images + Video + Files + Audio → Text- Agentic work
- Audio & video
- Pricing
- $1.25 / $4.25 per 1M tokens, in / out
- Context window
- 1M context
Ready on endueUse in agent - StepFun
Step 3.7 Flash
StepEfficient MoE with native image and video understanding.
MultimodalText + Images + Video → Text- Audio & video
- Low cost
- Images
- Pricing
- $0.2 / $1.15 per 1M tokens, in / out
- Context window
- 262K context
Ready on endueUse in agent - Perplexity
Sonar Pro Search
SonarSearches the web and reasons over results. Cannot call tools.
TextText + Images → Text- Web research
- Pricing
- $3 / $15 per 1M tokens, in / out
- Context window
- 200K context
Ready on endueUse in agent
No model matches these filters.