Audio
Models that can listen to audio you send. endue does not offer speech or music generation yet.
- models
- 72
- providers
- 15
Prices and specs checked 2026-09-20
- Google
Gemini 3.8 Flash
Gemini FlashLatest Gemini Flash. Reads text, images, audio and video; 1M context.
MultimodalText + Images + Video + Files + Audio → Text- Audio & video
- Agentic work
- Long documents
- Pricing
- $0.75 / $3.75 per 1M tokens, in / out
- Context window
- 1M context
Ready on endueUse in agent - Google
Gemini 3.5 Flash
Gemini FlashNear-Pro coding and reasoning at Flash cost, with parallel tool use.
MultimodalText + Images + Video + Files + Audio → Text- Coding
- Audio & video
- Agentic work
- Pricing
- $1.5 / $9 per 1M tokens, in / out
- Context window
- 1M context
Ready on endueUse in agent - Google
Gemini 3.6 Flash
Gemini FlashFlash model for coding and web/app development with fewer stray edits.
MultimodalText + Images + Video + Files + Audio → Text- Coding
- Audio & video
- Pricing
- $0.75 / $3.75 per 1M tokens, in / out
- Context window
- 1M context
Ready on endueUse in agent - Google
Gemini 3.5 Flash Lite
Gemini Flash-LiteCheap Gemini for sub-agents that run focused tasks.
MultimodalText + Images + Video + Files + Audio → Text- Low cost
- Audio & video
- Fast replies
- Pricing
- $0.3 / $2.5 per 1M tokens, in / out
- Context window
- 1M context
Ready on endueUse in agent - Google
Gemini 2.5 Pro
Gemini ProGemini with step-by-step thinking for math, science and code.
MultimodalText + Images + Files + Audio + Video → Text- Reasoning
- Audio & video
- Long documents
- Pricing
- $1.25 / $10 per 1M tokens, in / out
- Context window
- 1M context
Ready on endueUse in agent - Google
Gemini 3.1 Flash Lite
Gemini Flash-LiteLow-latency Gemini for high volume. Reads audio, video and PDF.
MultimodalText + Images + Video + Files + Audio → Text- Low cost
- Fast replies
- Audio & video
- Pricing
- $0.25 / $1.5 per 1M tokens, in / out
- Context window
- 1M context
Ready on endueUse in agent - Google
Gemini 2.5 Flash Lite
Gemini Flash-LiteAmong the cheapest multimodal models here, 1M context.
MultimodalText + Images + Files + Audio + Video → Text- Low cost
- Fast replies
- Audio & video
- Pricing
- $0.1 / $0.4 per 1M tokens, in / out
- Context window
- 1M context
Ready on endueUse in agent - Meta
Muse Spark 1.3
Muse SparkKeeps track of long tasks. Reads text, images, audio, video and PDF.
MultimodalText + Images + Video + Files + Audio → Text- Agentic work
- Audio & video
- Long documents
- Pricing
- $1.25 / $4.25 per 1M tokens, in / out
- Context window
- 1M context
Ready on endueUse in agent - Meta
Muse Spark 1.1
Muse SparkMultimodal reasoning for agent tasks, 1M context.
MultimodalText + Images + Video + Files + Audio → Text- Agentic work
- Audio & video
- Pricing
- $1.25 / $4.25 per 1M tokens, in / out
- Context window
- 1M context
Ready on endueUse in agent
No model matches these filters.