StepFun
Step 3.7 Flash
stepfun/step-3.7-flash
Efficient MoE with native image and video understanding.
MultimodalText, Images, Video → TextReady on endue
- Audio & video
- Low cost
- Images
Overview
Good for
Written by endue
- Audio & video
- Low cost
- Images
Not ideal for
From the model’s specs
Nothing in the specs rules out common agent work.
Capabilities
| Tool use | Supported |
|---|---|
| Structured output | Supported |
| JSON mode | Supported |
| Reasoning | Supported |
| Vision | Supported |
| Audio input | Not supported |
| Video input | Supported |
| File input | Not supported |
| Long context | Supported |
Specifications
- Model ID
stepfun/step-3.7-flash- Family
- Step
- Context window
- 262K tokens
- Input
- Text, Images, Video
- Output
- Text
- Released
- 2026-05-28