endueendue

StepFun

Step 3.7 Flash

stepfun/step-3.7-flash

Efficient MoE with native image and video understanding.

MultimodalText, Images, Video → TextReady on endue
  • Audio & video
  • Low cost
  • Images

Overview

Good for

Written by endue

  • Audio & video
  • Low cost
  • Images

Not ideal for

From the model’s specs

Nothing in the specs rules out common agent work.

Capabilities

Tool useSupported
Structured outputSupported
JSON modeSupported
ReasoningSupported
VisionSupported
Audio inputNot supported
Video inputSupported
File inputNot supported
Long contextSupported

Specifications

Model ID
stepfun/step-3.7-flash
Family
Step
Context window
262K tokens
Input
Text, Images, Video
Output
Text
Released
2026-05-28