Z.ai
GLM-5V Turbo
z-ai/glm-5v-turbo
GLM that reads images and video, for vision-based coding.
MultimodalImages, Text, Video → TextReady on endue
- Images
- Coding
- Audio & video
Overview
Good for
Written by endue
- Images
- Coding
- Audio & video
Not ideal for
From the model’s specs
Nothing in the specs rules out common agent work.
Capabilities
| Tool use | Supported |
|---|---|
| Structured output | Not supported |
| JSON mode | Supported |
| Reasoning | Supported |
| Vision | Supported |
| Audio input | Not supported |
| Video input | Supported |
| File input | Not supported |
| Long context | Supported |
Specifications
- Model ID
z-ai/glm-5v-turbo- Family
- GLM-5V
- Context window
- 203K tokens
- Input
- Images, Text, Video
- Output
- Text