Umans GLM 5.3 Flash (lab) Experimental
…
throughput · p50 · no requests
…
TTFT · p50 · no requests
100.00%
uptime · 24h
GLM-5.3-Flash as a Labs experiment, open for a short test window: temporary, not a permanent id. Served from the native weights (Z.ai's GLM-5.3-Flash): a 320B mixture-of-experts model with 18B active parameters per token, the first natively multimodal release in the GLM-5 series, built for fast coding and agentic tasks on a 1M context window. It always thinks at max effort by default; dial reasoning to low or high (thinking cannot be turned off). Access is seat-gated through the Labs page while an experiment is live. It is offered at limited capacity and low availability, so expect it to be flaky and to go down under load: crash it, give it a moment, and try again. For production work we recommend umans-coder or umans-deepseek-v4-pro-0813.
Trends
Speed over the last 90 days
peak 94.8 tok/s · Aug 27now 88.6 tok/s
90 days agotoday
now 2.98s
90 days agotoday
Changelog
Events for Umans GLM 5.3 Flash (lab)
Aug 262026
New Labs experiment: Umans GLM 5.3 Flash Testing
A new lab opened on umans-glm-5.3-flash-lab: GLM-5.3-Flash served from its native weights, a 320B mixture-of-experts model (18B active per token) and the first natively multimodal release in the GLM-5 series, with native image and video understanding and a 1M context window. It always thinks at max effort by default (dial to low or high; thinking cannot be turned off). Free and seat-gated while the experiment runs, served at limited capacity and low availability: it is a lab, so expect it to be flaky and to go down under load. Crash it, give it a moment, and try again. The window closes August 28, 2026.