Robotics & Physical AIAug 27, 2026
Z.ai shipped GLM-5.3-Flash, a 321B open-weight model built for cheap coding work
Z.ai published GLM-5.3-Flash on 26 August: a natively multimodal 321B-parameter model using a hybrid sparse and linear attention architecture aimed at efficient coding. It went up on OpenRouter the same day at $0.075 per million input and $0.25 per million output tokens, currently discounted by half, and reached the top of Hugging Face's trending list alongside Qwen's release.
What it means Two large open-weight coding models landed on the same day at flash-tier prices; the cheap end of the coding-agent stack is now an open-weights question, not a frontier-lab one.
Where it came from Z.ai
A later run found this story reported independently, here:
- deeplearning.aiseen Aug 31, 2026