Robotics & Physical AIAug 27, 2026

Z.ai shipped GLM-5.3-Flash, a 321B open-weight model built for cheap coding work

Z.ai published GLM-5.3-Flash on 26 August: a natively multimodal 321B-parameter model using a hybrid sparse and linear attention architecture aimed at efficient coding. It went up on OpenRouter the same day at $0.075 per million input and $0.25 per million output tokens, currently discounted by half, and reached the top of Hugging Face's trending list alongside Qwen's release.

What it means Two large open-weight coding models landed on the same day at flash-tier prices; the cheap end of the coding-agent stack is now an open-weights question, not a frontier-lab one.

Where it came from Z.ai

A later run found this story reported independently, here:

Back to the Stream