Robotics & Physical AIAug 28, 2026
Microsoft details Maia 200 at Hot Chips and says it is already serving in an Iowa Azure region
Microsoft used Hot Chips 2026 on 25 August to give the technical detail behind Maia 200, the inference accelerator it first announced in January 2026: TSMC 3nm, 140 billion transistors, six HBM stacks for 7TB/s of bandwidth, and over 10 petaFLOPS at FP4 within a 750W part. Microsoft says it is the most efficient inference system it has deployed, at 30% better performance per dollar than the current fleet generation, and that it is running in the US Central region near Des Moines with US West 3 near Phoenix to follow. The efficiency figures are Microsoft’s own.
What it means Inference cost per token is set by which silicon a cloud actually runs; a hyperscaler moving its own part into a live region is where the price pressure on rented GPUs starts.
Where it came from ServeTheHome, from Hot Chips 2026