Data & Decision ScienceJul 30, 2026

A 4B multimodal model claims 75% fewer visual tokens by encoding only what moves

Mage-VL, posted 27 July 2026, is a codec-native streaming multimodal foundation model that works at a 16x16 patch level and selectively encodes dynamic regions to keep spatiotemporal context. The authors report over 75% lower visual-token consumption and up to a 3.5x wall-clock inference speedup, with Mage-VL-4B matching Qwen3-VL-4B on static tasks and surpassing a 15B Phi-4-reasoning-vision baseline; pre-training used roughly 560M unlabelled images and 100M unlabelled video frames. Every figure here is the paper’s own — no independent benchmark run is attached.

What it means Token count is the bill for anything that watches a video feed all day; a 75% claim is worth a reproduction before it is worth a migration.

Where it came from arXiv

Back to the Stream