Robotics & Physical AISep 5, 2026

llama.cpp reaches v0.4.0 with new model support and per-slot context limits

llama.cpp published v0.4.0, adding initial Qwen3.8-Flash-Next and Nemotron-3-Puzzle support, on-demand tensor reading, per-slot server context limits, video input options and a ggml bump to 0.23.0.

What it means Per-slot context limits and on-demand tensor reading are the practical knobs for serving several users off one local box without running out of memory.

Where it came from ggml-org / llama.cpp

Back to the Stream