Robotics & Physical AISep 5, 2026
llama.cpp reaches v0.4.0 with new model support and per-slot context limits
llama.cpp published v0.4.0, adding initial Qwen3.8-Flash-Next and Nemotron-3-Puzzle support, on-demand tensor reading, per-slot server context limits, video input options and a ggml bump to 0.23.0.
What it means Per-slot context limits and on-demand tensor reading are the practical knobs for serving several users off one local box without running out of memory.
Where it came from ggml-org / llama.cpp