← Founder Notes
Archive

The inference server that everyone runs just made model restarts nearly free. vllm v0.30.0, out…

Yethikrishna ROriginal on Threads

the inference server that everyone runs just made model restarts nearly free. vllm v0.30.0, out september 22, keeps quantized weights resident in gpu memory via a per-gpu daemon, so a restarting engine maps them over cuda ipc instead of reloading from disk, and ships hybrid-attention paths for kimi k3, deepseek-v4.1-flash and qwen3.8-flash-next.

serving infra now treats weights like a cache instead of a cold start.

Provenance

The note above is reproduced unedited from the original post, first published on Threads on 24 September 2026 at 01:47 IST.

View the original post
Embed this note
<iframe src="https://founder.myndlabs.tech/notes/embed/the-inference-server-that-everyone-runs-just-made-DdpJh2ZiPpL" width="480" height="420" style="border:0;max-width:100%" loading="lazy" title="The inference server that everyone runs just made model restarts nearly free. vllm v0.30.0, out…"></iframe>

More notes