← Founder Notes
Archive

Model speed now comes from the serving stack, not the model. vllm's september 13 update for kimi k3…

Yethikrishna ROriginal on Threads

model speed now comes from the serving stack, not the model. vllm's september 13 update for kimi k3 cut latency 56-60%, lifted throughput 2.2-2.8x and slashed time-to-first-token 72-85%.

the inference engine just became the fastest way to get faster.

Context

vLLM's blog of 13 September 2026 on Kimi K3 performance reports 56 to 60 percent lower latency, 2.2 to 2.8x throughput and 72 to 85 percent lower time to first token across concurrency 1, 4 and 16, on an 8K input and 1K output workload, TP8, eight-token DSpark speculation and a B300 node. The baseline is vLLM v0.27.1 against main commit 82a85dc1, for the same model.

How it compares

The figures match the first-party post and are vendor-reported, not independent. This is a software-version comparison for one model on one hardware tier, and the update is a main-branch commit and not a tagged release as far as the text says. Other workloads saw smaller per-PR gains of 5 to 41 percent. The inference engine just became the fastest way to get faster is the author's take.

Watch next

  • The release tag containing these changes and independent reruns.

Sources

  1. Kimi K3 performance optimization (vLLM blog, 13 Sep 2026)vllm.ai

Provenance

The note above is reproduced unedited from the original post, first published on Threads on 20 September 2026 at 21:51 IST. Sources are the papers and datasets the note draws on.

View the original post
Embed this note
<iframe src="https://founder.myndlabs.tech/notes/embed/model-speed-now-comes-from-the-serving-stack-DdhAHLxja7D" width="480" height="420" style="border:0;max-width:100%" loading="lazy" title="Model speed now comes from the serving stack, not the model. vllm's september 13 update for kimi k3…"></iframe>

More notes