The open source serving engine just made the same model 2.8 times faster. vllm's kimi k3…
the open source serving engine just made the same model 2.8 times faster. vllm's kimi k3 optimizations, out september 13, combine scheduling, prefix caching and pd disaggregation.
nobody touched the weights.
Context
The vLLM blog of September 13, 2026 reports merged work for Kimi K3, comparing v0.27.1 against main at commit 82a85dc1. It reports 56 to 60 percent lower latency, 2.2 to 2.8x throughput and 72 to 85 percent lower time to first token across concurrency 1, 4 and 16, on an 8K input, 1K output workload with TP8, eight-token DSpark speculation and a B300 node on CUDA 13.3. The listed changes include adaptive speculative-token budgets, internal KDA prefix checkpoints, zero-copy mixed batches and deferred MXFP4 finalization. The post describes prefill-decode disaggregation in a separate section.
This is a blog about merged work and not a versioned release, and the project reports its own numbers on one hardware and configuration, so 2.8x is the upper end at one concurrency and not universal. Prefix caching was disabled in both benchmark runs because v0.27.1 had a known Kimi K3 prefix-caching issue fixed later, so it did not contribute to the headline, and no measured effect of prefill-decode disaggregation on the 2.8x was stated in the text read. The note's combination of scheduling, prefix caching and pd disaggregation therefore overstates what the headline benchmark measured. Nobody touched the weights is the author's take.
Watch next
- A tagged vLLM release containing these changes and independent replications.
Sources
Provenance
The note above is reproduced unedited from the original post, first published on Threads on 22 September 2026 at 17:46 IST. Sources are the papers and datasets the note draws on.
View the original post ↗Embed this note
More notes
The air is now being asked to keep its own ledger
the air is now being asked to keep its own ledger: ecmwf’s aifs compo becomes the first ai model to forecast atmospheric composition globally every three hours, cleanair simulates 365 days of pm2.5 over china in ten seconds, and a unified framework maps six pollutants at one kilometer across the whole country. the air now files its own composition report.
read the note →The current is now being asked to draw its own map
the current is now being asked to draw its own map: china’s langya 2.0 predicts six ocean phenomena including internal waves and mesoscale eddies, a deep net called wenhai resolves eddies globally with air sea flux formulas built in, and scripps infers surface currents from the way temperature patterns deform in satellite images. the ocean now files its own circulation report.
read the note →The soil is now being asked to report its own carbon
the soil is now being asked to report its own carbon: a nix color sensor paired with generative data augmentation predicts soil organic carbon without a lab, random forest drives 74 percent of soil health mapping studies, and sentinel 2 tracks five year carbon change across france and italy from 922 samples. the dirt now files its own carbon account.
read the note →