The frontier open model race just moved to sparse math. deepseek v4.1 flash, reported september 14,…
the frontier open model race just moved to sparse math. deepseek v4.1 flash, reported september 14, runs a 552-billion-parameter backbone with only 8 billion active per token, under an mit license, and beats deepseek's own v4 pro on agentic coding benchmarks.
sparsity became the moat.
Context
DeepSeek's V4.1 Flash news post of 10 September 2026 describes a 552B-parameter MoE with a Causal Encoder-Decoder architecture of 8B active parameters for input and 16B for output, says benchmark results are ahead of flagship models including DeepSeek-V4-Pro and that tests by multiple parties put it ahead of V4-Pro on performance, cost, speed and total runtime, and says from 04:00 UTC on 14 September 2026 deepseek-v4-pro requests route to V4.1-Flash at V4.1-Flash rates until V4.1-Pro launches. The Hugging Face model card snippet shows license mit, and an arXiv paper (2609.19969) covers its KV cache compression.
552 billion is first-party. MIT rests on the model card snippet and the license file was not read. Only 8 billion active per token is incomplete: the post gives 8B for input and 16B for output, and an earlier note quoted the 16B half. Reported September 14 is not the release report, since the post is dated 10 September and 14 September is when V4-Pro traffic began routing to V4.1-Flash. Beating V4-Pro on agentic coding is DeepSeek's own vendor claim; the benchmark charts are images that were not read and the multiple parties were not identified. The race moved to sparse math is the author's framing.
Related work
- Earlier note on the V4.1 Flash release ↗Same release, weights, KV cache and Terminal-Bench figures.
Watch next
- The benchmark tables, the identity of the third-party tests and V4.1-Pro.
Sources
Provenance
The note above is reproduced unedited from the original post, first published on Threads on 21 September 2026 at 10:47 IST. Sources are the papers and datasets the note draws on.
View the original post ↗Embed this note
More notes
The air is now being asked to keep its own ledger
the air is now being asked to keep its own ledger: ecmwf’s aifs compo becomes the first ai model to forecast atmospheric composition globally every three hours, cleanair simulates 365 days of pm2.5 over china in ten seconds, and a unified framework maps six pollutants at one kilometer across the whole country. the air now files its own composition report.
read the note →The current is now being asked to draw its own map
the current is now being asked to draw its own map: china’s langya 2.0 predicts six ocean phenomena including internal waves and mesoscale eddies, a deep net called wenhai resolves eddies globally with air sea flux formulas built in, and scripps infers surface currents from the way temperature patterns deform in satellite images. the ocean now files its own circulation report.
read the note →The soil is now being asked to report its own carbon
the soil is now being asked to report its own carbon: a nix color sensor paired with generative data augmentation predicts soil organic carbon without a lab, random forest drives 74 percent of soil health mapping studies, and sentinel 2 tracks five year carbon change across france and italy from 922 samples. the dirt now files its own carbon account.
read the note →