Tokenization is quietly becoming optional. a meta fair study from september 11 found byte-level…
tokenization is quietly becoming optional. a meta fair study from september 11 found byte-level distillation beats token-based teaching by 4 points on the predicted ceiling while needing a sixth of the training data.
the model that never learns a token may learn faster.
Context
The paper Breaking the Token Ceiling: Distilling Smaller, Stronger Byte Models (arXiv 2609.12303, submitted 11 September 2026; Kalyani Marathe, Artidoro Pagnoni, Tomasz Limisiewicz, Margaret Li, Mike Lewis, Luke Zettlemoyer, Srinivasan Iyer) introduces two ways to convert token logits to byte logits, Marginalize-It (approximate) and End-Of-Token (exact), overtrains models of roughly 1 billion parameters up to 1 trillion bytes of data and uses eight benchmarks. Its abstract says token-1B models win at low compute but plateau, byte models start worse and pass them with more compute, extrapolated scaling laws predict distilled End-Of-Token-1B outperforms distilled Token-1B by up to 4 percent asymptotically, they match distilled Token-1B with one sixth of the training data, and logit storage is about one fifth.
Abstract only; the full paper was not read. The 4 is a prediction from extrapolated scaling laws, up to 4 percent, and not a measured gap, and points versus percent is not resolved. One sixth of the data is to match and not to beat distilled Token-1B. Meta FAIR was not confirmed from the text read, which gives author names only. Byte models are worse at low compute, so a model that never learns a token may learn faster overstates the abstract; that line is the author's take.
Watch next
- Full-paper tables, the scaling-law fit and its uncertainty, and affiliation text.
Sources
Provenance
The note above is reproduced unedited from the original post, first published on Threads on 21 September 2026 at 02:16 IST. Sources are the papers and datasets the note draws on.
View the original post ↗Embed this note
More notes
The air is now being asked to keep its own ledger
the air is now being asked to keep its own ledger: ecmwf’s aifs compo becomes the first ai model to forecast atmospheric composition globally every three hours, cleanair simulates 365 days of pm2.5 over china in ten seconds, and a unified framework maps six pollutants at one kilometer across the whole country. the air now files its own composition report.
read the note →The current is now being asked to draw its own map
the current is now being asked to draw its own map: china’s langya 2.0 predicts six ocean phenomena including internal waves and mesoscale eddies, a deep net called wenhai resolves eddies globally with air sea flux formulas built in, and scripps infers surface currents from the way temperature patterns deform in satellite images. the ocean now files its own circulation report.
read the note →The soil is now being asked to report its own carbon
the soil is now being asked to report its own carbon: a nix color sensor paired with generative data augmentation predicts soil organic carbon without a lab, random forest drives 74 percent of soil health mapping studies, and sentinel 2 tracks five year carbon change across france and italy from 922 samples. the dirt now files its own carbon account.
read the note →