← Founder Notes
Archive

Tokenization is quietly becoming optional. a meta fair study from september 11 found byte-level…

Yethikrishna ROriginal on Threads

tokenization is quietly becoming optional. a meta fair study from september 11 found byte-level distillation beats token-based teaching by 4 points on the predicted ceiling while needing a sixth of the training data.

the model that never learns a token may learn faster.

Context

The paper Breaking the Token Ceiling: Distilling Smaller, Stronger Byte Models (arXiv 2609.12303, submitted 11 September 2026; Kalyani Marathe, Artidoro Pagnoni, Tomasz Limisiewicz, Margaret Li, Mike Lewis, Luke Zettlemoyer, Srinivasan Iyer) introduces two ways to convert token logits to byte logits, Marginalize-It (approximate) and End-Of-Token (exact), overtrains models of roughly 1 billion parameters up to 1 trillion bytes of data and uses eight benchmarks. Its abstract says token-1B models win at low compute but plateau, byte models start worse and pass them with more compute, extrapolated scaling laws predict distilled End-Of-Token-1B outperforms distilled Token-1B by up to 4 percent asymptotically, they match distilled Token-1B with one sixth of the training data, and logit storage is about one fifth.

How it compares

Abstract only; the full paper was not read. The 4 is a prediction from extrapolated scaling laws, up to 4 percent, and not a measured gap, and points versus percent is not resolved. One sixth of the data is to match and not to beat distilled Token-1B. Meta FAIR was not confirmed from the text read, which gives author names only. Byte models are worse at low compute, so a model that never learns a token may learn faster overstates the abstract; that line is the author's take.

Watch next

  • Full-paper tables, the scaling-law fit and its uncertainty, and affiliation text.

Sources

  1. Breaking the Token Ceiling: Distilling Smaller, Stronger Byte Models (arXiv 2609.12303)arxiv.org

Provenance

The note above is reproduced unedited from the original post, first published on Threads on 21 September 2026 at 02:16 IST. Sources are the papers and datasets the note draws on.

View the original post
Embed this note
<iframe src="https://founder.myndlabs.tech/notes/embed/tokenization-is-quietly-becoming-optional-a-meta-fair-DdheeZ_EYoN" width="480" height="420" style="border:0;max-width:100%" loading="lazy" title="Tokenization is quietly becoming optional. a meta fair study from september 11 found byte-level…"></iframe>

More notes