Nature study shows existing language models can be cheaply converted to read raw bytes
Researchers led by the Allen Institute for AI describe 'byteification', a two-stage method that retrofits standard language models to work on raw bytes instead of word-piece tokens. The converted models came close to their originals and beat earlier byte-level models of similar size.

A paper published in Nature on October 7 by researchers from the Allen Institute for AI, the University of Cambridge, the University of Washington, Imperial College London and LMU Munich introduces a method the authors call byteification. Instead of training a byte-level model from scratch, the method converts an existing subword-token model into one that operates directly on bytes.
Most language models split text into words or word fragments, which hides individual characters and can hurt performance on code, biological sequences and character-level tasks. Byte-level models avoid this, but the authors note they have lagged behind token-based models in practice.
The team retrofitted Olmo 3 7B, Qwen3 8B and Llama 3 8B into byte-level versions using less than 1% of a typical pretraining budget. The authors report that the resulting models outperformed earlier publicly available byte-level models of comparable size on average, came close to their source models, and were much stronger at character understanding.
The paper also shows that existing post-trained versions of the source model can be merged into the byte-level model to add capabilities such as instruction following without extra training. Training data and code have been released publicly.