Hevolve AI: Self-Evolving Multimodal AI Agents

Turn your domain expertise into AI agents that keep learning. Hevolve AI lets experts build multimodal AI systems by talking to them and correcting them in real time, with no code to write.

Key Features

Quick Links

© 2024 Hevolve AI Pvt Ltd. All rights reserved.

← All research
Artificial Intelligence
arXiv (cs.CL) · July 17, 2026

Rate-Utility Frontiers for Language Encodings: Comparing Tokens, Bytes, and Pixels Under Controlled Linguistic Content

Ingo Ziegler, Martin Krebs, Desmond Elliott

A language model has to turn text into something it can process, and there are three broad answers: subword tokens, raw bytes, or images of rendered text. Comparisons between them are usually unfair in a way that is hard to see, because each encoding exposes a different amount of actual linguistic content depending on the language.

That unfairness is not evenly distributed. Tokenisers trained mostly on English chop other scripts into far more pieces, so a model reading Hindi or Amharic sees a smaller slice of the sentence within the same budget. Comparing encodings without controlling for this measures the tokeniser's training data as much as the encoding itself.

The study controls both content and downstream capacity, using verified parallel sentences across thirteen languages and five scripts, so each encoding is judged on what it preserves rather than on how much it happened to be handed. Framing it as a rate-utility frontier is apt: the question was never which encoding is best, but what each costs for what it keeps.

From the arXiv (cs.CL) abstract

Language models encode text as subword tokens, raw bytes, or rendered pixels, but these encodings are usually compared under modeling constraints that expose different amounts of linguistic content to models across different languages. We instead ask what each encoding preserves when both the content and the downstream capacity are controlled. Using verified parallel sentences across thirteen languages and five scripts, we compare tokens, bytes, and pixels through a shared…


More Artificial Intelligence papers