Tokenisation via Convex Relaxations

Published
Source
arXiv
Paper number
219
Field
LLMs / NLP
arXiv ID
2605.22821

Key points

  • For the integer-only approach, linear programming selects only tokens it is almost certain about, meaning tokens with values of 0.999 or higher. This often yields a vocabulary smaller than the allowed budget.
  • Proven optimality limit: now we know how much room remains to squeeze tokenization for compression. Since BPE is already within 1% of the limit at scale, researchers can focus more on other properties such as language morphology or robustness rather than compression.
  • Non-greedy alternative: ConvexTok builds a vocabulary by considering the entire dataset at once, which makes better use of the vocabulary budget.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)