Tokenisation via Convex Relaxations
- Published
- Source
- arXiv
- Paper number
- 219
- Field
- LLMs / NLP
- arXiv ID
- 2605.22821
Key points
- For the integer-only approach, linear programming selects only tokens it is almost certain about, meaning tokens with values of 0.999 or higher. This often yields a vocabulary smaller than the allowed budget.
- Proven optimality limit: now we know how much room remains to squeeze tokenization for compression. Since BPE is already within 1% of the limit at scale, researchers can focus more on other properties such as language morphology or robustness rather than compression.
- Non-greedy alternative: ConvexTok builds a vocabulary by considering the entire dataset at once, which makes better use of the vocabulary budget.
Paper links
External research summaries. These are not HDATF publications or measured product results.