The Price of Token Boundaries: Compression Certificates and Prediction
arXiv cs.AIen
arXiv cs.AI
AI Global WirearXiv:2609.35869v1 Announce Type: new Abstract: Pre-tokenisation restricts which text fragments can become prediction units, but its compression cost is obscured when tokenisers are compared only under the same boundaries. We measure this cost by bounding the minimum token count from both sides, with and without a regular-expression boundary rule. Nonnegative prices on token occurrences yield a lower bound through shortest paths and vocabulary-budget selection; maximising over all prices recovers the linear programming relaxation, and an independent integer checker certifies the reported values. On English Wikipedia, boundaries increase the optimal token count by 28.3--36.8\%. Byte pair enco
This is a short summary published by AI Global Wire. The full article is owned and hosted by arXiv cs.AI — open it there to read it in full.
Read the full story at arXiv cs.AI- Forskning
Related AI news
- Chinese firms trail global peers on profits, but AI power boom offers bright spot: NatixisSCMP Tech · September 30, 2026
- More Features Are Not More Evidence: Limits of Training-Free Human Activity Recognition with JevarXiv cs.AI · September 30, 2026
- Towards Mitigating Deceptive Safety Alignment in Large Reasoning ModelsarXiv cs.AI · September 30, 2026
- GeoOutageBench: Benchmarking Ambiguity-aware, Ontology-grounded Geospatiotemporal KGQA for Multimodal Power Outage and Resilience AnalysisarXiv cs.AI · September 30, 2026
- The Layer Mystery of VLA: An Information-Theoretical Analysis of VLA Latent InterfacearXiv cs.AI · September 30, 2026
- An Empirical Study and Assessment of EU AI Act Compliance CheckersarXiv cs.AI · September 30, 2026