Less Is More: Reducing Token Counts Without Compromising Performance
arXiv:2506.15138v2 Announce Type: replace-cross Abstract: Tokenization directly affects the inference efficiency of large language models, since fragmented tokenization increases sequence length and g