You Can Learn Tokenization End-to-End with Reinforcement Learning
DGX agentarXiv:2602.13940v2 Announce Type: replace-cross Abstract: Tokenization is a hardcoded compression step which remains in the training pipeline of Large Language Models (LLMs), despite a general trend t