P-MTP: Efficient Document Parsing via Multi-Token Prediction with Progressive Depth Scaling
arXiv:2606.24447v1 Announce Type: new Abstract: Vision-Language Models (VLMs) have revolutionized document parsing by enabling end-to-end mapping from images to structured text, imposing a significant