Towards Real-World Document Parsing via Realistic Scene Synthesis and Document-Aware Training
DGX agentarXiv:2603.23885v3 Announce Type: replace Abstract: Document parsing has recently advanced with multimodal large language models (MLLMs) that directly map document images to structured outputs. Tradit