Building a Multimodal Dataset of Academic Paper for Keyword Extraction
DGX agentarXiv:2606.31069v1 Announce Type: new Abstract: Up to this point, keyword extraction task typically relies solely on textual data. Neglecting visual details and audio features from image and audio mod