Agents
Don't Just Listen, Try Planning: Graph-based Retrieval-Generation Agent for Long-form Audio Meeting Understanding
arXiv:2608.24048v1 Announce Type: cross Abstract: While long-form audio meeting understanding (LAMU) is garnering growing attention, task-specific question answering (QA) datasets remain scarce. Exist
arXiv:2608.24048v1 Announce Type: cross Abstract: While long-form audio meeting understanding (LAMU) is garnering growing attention, task-specific question answering (QA) datasets remain scarce. Existing speech QA paradigms and state-of-the-art Speech LLMs suffer from acoustic information loss and poor long-term context memory. To address these issues, we construct the LongAudioQA dataset and propose the GRGA model, which models heterogeneous audio features into a multi-dimensional graph and leverages agent planning for retrieval and answer generation.
Related
- MemGraphRAG: Memory-based Multi-Agent System for Graph Retrieval-Augmented Generation
- LegalGraphRAG: Multi-Agent Graph Retrieval-Augmented Generation for Reliable Legal Reasoning
- A2RAG: Adaptive Agentic Graph Retrieval for Cost-Aware and Reliable Reasoning
- TechGraphRAG: An Agentic Graph-Augmented RAG Framework for Technical Literature Reasoning
Source: arXiv cs.AI | 2026-08-26