Integrating Facial Generation into Full-Duplex Spoken Dialogue Systems
DGX agentarXiv:2606.21970v1 Announce Type: cross Abstract: Full-duplex spoken dialogue models, such as Moshi, enable natural, low-latency voice conversations. However, they remain limited to the audio modality