A Reliability Assessment of LALM Audio Judges for Full-Duplex Voice Agents
arXiv:2607.07985v1 Announce Type: cross Abstract: We report the empirical reliability of Gemini models as audio judges that score full-duplex agent conversations directly from the raw stereo waveform,