Research
StereoFoley: Object-Aware Stereo Audio Generation from Video
We present StereoFoley, a video-to-audio generation framework that produces semantically aligned, temporally synchronized, and spatially accurate stereo sound at 48 kHz. While recent generative video-
We present StereoFoley, a video-to-audio generation framework that produces semantically aligned, temporally synchronized, and spatially accurate stereo sound at 48 kHz. While recent generative video-to-audio models achieve strong semantic and temporal fidelity, they largely remain limited to mono or fail to deliver object-aware stereo imaging, constrained by the lack of professionally mixed, spatially accurate video-to-audio datasets. First, we develop and train a base model that generates stereo audio from video, achieving state-of-the-art in both semantic accuracy and synchronization. Next…
Related
- UniSonate: A Unified Model for Speech, Music, and Sound Effect Generation with Text Instructions
- MMAudioSep: Taming Video-to-Audio Generative Model Towards Video/Text-Queried Sound Separation
- FoleyDirector: Fine-Grained Temporal Steering for Video-to-Audio Generation via Structured Scripts
- Don't Let the Video Speak: Audio-Contrastive Preference Optimization for Audio-Visual Language Models
- Exploring Audio Hallucination in Egocentric Video Understanding
Source: Apple ML Research | 2026-04-28