Model Releases
Finetuning Strategies for Querying Sounds by Vocal Imitation
arXiv:2608.19174v1 Announce Type: cross Abstract: This technical report describes our winning submission to the AES AIMLA 2025 Challenge on querying sound effects by vocal imitation. We investigate tw
arXiv:2608.19174v1 Announce Type: cross Abstract: This technical report describes our winning submission to the AES AIMLA 2025 Challenge on querying sound effects by vocal imitation. We investigate two complementary fine-tuning strategies: contrastive learning with a frozen, pretrained CED encoder, and joint contrastive-triplet learning with semi-hard negatives using a MobileNetV3 encoder. This report has been updated for posterity to include details released after the challenge.
Related
- Unified Synthesis of Compositional Speech and Sound from Free-Form Text Prompts
- BAT: Learning to Reason about Spatial Sounds with Large Language Models
- ZONOS2 Technical Report
- Raon-Speech Technical Report
Source: arXiv cs.AI | 2026-08-20