Sound Design AI
🎤 Oriol NietoAdobe Research
Papers (10)
- TAC: Timestamped Audio Captioning NeurIPS 2026
- Generative Audio Extension and Morphing ICASSP 2026
- Mix2Morph: Learning Sound Morphing From Noisy Mixes ICASSP 2026
- PromptSep: Generative Audio Separation Via Multimodal Prompting ICASSP 2026
- AudioCards: Structured Metadata Improves Audio Language Models For Sound Design ICASSP 2026
- SoundStager: Interactive Design of Story-Driven GenAI Soundscapes for Video CHI 2026
- SILA: Signal-to-Language Augmentation for Enhanced Control in Text-to-Audio Generation WASPAA 2025
- FLAM: Frame-Wise Language-Audio Modeling ICML 2025
- Visual Description Grounding Reduces Hallucinations and Boosts Reasoning in LVLMs ICLR 2025
- Sketch2Sound: Controllable Audio Generation via Time-Varying Signals and Sonic Imitations ICASSP 2025
Abstract
Sound design has traditionally been expensive, slow, and daunting, with creators sifting through large sound-effect libraries to find and time the right audio. This talk presents the recent work of Adobe Research's Sound Design AI (SODA) team, which is building generative tools that make professional sound design faster, more affordable, and more intuitive. We trace a progression of increasingly expressive controls: from text-to-SFX, to voice-to-SFX driven by vocal imitations (Sketch2Sound), to video-guided Foley generation with multimodal controls (MultiFoley) that syncs sound to on-screen action. We then look toward fully agentic, story-driven soundscape design (Sound Stager), alongside supporting technologies for understanding, editing, and reshaping audio: timestamped audio captioning (TAC), prompt-based sound separation (PromptSep), generative sound morphing (Mix2Morph), generative audio extension (GenExtend), and structured metadata for sound libraries (AudioCards). Together, these advances point toward a future where anyone can create rich, well-timed, and highly controllable soundscapes with minimal cost and effort.
Speaker bio
Oriol is a Senior Research Engineer 2 at Adobe Research, where he focuses on human-centered AI for audio creativity, encompassing everything from music to audiobooks, video editing, and, more recently, sound design. He holds a PhD in Music Technology from MARL, NYU, a Master's in Music, Science, and Technology from Stanford University, and a Master's in Information Technologies from Pompeu Fabra University. Highly involved with the Music Information Retrieval community, he served as General Chair for ISMIR 2024 in San Francisco. Oriol has authored relevant open-source MIR packages such as librosa, mir-eval, and MSAF; contributed to PyTorch; and plays guitar, violin, cajón, and sings (and screams) in his spare time.
