← Pooneh Mousavi

Conversational AI Reading Group

Thursdays · 11:00–12:00 ET Online, on Zoom Open to everyone

🎙️ Upcoming talks

Oct15
● Next talk

Sound Design AI

🎤 Oriol NietoAdobe Research

Oct29
Thursday

Scaling Speech LLMs Responsibly: Languages, Privacy and Decentralization

🎤 Alessio BruttiFondazione Bruno Kessler (FBK)

📼 Past talks

55 talks
Fall 20261 talks⌄
Oct 8

Towards Fully Open General Audio Intelligence

🎤 Sreyan Ghosh · Google DeepMind

Spring 202615 talks⌄
May 14

The Grand Design Challenge of Music GenAI

🎤 Nicholas J. Bryan · Adobe Research

Apr 30

Large Audio-Language Models: Algorithms and Applications

🎤 Wenwu Wang · University of Surrey

Apr 23

Improving spoken language identification for non-native speech

🎤 Tanel Alumäe · TalTech

Apr 16

Toward Trustworthy Voice Agents: Speech-Native Evaluation, Safety, and Security

🎤 Amir Ivry · Technion, incoming CMU

Apr 9

Faithful and Grounded Audio Language Models

🎤 Cem Subakan · Laval University & Mila

Apr 2

Voice-Based Psychiatric Assessment in the Real World and the Case for Production-Aligned Speech Frontends

🎤 George Fairs & Stefano Goria · thymia

Mar 26

Prompt-Conditioned Unified Audio Modeling: From Source Separation to Source-Aware Codecs

🎤 Jonathan Le Roux · Mitsubishi Electric Research Laboratories (MERL)

Mar 19

SAM Audio: Segment Anything in Audio

🎤 Andros Tjandra & Bowen Shi · FAIR (Meta AI)

Mar 12

Advancing the Linguistic Capabilities of Speech Language Models

🎤 Ricard Marxer · Université de Toulon

Mar 5

Auden: Where is the “GPT moment” for audio?

🎤 Yiwen Shao · Tencent AI Lab

Feb 26

Conversational Speech Processing: Challenges & Opportunities

🎤 Samuele Cornell · Carnegie Mellon University

Feb 19

Amplitude Modulation Spectral Analysis: From Conventional Audio Feature Engineering to Speech Foundation Models

🎤 Tiago H. Falk · INRS

Feb 12

Towards Accountable Conversational Agents for Task Completion

🎤 Dilek Hakkani-Tür · Univ. of Illinois at Urbana-Champaign

Feb 5

Model-based audio deep learning with application to source separation and dereverberation

🎤 Gaël Richard · Télécom Paris

Jan 29

Recomposer: Event-roll-guided Audio Editing

🎤 Daniel P. W. Ellis · Google Deepmind

Fall 202513 talks⌄
Dec 18

Data as Leverage: Improving Foundation Models Beyond Scaling

🎤 Daniel D'souza · Cohere Labs

Dec 11

Low-latency Conversational Agent

🎤 Tatiana Likhomanenko · Apple

Nov 27

Multimodal analysis of Parkinson’s disease symptoms

🎤 Juan Rafael Orozco-Arroyave · Universidad de Antioquia

Nov 20

Neural Target Speech and Sound Extraction

🎤 Marc Delcroix · NTT Communication Science Laboratories

Nov 13

Efficient and resource-constrained dynamic neural networks

🎤 Simone Scardapane · Sapienza University of Rome

Nov 6

Machine learning paradigms for music and audio understanding

🎤 Emmanouil Benetos · Queen Mary University of London

Oct 30

Gemini Voice Agent: A Natively Multimodal Dialog Model with Advanced Reasoning and Tool Use

🎤 Michael Han · Google DeepMind

Oct 23

Supervised contrastive learning from weakly-labeled audio segments for musical version matching

🎤 Joan Serrà · Sony AI

Oct 16

The development of spoken LM

🎤 Jinyu Li · Microsoft

Oct 9

AI-based Spatial Audio

🎤 Fabio Antonacci · Politecnico di Milano

Oct 2

A safety case for the Scientist AI

🎤 Yoshua Bengio · LawZero

Sep 25

Advances in Speaker Recognition: Pruning, Deepfake Detection, and Learning without Temporal Labels

🎤 Themos Stafylakis · Athens University of Economics and Business

Sep 18

Discrete Audio Tokens: More Than a Survey!

🎤 Pooneh Mousavi · Mila - Concordia

Summer 20253 talks⌄
Jun 26

Audio Processing in the Age of Large Language Models

🎤 Dinesh Manocha · University of Maryland

Jun 19

AI for Creators: Pushing Creative Abilities to the Next Level

🎤 Yuki Mitsufuji · SonyAI

Jun 12

Meta Audiobox Aesthetics: Unified Automatic Quality Assessment for Speech, Music, and Sound

🎤 Andros Tjandra · FAIR (Meta AI)

Spring 202512 talks⌄
Jun 5

Voice conversion and the geometry of self-supervised speech representations

🎤 Herman Kamper · Stellenbosch University

May 29

On The Landscape of Spoken Language Models

🎤 Yossi Adi · Hebrew University and Meta

May 22

Speaker diarization, a love loss story

🎤 Hervé Bredin · pyannoteAI

May 15

Automatic Quality Assessment for Speech and Beyond

🎤 Wen-Chin Huang · Nagoya University

May 8

The Voicebox Model and Its Applications

🎤 Leda Sari · Otter.ai

May 1

CR-CTC: Consistency regularization on CTC for improved speech recognition

🎤 Daniel Povey · Xiaomi Corp

Apr 24

GenAI for Sound Design

🎤 Oriol Nieto · Adobe Research

Apr 17

Unsupervised on-device adaptation of a speech recogniser and the Pitfalls of "SpeechLLM" evaluation

🎤 Titouan Parcollet · Samsung AI Center Cambridge

Apr 10

Toward Understanding Sign Language in the Real World

🎤 Karen Livescu · TTIC

Apr 3

Improving Multilingual Speech Recognition and Language Identification

🎤 Min Ma · Google DeepMind

Mar 27

Learning Source Disentanglement in Neural Audio Codec

🎤 Xiaoyu Bie · Télécom Paris

Mar 20

Making transformers work for audio coding

🎤 Julian Parker · Stability AI

Winter 20258 talks⌄
Mar 13

Moshi: a speech-text foundation model for real-time dialogue

🎤 Alexandre Defossez · Kyutai

Feb 27

Open Whisper-Style Speech Models: Transparency, Scalability, and Advancing Explainability

🎤 Shinji Watanabe · Carnegie Mellon University

Feb 20

Singing Voice Synthesis: Data curation, Modeling, and Evaluation

🎤 Jiatong Shi · Carnegie Mellon University

Feb 13

Scalable and Efficient Speech Enhancement

🎤 Minje Kim · University of Illinois at Urbana-Champaign

Feb 6

Teaching Foundation Models New Skills: Insights and Experiences

🎤 Hung-yi Lee · National Taiwan University

Jan 23

Foundational Speech Models and Their Efficient Training with NVIDIA NeMo

🎤 Piotr Żelasko · Nvidia

Jan 16

Improving Universal Access to Modern Speech Technology

🎤 Martijn Bartelds · Stanford University

Jan 9

Neural Audio Codecs in the Era of Speech LMs

🎤 Haibin Wu · Microsoft

Fall 20243 talks⌄
Dec 19

Discrete Audio Tokens for Multimodal LLMs

🎤 Mirco Ravanelli · Concordia University - Mila

Dec 5

Posthoc Explanations for Audio Models

🎤 Cem Subakan · Université Laval - Mila

Nov 21

PARAMETER AVERAGING IS ALL YOU NEED TO PREVENT FORGETTING

🎤 Peter Plantinga · McGill University