👩🏻‍💻About me

I work on multimodal foundation models, audio representation learning, and interpretability.

My research focuses on bridging the gap between audio and large language models, from discrete audio representations to benchmarks and evaluation. I'm finishing my PhD at Mila and Concordia University with Mirco Ravanelli and Cem Subakan. I recently completed an internship at Epic Games, working on expressive generative models. I'm a core contributor to SpeechBrain and founded the Conversational AI Reading Group.

Conversational AI Reading GroupBack on October 8! Weekly talks on speech and conversational AI, Thursdays at 11:00 ET, open to all.
See talks →
📢News
📚Selected publications
  1. 2026

    Investigating Faithfulness in Large Audio Language Models

    Pooneh Mousavi, Lovenya Jain, Mirco Ravanelli, Cem Subakan

    Interspeech 2026 · Oral Paper Website
  2. 2026

    Enhancing Audio Reasoning via Semantic Summary Prediction

    Francesco Bonzi, Pooneh Mousavi, Cem Subakan, Mirco Ravanelli

    Interspeech 2026 Paper Website
  3. 2026

    DASB: Discrete Audio and Speech Benchmark

    Pooneh Mousavi, Jarod Duret, Darius Petermann, … Mirco Ravanelli

  4. 2026

    Listen First, Then Answer: Timestamp-Grounded Speech Reasoning

    Jihoon Jeong, Pooneh Mousavi, Mirco Ravanelli, Cem Subakan

    Preprint, under review Paper
  5. 2026

    ALAS: An Automatic Latent Alignment Score for Audio Language Models

    Pooneh Mousavi, Yingzhi Wang, Mirco Ravanelli, Cem Subakan

    MLSP 2026 · also ICML 2025 ML4Audio Workshop Paper Code
  6. 2025

    Discrete Audio Tokens: More Than a Survey!

    Pooneh Mousavi†, Gallil Maimon*, Adel Moumen*, … Mirco Ravanelli

    TMLR 2025 Paper Website
  7. 2025

    LiSTEN: Learning Soft Token Embeddings for Neural Audio LLMs

    Pooneh Mousavi*, Shubham Gupta*, Cem Subakan, Mirco Ravanelli

    Interspeech 2025 Paper
  8. 2025

    What Are They Doing? Joint Audio-Speech Co-Reasoning

    Yingzhi Wang, Pooneh Mousavi, Artem Ploujnikov, Mirco Ravanelli

    ICASSP 2025 Paper
  9. 2024

    How Should We Extract Discrete Audio Tokens from Self-Supervised Models?

    Pooneh Mousavi, Jarod Duret, Salah Zaiem, … Mirco Ravanelli

    Interspeech 2024 · Oral Paper Code

* equal contribution · † lead author

Google Scholar →
💼Experience
  • 🎮 Research Scientist Intern, Epic GamesGenerative & expressive TTS · Montréal · 2025–2026
  • Scientist in Residence, MilaAudio representations & audio LLMs · 2022–present
  • ML Scientist in Residence, Draft & GoalLLM fine-tuning with QLoRA · 2023
  • Research Scientist, UKP Lab, TU DarmstadtAI for digital mental health, affective analysis · 2021–2022
🎓Education
  • PhD Candidate, Computer ScienceMila / Concordia · 2022–present
  • MS, Computer ScienceUT Dallas · 2018–2021
  • BSc, Computer EngineeringK. N. Toosi University of Technology · 2009–2013
🏆Honors
  • Oral presentations, Interspeech2024, 2026
  • Mila travel scholarship, WiML2025
  • Concordia graduate travel grantInterspeech 2024, 2025, 2026
👩🏻‍🏫Teaching & service