LISTEN-to-Reason: Listen with Experts, Retrieve over a Graph, Reason with LLMs

Listen-to-Reason (L2R) is an interpretable-by-design pipeline that passes audio to the LLM through an explicit, human-readable tree: small heads on frozen expert encoders map each chunk of a clip to semantically meaningful nodes on the tree (for speech, music and environmental sound), and a frozen text-only LLM answers from these nodes and an ASR transcript without hearing the clip. Every answer can therefore be traced to the nodes and transcript it read.

39 attributes · 363 values · 668 leavesin the audio tree, over speech, music and environmental sound
7.7M trained parameterssmall heads on frozen encoders; the reader LLM is never fine-tuned
Any text LLM as readerthe same cached description is read by 26 LLMs from 0.5B to 72B
One head per new domainbirds and marine mammals added from a few labelled clips per species

How it works

From waveform to answer, through an explicit audio tree

The same figure as in the paper, following one question through the system.

Audio clip3 s chunks,1.5 s hopRouterregion gate (per chunk)CLAPBEATsMuQ-MuLanWavLM-SVemotion2vecWhisper enc.one frozenencoder perattributefamily; smalltrained headsAudio treeclipspeechmusicsoundgenderanimalfemale voicedogBarkbird speciesnew domainText contextspeech gender: female voice;sound animal: dog (Bark)Transcript:[0:00.0] "Come here, boy!"nodes + transcriptif speech or singing is heardAny frozen LLMtext only; never hears the clipQuestion“Who calls the dog?”(b) a womananswer, traceableto the nodes aboveASR transcript, with timestampsfrozentrained (small heads only)paths reached by the routeradded for a new domain

The audio tree

Every node the router can reach

Click a node to open or close it. Regions hold attributes, each attribute has a fixed menu of values, and many values have more specific leaves. The label next to an attribute names the frozen encoder that reads it. Search to open the paths to every matching node.

Examples

Follow the router's decisions on real benchmark clips

Single-domain and mixed clips from MMAU, MMAR and SAKURA, with the router's real output: the branches each chunk opened, every node it reached with its probability, and the exact text the reader received. Use Next to go one decision at a time.

0 s