Listen-to-Reason (L2R) is an interpretable-by-design pipeline that passes audio to the LLM through an explicit, human-readable tree: small heads on frozen expert encoders map each chunk of a clip to semantically meaningful nodes on the tree (for speech, music and environmental sound), and a frozen text-only LLM answers from these nodes and an ASR transcript without hearing the clip. Every answer can therefore be traced to the nodes and transcript it read.
The same figure as in the paper, following one question through the system.
Click a node to open or close it. Regions hold attributes, each attribute has a fixed menu of values, and many values have more specific leaves. The label next to an attribute names the frozen encoder that reads it. Search to open the paths to every matching node.
Single-domain and mixed clips from MMAU, MMAR and SAKURA, with the router's real output: the branches each chunk opened, every node it reached with its probability, and the exact text the reader received. Use Next to go one decision at a time.