Enhancing Audio Reasoning via Semantic Summary Prediction

Francesco Bonzi1, 2, , Pooneh Mousavi1, 2, Cem Subakan3, 2, Mirco Ravanelli1, 2
1 Concordia University, Canada   |   2 Mila - Quebec AI Institute, Canada   |   3 Université Laval, Canada
GitHub

Abstract

Large Audio Language Models (LALMs) perform well on complex question answering but often show a reasoning gap, where explicit Chain-of-Thought (CoT) reduces accuracy compared to direct answers. We hypothesize that long reasoning sequences shift attention away from the audio input. To address this, we propose SPARE (Semantic Prediction for Audio REasoning), which introduces a register token aligned with the final conclusion using a cosine similarity loss with a Sentence-BERT embedding. This conditions the model’s latent space with the target semantic goal before reasoning begins. Experiments on MMAU and MMAR with SALMONN show improved zero-shot reasoning and stronger early attention to audio without additional inference cost.

SPARE Pipeline

SPARE Pipeline Method

Audio & reasoning demos

Curated MMAU-style clips where SPARE matches the gold label and SALMONN (pretrained), SALMONN (SFT), and Audio MuToR do not, mirroring the qualitative analysis in our CoT write-up. In each card, compare the CAPTION text: SPARE typically extracts a longer and/or more meaningful, audio-grounded caption than the baselines—more concrete vocabulary tied to what is actually heard—before reasoning and conclusion. For each clip, model outputs are shown in that order; open a row to read the full CAPTION, REASONING, and CONCLUSION.

Multi-speaker role mapping

🎙️ Speech

Task: Identify how the two speakers are connected.

SPARE’s CAPTION is longer and more grounded in the audio than the baselines here: it encodes what is said (rising healthcare costs) and who speaks when (male, then neutral-toned female). SFT and Audio MuToR stay with short, generic scene labels (“speech in English” or a single neutral male), so the chain never locks onto the policy content that supports politician–voter.

SALMONN (pretrained)WRONG
<CAPTION>
[Not emitted: this pretrained run has no explicit CAPTION tag.]
</CAPTION>

<REASONING>
[Not emitted: this pretrained run has no explicit REASONING tag.]
</REASONING>

<CONCLUSION>
bar owner-regular customer
</CONCLUSION>

— Verbatim output —
<s> 
bar owner-regular customer</s><unk><unk><unk><unk>
SALMONN (SFT)WRONG
<CAPTION>
A male voice delivers a speech in English.
</CAPTION>

<REASONING>
The two speakers are connected as a bar owner and a regular customer. The bar owner provides goods or services to the regular customer, establishing a business-customer relationship. This aligns with option (A) bar owner-regular customer.
</REASONING>

<CONCLUSION>
bar owner-regular customer
</CONCLUSION>
Audio MuToRWRONG
<CAPTION>
A male voice speaks in a neutral tone.
</CAPTION>

<REASONING>
The two speakers are connected as a pet groomer (speaker 1) and a pet owner (speaker 2). The groomer provides services to the pet owner, establishing a professional-client relationship.
</REASONING>

<CONCLUSION>
pet groomer-pet owner
</CONCLUSION>
SPARE (ours)CORRECT
<CAPTION>
A male voice expresses concern about rising healthcare costs, followed by a neutral-toned female voice.
</CAPTION>

<REASONING>
The speakers are a politician and a voter. The politician addresses the voter's concerns about rising healthcare costs, indicating a professional-voter relationship. This aligns with option (D) politician-voter.
</REASONING>

<CONCLUSION>
politician-voter
</CONCLUSION>

Emotion state summarisation

🎙️ Speech

Task: Match each speaker’s emotion to the labeled option.

SPARE’s CAPTION is more specific and better tied to the heard affect than the baselines: it states frustration vs neutral directly. Audio MuToR’s caption is short but misbinds genders to happy/sad; SFT’s line is even sparser (“man and woman … conversation”), which invites a stereotyped sad/happy template—none of which matches the benchmark’s pragmatic affect.

SALMONN (pretrained)WRONG
<CAPTION>
[Not emitted: this pretrained run has no explicit CAPTION tag.]
</CAPTION>

<REASONING>
[Not emitted: this pretrained run has no explicit REASONING tag.]
</REASONING>

<CONCLUSION>
first speaker shows sadness and second speaker shows happiness.
</CONCLUSION>

— Verbatim output —
<s> 
first speaker shows sadness and second speaker shows happiness.</s><unk><unk><unk><unk><unk><unk><unk>
SALMONN (SFT)WRONG
<CAPTION>
A man and a woman engage in a conversation.
</CAPTION>

<REASONING>
The first speaker's tone, pitch, and word choice indicate sadness, while the second speaker's tone, pitch, and word choice indicate happiness. These emotional cues align with option (B), where the first speaker shows sadness and the second speaker shows happiness.
</REASONING>

<CONCLUSION>
first speaker shows sadness and second speaker shows happiness
</CONCLUSION>
Audio MuToRWRONG
<CAPTION>
A woman expresses sadness and a man expresses happiness.
</CAPTION>

<REASONING>
The first speaker's vocal tone, pitch, and volume indicate sadness, while the second speaker's vocal tone, pitch, and volume indicate happiness.
</REASONING>

<CONCLUSION>
first speaker shows happiness and second speaker shows sadness.
</CONCLUSION>
SPARE (ours)CORRECT
<CAPTION>
Two individuals converse, one expressing frustration and the other remaining neutral.
</CAPTION>

<REASONING>
The first speaker's frustration is indicated by their tone, pitch, and word choice, while the second speaker's neutrality is reflected in their tone, pitch, and word choice. These emotional cues align with option (D), as the first speaker shows frustration and the second speaker shows neutrality.
</REASONING>

<CONCLUSION>
first and second speaker show frustration and neutral respectively.
</CONCLUSION>

Primary instrumentation

🎵 Music

Task: Choose the option that best lists the primary instruments or layers.

SPARE’s CAPTION names richer, production-level detail that tracks the mix—groovy bass, shimmering hi-hats, punchy snares, keyboard harmonies—i.e. a longer, more grounded read of the bed. Audio MuToR and SFT fall back on a shorter rock-band inventory (piano, drums, guitar) even when their own reasoning mentions hi-hats, so the perceptual summary under-specifies the texture the gold option targets.

SALMONN (pretrained)WRONG
<CAPTION>
[Not emitted: this pretrained run has no explicit CAPTION tag.]
</CAPTION>

<REASONING>
[Not emitted: this pretrained run has no explicit REASONING tag.]
</REASONING>

<CONCLUSION>
n/a
</CONCLUSION>

— Verbatim output —
<s> n/a</s><unk><unk><unk><unk>
SALMONN (SFT)WRONG
<CAPTION>
Male vocalist performs emotionally over a medium-tempo pop track with piano, drums, guitar, and bass.
</CAPTION>

<REASONING>
The audio features a piano, synth strings, and shimmering hi hats. Among these, the piano is the most prominently featured instrument.
</REASONING>

<CONCLUSION>
Piano, Drums, Guitar
</CONCLUSION>
Audio MuToRWRONG
<CAPTION>
Male vocalist performs emotionally over medium-tempo pop track with piano, drums, guitar, bass, and synth strings.
</CAPTION>

<REASONING>
The audio features a piano, synth strings, and shimmering hi hats. Among these, the piano and synth strings are the most prominently featured instruments.
</REASONING>

<CONCLUSION>
Piano, Drums, Guitar
</CONCLUSION>
SPARE (ours)CORRECT
<CAPTION>
Male vocalist performs emotionally over medium-tempo pop track with groovy bass, shimmering hi hats, punchy snares, and keyboard harmonies.
</CAPTION>

<REASONING>
The audio features a piano, synth strings, and shimmering hi hats. Among these, the piano is the most prominent instrument, as it is the primary melodic and harmonic foundation. Synth strings and shimmering hi hats contribute to the overall texture, but the piano is the central element.
</REASONING>

<CONCLUSION>
Tinny bells, Synth strings, Shimmering hi hats
</CONCLUSION>

Musical texture

🎵 Music

Task: Identify which background sound effect best matches a rain-like description.

SPARE’s CAPTION explicitly labels the background layer the question cares about—white noise behind the vocal—so it is both more meaningful and better grounded than alternatives. Audio MuToR substitutes wind; SFT substitutes piano and steers reasoning toward thunder. Only SPARE’s caption anchors the correct texture term before the reasoning step.

SALMONN (pretrained)WRONG
<CAPTION>
[Not emitted: this pretrained run has no explicit CAPTION tag.]
</CAPTION>

<REASONING>
[Not emitted: this pretrained run has no explicit REASONING tag.]
</REASONING>

<CONCLUSION>
Thunder
</CONCLUSION>

— Verbatim output —
<s> Thunder</s><unk><unk><unk>
SALMONN (SFT)WRONG
<CAPTION>
A female voice sings a melancholic folk tune, backed by soothing piano music, creating a calming atmosphere.
</CAPTION>

<REASONING>
Rain is characterized by a continuous, rhythmic pattern of droplets hitting surfaces. Among the options, thunder (A) is the sound most directly associated with rain, as thunder is caused by lightning striking the atmosphere. Birds chirping (B) and wind blowing (D) produce distinct, non-rain-like sounds, while white noise (C) lacks the rhythmic pattern typical of rain. Thus, thunder is the closest match to the rain-like sound effect.
</REASONING>

<CONCLUSION>
Thunder
</CONCLUSION>
Audio MuToRWRONG
<CAPTION>
A female voice sings a melancholic folk tune, accompanied by soothing wind sounds.
</CAPTION>

<REASONING>
Rain is characterized by rhythmic droplets and rustling leaves. Among the options, wind blowing (D) is the closest match to these sounds, as it can produce a similar rhythmic, rustling effect. Thunder (A) and birds chirping (B) lack the rhythmic droplet pattern, while white noise (C) lacks the rustling leaves. Thus, wind blowing is the most accurate match.
</REASONING>

<CONCLUSION>
Wind blowing
</CONCLUSION>
SPARE (ours)CORRECT
<CAPTION>
A female voice sings a melancholic folk song, accompanied by soothing white noise in the background.
</CAPTION>

<REASONING>
Rain is characterized by a continuous, rhythmic pattern of droplets hitting surfaces. Among the options, white noise (C) is a consistent, unvarying sound that resembles the steady, low-frequency nature of raindrops. Thunder (A) is intermittent and loud, birds chirping (B) is melodic and varied, and wind blowing (D) is irregular and high-pitched. White noise aligns most closely with the steady, low-frequency characteristics of raindrops.
</REASONING>

<CONCLUSION>
White noise
</CONCLUSION>

Continuous sound effect

🔊 Sound

Task: Pick the best description of the sustained / continuous effect among the choices.

SPARE’s CAPTION is a more faithful, grounded scene read than Audio MuToR and SFT: it stresses deep, resonant bass and no speech instead of latching onto a spurious “car horn after the explosion” script. That longer, more spectrally specific summary supports treating the sustained energy as a cinematic intro rather than the wrong intermittent cue.

SALMONN (pretrained)WRONG
<CAPTION>
[Not emitted: this pretrained run has no explicit CAPTION tag.]
</CAPTION>

<REASONING>
[Not emitted: this pretrained run has no explicit REASONING tag.]
</REASONING>

<CONCLUSION>
A single car horn beep
</CONCLUSION>

— Verbatim output —
<s> 
A single car horn beep</s><unk><unk>
SALMONN (SFT)WRONG
<CAPTION>
Loud, low-pitched explosion followed by a single car horn beep.
</CAPTION>

<REASONING>
The audio contains a continuous sound effect. Among the options, a single car horn beep (C) is the only continuous sound effect. Other options (A, B, D) involve intermittent or brief sounds, which do not align with the description of a continuous effect.
</REASONING>

<CONCLUSION>
A single car horn beep
</CONCLUSION>
Audio MuToRWRONG
<CAPTION>
Loud, low-pitched explosion followed by a single car horn beep.
</CAPTION>

<REASONING>
The continuous sound effect in the audio is a single car horn beep. This sound is characterized by its repetitive and consistent nature, which aligns with the definition of a continuous sound effect. Other options like bird chirping (intermittent), cinematic intro (continuous), or dog bark (intermittent) do not match the description of a continuous sound effect.
</REASONING>

<CONCLUSION>
A single car horn beep
</CONCLUSION>
SPARE (ours)CORRECT
<CAPTION>
Loud, intense explosion with deep, resonant bass; no speech.
</CAPTION>

<REASONING>
The audio contains a continuous sound effect. Among the options, a cinematic intro (A) is the only one that aligns with a sustained, non-intermittent sound. The other options (bird chirping, car horn beep, dog bark) are either intermittent (bird chirping, car horn beep) or brief (dog bark). Thus, the continuous sound effect is most likely a cinematic intro.
</REASONING>

<CONCLUSION>
Likely a cinematic intro
</CONCLUSION>

Same caption, different scene

🔊 Sound

Task: Choose the setting whose ongoing and sudden sounds best fit the clip.

Control example: Audio MuToR, SALMONN (SFT), and SPARE share the same CAPTION, so the win here is not from a longer SPARE caption—in the other demos, SPARE’s CAPTION tends to be longer and/or more meaningfully grounded in the audio. With caption parity, the difference is entirely in REASONING: Audio MuToR and SFT argue a subway platform, while SPARE’s reasoning composes the same train cues into an outdoor festival near a railway, matching the gold label in the document.

SALMONN (pretrained)WRONG
<CAPTION>
[Not quoted in the CoT bonus subsection for this clip.]
</CAPTION>

<REASONING>
[Not quoted in the CoT bonus subsection for this clip.]
</REASONING>

<CONCLUSION>
[Not quoted in the CoT bonus subsection for this clip.]
</CONCLUSION>
SALMONN (SFT)WRONG
<CAPTION>
Train horn blares, followed by rhythmic clacking of wheels on tracks.
</CAPTION>

<REASONING>
The ongoing sounds in the audio are the continuous hum of a train engine and background noise. The sudden sounds are the train horn and crowd reactions. Among the options, a subway platform during rush hour (B) aligns with both ongoing (train movement) and sudden (train horn, crowd reactions) sounds.
</REASONING>

<CONCLUSION>
Subway platform during rush hour
</CONCLUSION>
Audio MuToRWRONG
<CAPTION>
Train horn blares, followed by rhythmic clacking of wheels on tracks.
</CAPTION>

<REASONING>
The ongoing sounds include a train horn and the sound of a train moving. These sounds are most characteristic of a subway platform during rush hour, where trains operate frequently and horns are used to signal arrivals and departures. The sudden sounds of a train horn and movement align with the context of a subway platform, distinguishing it from other settings like outdoor festivals, concert halls, or forest trails.
</REASONING>

<CONCLUSION>
Subway platform during rush hour
</CONCLUSION>
SPARE (ours)CORRECT
<CAPTION>
Train horn blares, followed by rhythmic clacking of wheels on tracks.
</CAPTION>

<REASONING>
The ongoing sounds in the audio are the sound of a train horn and the rhythmic clacking of wheels on tracks. These sounds are characteristic of a railway setting. The sudden sound of a train horn is a common feature in such environments. Among the options, an outdoor festival near a railway (A) aligns with both ongoing (train movement) and sudden (train horn) sounds, making it the most fitting setting.
</REASONING>

<CONCLUSION>
Outdoor festival near a railway
</CONCLUSION>

When SPARE overthinks the audio Only 0.13% of cases

On the strict subset of MMAU items where SALMONN (pretrained), SALMONN (SFT), and Audio MuToR all match the gold label, but SPARE does not, we find the following cases.

Over-interpreting lyrical stance

🎵 Music

Task: Lyrical reasoning — what is the speaker’s attitude toward their decision to show up?

SPARE’s caption leans pop ballad with sadness / longing, but the reasoning then re-frames the line toward a positive decision (“best decision”). Everyone else stays on regret / mistake; a case where richer CoT can over-commit to a coherent narrative that conflicts with the gold reading.

SALMONN (pretrained)CORRECT
<CAPTION>
[Not emitted: this pretrained run has no explicit CAPTION tag.]
</CAPTION>

<REASONING>
[Not emitted: this pretrained run has no explicit REASONING tag.]
</REASONING>

<CONCLUSION>
It was a mistake.
</CONCLUSION>

— Verbatim output —
<s> 
It was a mistake.</s><unk><unk><unk><unk><unk><unk><unk><unk><unk><unk><unk><unk><unk>
SALMONN (SFT)CORRECT
<CAPTION>
A female vocalist performs a melancholic rock song with guitar, bass, and drums, expressing sadness and longing.
</CAPTION>

<REASONING>
The speaker's emotional state is described as melancholic, which suggests a negative or dissatisfied mood. This aligns with option (A) 'It was a mistake' as it implies regret or dissatisfaction with the decision to show up.
</REASONING>

<CONCLUSION>
It was a mistake.
</CONCLUSION>
Audio MuToRCORRECT
<CAPTION>
A female vocalist performs a melancholic pop song with piano, bass, and drums.
</CAPTION>

<REASONING>
The speaker's tone and language convey a sense of regret or dissatisfaction, which aligns with option (A) 'It was a mistake.'
</REASONING>

<CONCLUSION>
It was a mistake.
</CONCLUSION>
SPARE (ours)WRONG
<CAPTION>
Female vocalist performs emotionally in a pop ballad with piano, bass, and drums, conveying sadness and longing.
</CAPTION>

<REASONING>
The speaker's emotional state is conveyed through their words and actions, which suggest a positive attitude towards their decision to show up. The phrase 'I knew you didn't care' implies that the speaker was aware of the other person's indifference, but their own attitude was positive. This aligns with the option 'It was the best decision.'
</REASONING>

<CONCLUSION>
It was the best decision.
</CONCLUSION>

Missing fact retrieval

🎙️ Speech

Task: Event-based knowledge retrieval — identify which organ was transplanted in the procedure described in the audio.

Both kidney and liver are plausible in medical talk; SPARE’s reasoning invents a “most commonly transplanted” story and lands on liver even though SFT’s chain quotes the line about the first successful kidney transplant. CoT does not grant reliable anatomy when the evidence has to be pinned to an exact spoken fact.

SALMONN (pretrained)CORRECT
<CAPTION>
[Not emitted: this pretrained run has no explicit CAPTION tag.]
</CAPTION>

<REASONING>
[Not emitted: this pretrained run has no explicit REASONING tag.]
</REASONING>

<CONCLUSION>
Kidney
</CONCLUSION>

— Verbatim output —
<s> Kidney</s><unk><unk><unk><unk><unk><unk><unk>
SALMONN (SFT)CORRECT
<CAPTION>
A male speaker discusses a successful organ transplant, focusing on its significance in medical history.
</CAPTION>

<REASONING>
The speaker explicitly states 'the first successful organ transplant was a kidney transplant.' This directly identifies the organ transplanted as a kidney.
</REASONING>

<CONCLUSION>
Kidney
</CONCLUSION>
Audio MuToRCORRECT
<CAPTION>
A neutral-toned male speaker discusses a successful organ transplant.
</CAPTION>

<REASONING>
First, identify the organs commonly transplanted (heart, lung, kidney, liver). Next, determine which organ is explicitly mentioned in the audio. Finally, match the mentioned organ to the correct option.
</REASONING>

<CONCLUSION>
Kidney
</CONCLUSION>
SPARE (ours)WRONG
<CAPTION>
A female speaker discusses a successful organ transplant in medical history.
</CAPTION>

<REASONING>
The speaker mentions a 'successful organ transplant,' indicating the transplant was successful. Among the listed organs, the liver is most commonly transplanted due to its high availability and compatibility with donors. Therefore, the correct answer is (D) Liver.
</REASONING>

<CONCLUSION>
Liver
</CONCLUSION>

Citation (arXiv)

@article{your_name202Xspare,
  title={Enhancing Audio Reasoning via Semantic Summary Prediction},
  author={Firstname Lastname and Firstname Lastname},
  journal={arXiv preprint arXiv:XXXX.XXXXX},
  year={2026}
}