🎧 BAEmpBench

Benchmarking Audio-Based Empathetic Response Judgment in Large Audio Language Models

📊 Dataset Overview

🔊 Sample Audio with Annotations

Example 1: Anno-MI (Therapy Session)

Context: "Okay. Um, you don't sound very much like you think that's gonna work."
Speaker: " Yeah. It's, you know, it's one thing to control my own drinking, and it's another thing when everyone else around me is drinking and then there's like that pressure -of, "Oh, they're gonna think I look stupid.""
Listener: "Okay. So, how are you gonna quit drinking without your friends thinking you look stupid?"

Annotations: Emotional Reaction: 0 Exploration: 1 Interpretation: 1 Genuineness: 1 Prosody: 1 Warmth: 1

Example 2: MESC (Scripted Therapy)

Context: "You wanna write everybody's story for them?"
Speaker: "I really think I'd be a shitty writer. I'd want to give everybody a happy ending."
Listener: "I'm very sorry about your patient."

Annotations: Emotional Reaction: 2 Exploration: 0 Interpretation: 2 Genuineness: 2 Prosody: 2 Warmth: 2

Example 3: AvaMERG (Synthetic)

Context: "Did you call the police?"
Speaker: "No but if it happens again I will. I ordered a camera for my porch. I couldn't get any sleep that night."
Listener: "Do you live in a house or an apartment?"

Annotations: Emotional Reaction: 0 Exploration: 1 Interpretation: 0 Genuineness: 1 Prosody: 1 Warmth: 0

📥 Availability

The benchmark, prompts, and evaluation scripts will be publicly released upon publication.

📥 Download

📦 Download Dataset (ZIP)(placeholder only)
💻 GitHub Repository(placeholder only)

📄 Citation

@article{anonymous2025baempbench,
    title={BAEmpBench: Benchmarking Audio-Based Empathetic Response Judgment in Large Audio Language Models},
    author={Anonymous},
    journal={arXiv preprint},
    year={2025}
}
    

Note: This is an anonymized preview for review. Full dataset will be released upon publication.