Benchmarking Audio-Based Empathetic Response Judgment in Large Audio Language Models
Context: "Okay. Um, you don't sound very much like you think that's gonna work."
Speaker: " Yeah. It's, you know, it's one thing to control my own drinking, and it's another thing when everyone else around me is drinking and then there's like that pressure -of, "Oh, they're gonna think I look stupid.""
Listener: "Okay. So, how are you gonna quit drinking without your friends thinking you look stupid?"
Annotations: Emotional Reaction: 0 Exploration: 1 Interpretation: 1 Genuineness: 1 Prosody: 1 Warmth: 1
Context: "You wanna write everybody's story for them?"
Speaker: "I really think I'd be a shitty writer. I'd want to give everybody a happy ending."
Listener: "I'm very sorry about your patient."
Annotations: Emotional Reaction: 2 Exploration: 0 Interpretation: 2 Genuineness: 2 Prosody: 2 Warmth: 2
Context: "Did you call the police?"
Speaker: "No but if it happens again I will. I ordered a camera for my porch. I couldn't get any sleep that night."
Listener: "Do you live in a house or an apartment?"
Annotations: Emotional Reaction: 0 Exploration: 1 Interpretation: 0 Genuineness: 1 Prosody: 1 Warmth: 0
The benchmark, prompts, and evaluation scripts will be publicly released upon publication.
📦 Download Dataset (ZIP)(placeholder only)
💻 GitHub Repository(placeholder only)
@article{anonymous2025baempbench,
title={BAEmpBench: Benchmarking Audio-Based Empathetic Response Judgment in Large Audio Language Models},
author={Anonymous},
journal={arXiv preprint},
year={2025}
}
Note: This is an anonymized preview for review. Full dataset will be released upon publication.