TSRGet The Daily Report
The Singularity Report

TSR Desk · science · 1 October 2026, 01:00 UTC

HEAR: Real Voices, Real Bias: A Large-Scale Human-Recorded, Demographically Diverse Benchmark

What
HEAR: Real Voices, Real Bias: A Large-Scale Human-Recorded, Demographically Diverse Benchmark for Audio Language Models
Who
arxiv.org
When
30 September 2026, 04:00 UTC
Category
Science
Primary source
https://arxiv.org/abs/2609.35952
What is not known
This brief does not claim independent replication. Claims that appear only on X and not in the primary source stay unknown.

To our knowledge, this is the first large-scale voice benchmark grounded entirely in authentic human speech. It comes from a paper posted to arXiv on 30 September 2026. We introduce HEAR (Human-recorded Evaluation of Audio-LLM bias by Real speakers), a large-scale, ecologically valid benchmark comprising 87k real human audio samples from 843 demographically diverse participants. HEAR enables comprehensive evaluation through Multiple Choice Question Answering (MCQA) and open-ended long-form tasks. We evaluate model behavior across both real-time speech-to-speech and speech-to-text architectures. Our results reveal that voice-conditioned bias is a model-specific property. Furthermore, we demonstrate that personalization instructions consistently exacerbate demographic disparities. Our findings establish that voice bias is a controllable model characteristic, providing a foundational framework for future bias mitigation and evaluation in Audio-LLM development.

Why it counts

To our knowledge, this is the first large-scale voice benchmark grounded entirely in authentic human speech.

Sources

Primary source: primary source

What is not known

This brief does not claim independent replication. Claims that appear only on X and not in the primary source stay unknown.

No clip. The article still stands.