TSRGet The Daily Report
The Singularity Report

TSR Desk · science · 24 September 2026, 01:00 UTC

Queer inclusion in speech datasets: An audit and taxonomy of practical tensions

What
Queer inclusion in speech datasets: An audit and taxonomy of practical tensions
Who
arxiv.org
When
23 September 2026, 04:00 UTC
Category
Science
Primary source
https://arxiv.org/abs/2609.25491
What is not known
This brief does not claim independent replication. Claims that appear only on X and not in the primary source stay unknown.

In this paper, we examine speech datasets for their inclusion of LGBTQIA+, or queer, voices and provide a taxonomy of tensions to better understand why there is a lack of such voices in current speech technology datasets. It comes from a paper posted to arXiv on 23 September 2026. Through an audit of six diverse speech datasets, we find that measurable queer representation is low (0-1.4% of speakers) - insufficient for robust disparity measurement. We take this community as a case study to consider what challenges and tensions are associated with collecting speech data from marginalized communities. For comparison, we audit an additional two datasets from the speech sciences that were created by, for, and with the queer community. We note that many customs in speech dataset collection efforts in AI and speech technology research may conflict with values emphasized in participatory approaches with marginalized communities, and provide a taxonomy describing these tensions.

Why it counts

Through an audit of six diverse speech datasets, we find that measurable queer representation is low (0-1.4% of speakers) - insufficient for robust disparity measurement.

Sources

Primary source: primary source

What is not known

This brief does not claim independent replication. Claims that appear only on X and not in the primary source stay unknown.

No clip. The article still stands.