Queer inclusion in speech datasets: An audit and taxonomy of practical tensions
arXiv cs.AIen
arXiv:2609.25491v1 Announce Type: new Abstract: In this paper, we examine speech datasets for their inclusion of LGBTQIA+, or queer, voices and provide a taxonomy of tensions to better understand why there is a lack of such voices in current speech technology datasets. Through an audit of six diverse speech datasets, we find that measurable queer representation is low (0-1.4% of speakers) - insufficient for robust disparity measurement. We take this community as a case study to consider what challenges and tensions are associated with collecting speech data from marginalized communities. For comparison, we audit an additional two datasets from the speech sciences that were created by, for, a
This is a short summary published by AI Global Wire. The full article is owned and hosted by arXiv cs.AI — open it there to read it in full.
Read the full story at arXiv cs.AI- Forskning
Related AI news
- 4DGS-JEPA: Temporally Compositional Joint-Embedding Prediction for Dynamic Gaussian SplattingarXiv cs.AI · September 23, 2026
- Lean Pool: An AI-Maintained Archive of Formalized MathematicsarXiv cs.AI · September 23, 2026
- Making Agents More Consistent: Skills Should Form Habits for Repeat TasksarXiv cs.AI · September 23, 2026
- Efficient Iterative Retrieval with Heterogeneous BatchingarXiv cs.AI · September 23, 2026
- Attention as a Routing Graph: Live Circuit Extraction from a Single Forward PassarXiv cs.AI · September 23, 2026
- SMTB: Fast Structure-Mapping with Tight BoundsarXiv cs.AI · September 23, 2026