TSRGet The Daily Report
The Singularity Report

TSR Desk · science · 2 October 2026, 01:00 UTC

Who Said What, and Will It Be Remembered? Evaluating Persistent Speaker Attribution Across

What
Who Said What, and Will It Be Remembered? Evaluating Persistent Speaker Attribution Across Meetings
Who
arxiv.org
When
1 October 2026, 04:00 UTC
Category
Science
Primary source
https://arxiv.org/abs/2609.39344
What is not known
This brief does not claim independent replication. Claims that appear only on X and not in the primary source stay unknown.

The benchmark covers five commercial diarize-then-identify cascades, two open academic baselines, and ThyVoice on the full 129-meeting CHiME-8 NOTSOFAR evaluation set in clean and noiseaugmented form, plus CHiME-6. It comes from a paper posted to arXiv on 1 October 2026. Speech transcripts used as long-term memory must preserve both words and stable speaker identities. Existing meeting-transcription metrics either ignore speakers or remap anonymous speakers independently in each recording, so they cannot measure whether the same person retains one identity across meetings. We evaluate persistent speaker attribution with Speaker Identified cpWER (SI-cpWER), which scores a corpus under one global speaker-ID assignment. ThyVoice is our end-to-end reference system; it repairs overlap and gates the evidence used to create and update voiceprints. Requiring persistent identity changes the commercial ranking: ThyVoice records lower SI-cpWER than every evaluated commercial cascade in all three conditions and the lowest mean in the full panel, 47.13 versus 54.75 for the next system. Complementary lexical, diarization, per-recording attribution, and speaker-clustering diagnostics characterize upstream error surfaces in the final attributed record. These results show why persistent attribution must be evaluated directly in systems that reuse conversations across time.

Why it counts

The benchmark covers five commercial diarize-then-identify cascades, two open academic baselines, and ThyVoice on the full 129-meeting CHiME-8 NOTSOFAR evaluation set in clean and noiseaugmented form, plus CHiME-6.

Sources

Primary source: primary source

What is not known

This brief does not claim independent replication. Claims that appear only on X and not in the primary source stay unknown.

No clip. The article still stands.

Who Said What, and Will It Be Remembered? Evaluating Persistent Speaker Attribution Across · The Singularity Report