TSRGet The Daily Report
The Singularity Report

TSR Desk · science · 17 September 2026, 01:00 UTC

FriendBench: Benchmarking Dyadic Familiarity Inference in Humans and Multimodal Large Language

What
FriendBench: Benchmarking Dyadic Familiarity Inference in Humans and Multimodal Large Language Models
Who
arxiv.org
When
16 September 2026, 04:00 UTC
Category
Science
Primary source
https://arxiv.org/abs/2607.29602
What is not known
This brief does not claim independent replication. Claims that appear only on X and not in the primary source stay unknown.

Richer channels help both unequally, and only humans gain from visible behavior on top of speech. It comes from a paper posted to arXiv on 16 September 2026. Reading a social situation often depends on behavior, not words alone. We introduce FriendBench, a benchmark for inferring whether two people are already familiar or are meeting as strangers, from a 20-second clip of a dyadic ice-breaker conversation. Every pair answers the same type of prompt, so only the manner of interaction can reveal the answer. Across text, audio, and video, we compare 26 models from seven companies against matched human panels over 96 balanced dyads. The best model and the human crowd are statistically indistinguishable on accuracy in every modality, but reach it differently: humans stay balanced across the two answers, while the strongest models favor ``stranger.'' This is a difference in effective prior, not in discrimination. We release the stimuli, human ratings, and model predictions.

Why it counts

Richer channels help both unequally, and only humans gain from visible behavior on top of speech.

Sources

Primary source: primary source

What is not known

This brief does not claim independent replication. Claims that appear only on X and not in the primary source stay unknown.

No clip. The article still stands.