TSRGet The Daily Report
The Singularity Report

TSR Desk · science · 19 September 2026, 01:00 UTC

FitAQA: A Benchmark of Fitness Action Quality Assessment for Multimodal Large Language Models

What
FitAQA: A Benchmark of Fitness Action Quality Assessment for Multimodal Large Language Models
Who
arxiv.org
When
18 September 2026, 04:00 UTC
Category
Science
Primary source
https://arxiv.org/abs/2608.08736
What is not known
This brief does not claim independent replication. Claims that appear only on X and not in the primary source stay unknown.

Controlled experiments further indicate that visual perception is a key bottleneck, as judgement performance improves substantially when ground-truth perceptual evidence is provided. It comes from a paper posted to arXiv on 18 September 2026. Fitness Action Quality Assessment (AQA) is important for intelligent sports training, yet the capabilities of Multimodal Large Language Models (MLLMs) in this setting remain underexplored. Existing benchmarks rely on action-specific annotation schemes and focus primarily on final assessment outputs, offering limited insight into how models assess exercise quality. We introduce FitAQA, a systematic benchmark for evaluating MLLMs in fitness AQA, containing 2,219 videos and 5,512 QA instances across 30 bodyweight exercises. In collaboration with experts in sports science, we develop a unified form error taxonomy that defines 38 recurring form errors within six complementary quality dimensions: alignment, symmetry, stability, coordination, tempo, and completeness. This taxonomy provides a shared assessment framework across different exercises. FitAQA further formulates three evaluation tasks: perception for recognizing relevant visual evidence, judgement for combining that evidence with domain knowledge to assess execution correctness, and temporal grounding for localizing form errors over time. Extensive evaluation shows that current MLLMs still struggle to assess exercise quality comprehensively and localize form errors precisely. The dataset is available at https://huggingface.co/datasets/Kelly0510/FitAQA.

Why it counts

Controlled experiments further indicate that visual perception is a key bottleneck, as judgement performance improves substantially when ground-truth perceptual evidence is provided.

Sources

Primary source: primary source

What is not known

This brief does not claim independent replication. Claims that appear only on X and not in the primary source stay unknown.

No clip. The article still stands.