TSR Desk · science · 19 September 2026, 01:00 UTC
A Neuropsychologically Grounded Evaluation of LLM Cognitive Abilities
- What
- A Neuropsychologically Grounded Evaluation of LLM Cognitive Abilities
- Who
- arxiv.org
- When
- 18 September 2026, 04:00 UTC
- Category
- Science
- Primary source
- https://arxiv.org/abs/2603.02540
- What is not known
- This brief does not claim independent replication. Claims that appear only on X and not in the primary source stay unknown.
Furthermore, we observe that complex reasoning is not universally beneficial, whereas simple, human-like strategies yield partial gains. It comes from a paper posted to arXiv on 18 September 2026. Large language models (LLMs) display a unified "general factor" of capability across 10 benchmarks (a finding confirmed by our factor analysis of 156 models), yet they still struggle with simple, trivial tasks for humans. This is because current benchmarks focus on task completion, failing to probe the foundational cognitive abilities that highlight these behaviors. We address this by introducing the NeuroCognition benchmark, grounded in three adapted neuropsychological tests targeting distinct foundational cognitive components: Raven's Progressive Matrices (abstract relational reasoning), Spatial Working Memory (goal-directed spatial updating), and the Wisconsin Card Sorting Test (cognitive flexibility). Our evaluation reveals that while models perform strongly on text, their performance degrades for images and with increased complexity. Comparison with a human baseline shows that LLMs and humans fail at different parts of the same tasks. We also find that NeuroCognition correlates positively with standard general-capability benchmarks, while still measuring distinct cognitive abilities beyond them. Overall, NeuroCognition emphasizes where current LLMs align with human-like intelligence and where they lack core adaptive cognition, showing the potential to serve as a verifiable, scalable source for improving LLMs.
Why it counts
Furthermore, we observe that complex reasoning is not universally beneficial, whereas simple, human-like strategies yield partial gains.
Sources
Primary source: primary source
What is not known
This brief does not claim independent replication. Claims that appear only on X and not in the primary source stay unknown.
No clip. The article still stands.