A Neuropsychologically Grounded Evaluation of LLM Cognitive Abilities
cs.AI
Submitted: 2026-03-03
Updated: 2026-09-17
Comments: 33 pages, 7 figures, 10 tables
Code: https://github.com/reggans/CognitiveEval
License: http://creativecommons.org/licenses/by-sa/4.0/
The gist: Large language models (LLMs) display a unified "general factor" of capability across 10 benchmarks (a finding confirmed by our factor analysis of 156 models), yet they still struggle with simple,
Terminology
Abstract
Large language models (LLMs) display a unified "general factor" of capability across 10 benchmarks (a finding confirmed by our factor analysis of 156 models), yet they still struggle with simple, trivial tasks for humans. This is because current benchmarks focus on task completion, failing to probe the foundational cognitive abilities that highlight these behaviors. We address this by introducing the NeuroCognition benchmark, grounded in three adapted neuropsychological tests targeting distinct foundational cognitive components: Raven's Progressive Matrices (abstract relational reasoning), Spatial Working Memory (goal-directed spatial updating), and the Wisconsin Card Sorting Test (cognitive flexibility). Our evaluation reveals that while models perform strongly on text, their performance degrades for images and with increased complexity. Comparison with a human baseline shows that LLMs and humans fail at different parts of the same tasks. Furthermore, we observe that complex reasoning is not universally beneficial, whereas simple, human-like strategies yield partial gains. We also find that NeuroCognition correlates positively with standard general-capability benchmarks, while still measuring distinct cognitive abilities beyond them. Overall, NeuroCognition emphasizes where current LLMs align with human-like intelligence and where they lack core adaptive cognition, showing the potential to serve as a verifiable, scalable source for improving LLMs.
Sources
- Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
- Qwen3-VL Technical Report
- Qwen2.5-VL Technical Report
- GPT-4 Technical Report
- Revealing the structure of language model capabilities
- On the Measure of Intelligence
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- LLM Task Interference: An Initial Study on the Impact of Task-Switch in Conversational History
- Visual Large Language Models Exhibit Human-Level Cognitive Flexibility in the Wisconsin Card Sorting Test
- Language models show human-like content effects on reasoning tasks
- Cartridges: Lightweight and general-purpose long context representations via self-study
- Large Language Models Are Not Strong Abstract Reasoners
- LLMs Get Lost In Multi-Turn Conversation
- MSCoRe: A Benchmark for Multi-Stage Collaborative Reasoning in LLM Agents
- Working Memory Capacity of ChatGPT: An Empirical Study
- LatestEval: Addressing Data Contamination in Language Model Evaluation through Dynamic and Time-Sensitive Test Construction
- CriticBench: Benchmarking LLMs for Critique-Correct Reasoning
- MMDU: A Multi-Turn Multi-Image Dialog Understanding Benchmark and Instruction-Tuning Dataset for LVLMs
- OpenAI GPT-5 System Card
- Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection