Uncovering the Computational Ingredients of Human-Like Representations in LLMs
cs.AI
Submitted: 2025-10-01
Updated: 2026-08-31
Comments: 9 pages
License: http://creativecommons.org/licenses/by/4.0/
The gist: The human ability to translate diverse perceptual and linguistic inputs into structured behavior has been thought to rest on learning robust representations of concepts.
Terminology
Abstract
The human ability to translate diverse perceptual and linguistic inputs into structured behavior has been thought to rest on learning robust representations of concepts. The rapid advancement of transformer-based large language models (LLMs) has surfaced a diversity of computational ingredients relevant for model building - architectures, fine-tuning methods, and training datasets among others - yet it remains unclear which are most crucial for developing human-like conceptual representations. Further, most current benchmarks are ill-suited to measuring representational alignment, making LLMs' scores on them unreliable for assessing whether they are progressing as cognitive models. We address these limitations by evaluating over 75 models on a triplet similarity task, a method well established in cognitive science for measuring conceptual representations, using concepts from the THINGS database. We find that instruction fine-tuning and larger attention head dimensionality are among the strongest predictors of human alignment, while activation function choice, multimodal pretraining, and parameter size have limited influence on alignment. Correlations between alignment scores and existing benchmark scores reveal that while some benchmarks (e.g., BigBenchHard) better capture representational alignment than others (e.g., MUSR), none fully accounts for the variance in human-model alignment, demonstrating their insufficiency. Taken together, our findings highlight key computational ingredients for advancing LLMs as models of human conceptual representation and address a key gap in LLM evaluation.
Sources
- Yi: Open Foundation Models by 01.AI
- Representation Topology Divergence: A Method for Comparing Neural Network Representations
- Evaluating Large Language Models Trained on Code
- Scaling Instruction-Finetuned Language Models
- Mamba: Linear-Time Sequence Modeling with Selective State Spaces
- Scaling Vision Transformers to 22 Billion Parameters
- Adversarial Robustness as a Prior for Learned Representations
- OLMo: Accelerating the Science of Language Models
- Textbooks Are All You Need
- Language Models Represent Space and Time
- I-CTRL: Imitation to Control Humanoid Robots Through Constrained Reinforcement Learning
- Measuring Massive Multitask Language Understanding
- Do Androids Laugh at Electric Sheep? Humor "Understanding" Benchmarks from The New Yorker Caption Contest
- The Platonic Representation Hypothesis
- Finite Sample Prediction and Recovery Bounds for Ordinal Embedding
- Active Ranking using Pairwise Comparisons
- SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
- Scaling Laws for Neural Language Models
- Dissociating language and thought in large language models
- Predicting Human Similarity Judgments Using Large Language Models
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection