Available but Unclaimed: An Empirical Study of Human-AI Synergy
cs.HC, cs.AI
Submitted: 2026-09-15
Updated: 2026-09-15
Comments: 31 pages, including appendices
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
The gist: People increasingly reason with large language models (LLMs), yet complementary capabilities do not guarantee outperforming both components.
Terminology
Abstract
People increasingly reason with large language models (LLMs), yet complementary capabilities do not guarantee outperforming both components. In a between-subjects study, participants (N=535) solved a 40-item battery of matrix reasoning, mental rotation, syllogisms, and letter-string analogies, unaided or with GPT-5.6-Luna, Claude Opus 4.8, Gemini 3.6 Flash, or Kimi K3. Each assisted trial required consultation with the model. Each model answered every item alone 100 times under matched elicitation. The assisted-unaided accuracy difference increased with item-level LLM competence. Deference varied across tasks and increased with competence within tasks. Post-advice confidence distinguished correct from incorrect answers less strongly than unaided confidence. In a reference comparison, about half the increase in LLM accuracy carried through to assisted accuracy. How much of that accuracy gain reached participants differed across the models. These findings motivate evaluating LLMs in interaction with humans and designing support for selective deference that preserves independent reasoning.
Sources
- Measuring what Matters: Construct Validity in Large Language Model Benchmarks
- On the Measure of Intelligence
- Adding Error Bars to Evals: A Statistical Approach to Language Model Evaluations
- Does Spatial Cognition Emerge in Frontier Models?
- People over trust AI-generated medical responses and view them to be as valid as doctors, despite low accuracy
- Emergent Analogical Reasoning in Large Language Models
Related papers
- EduGage: A Multimodal Dataset and Benchmark for Sensor-Based Momentary Assessment of Engagement in Self-Guided Video Learning
- EvoDesign: Agentic Editable Diagram Creation via Design Expertise Evolution
- HAGI++: Head-Assisted Gaze Imputation and Generation
- Linking Behaviour and Perception to Evaluate Meaningful Human Control over Partially Automated Driving
- Review of Explainable Decision Support and Adaptive Human-Machine Interfaces for Automation Transparency in Maritime Autonomous Surface Ships
- Towards Cognitive Process-Aware Proactive Writing Support