Using profiles of cognitive capability to assess AI suitability for workplace tasks
cs.AI, cs.CY, cs.HC
Submitted: 2026-08-26
Updated: 2026-08-26
Code: https://github.com/Kinds-of-Intelligence-CFI/Task-Suitability-Profiles
Terminology
Sources
- Towards a Science of AI Agent Reliability
- Measuring Agents in Production
- The Illusion of Readiness in Health AI
- Line Goes Up? Inherent Limitations of Benchmarks for Evaluating Large Language Models
- Artificial Intelligence Index Report 2025
- AI and the Everything in the Whole Wide World Benchmark
- Cognitive Dark Matter: Measuring What AI Misses
- Inferring Capabilities from Task Performance with Bayesian Triangulation
- Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks
- Visuospatial Perspective Taking in Multimodal Language Models
- From Human-Level AI Tales to AI Leveling Human Scales
- Benchmark Data Contamination of Large Language Models: A Survey
- Which Economic Tasks are Performed with AI? Evidence from Millions of Claude Conversations
- LLM-BabyBench: Can Language Models Plan in Worlds They Can Simulate?
- Capabilities Ain't All You Need: Measuring Propensities in AI
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection