PhenoBench: Mapping What a Deeply Phenotyped Human Cohort Can Tell Us
cs.LG, cs.AI
Submitted: 2026-09-05
Updated: 2026-09-23
Comments: 33 pages; 5 main figures and 5 supplementary figures. Project website: https://galsapir.github.io/phenobench-benchmark/ . Code and benchmark materials: https://github.com/galsapir/phenobench-benchmark
Code: https://github.com/galsapir/phenobench-benchmark
Project page: https://galsapir.github.io/phenobench-benchmark
License: http://creativecommons.org/licenses/by/4.0/
The gist: Deeply phenotyped cohorts combine clinical, imaging, molecular, and wearable observations across timescales from seconds to years.
Terminology
Abstract
Deeply phenotyped cohorts combine clinical, imaging, molecular, and wearable observations across timescales from seconds to years. This breadth can reveal which measurements inform which health-related questions, but heterogeneous analyses are not directly comparable. We present PhenoBench, an executable benchmark built around the Human Phenotype Project, in which more than 13,000 participants have completed the initial visit. Each question fixes the target, eligible population, timing, and allowed information; its evaluation contract specifies the split, metric, baseline, and claim boundary. The benchmark defines 90 clinically grounded tasks across 15 domains and 26 input modalities. Measurements showed question- and representation-dependent predictive value, including positive, near-zero, and negative changes in held-out performance relative to matched baselines. We used PhenoBench to evaluate emerging tabular foundation models across 160 matched regression comparisons spanning 52 tasks. These models ranked above standard task-specific models in aggregate but, averaged across the three pretrained models within each cell, improved on ridge by a median of only 0.004 R squared (95% CI, 0.002--0.006). We then used the same cohort data and evaluation contracts to evaluate 14 language models, collectively covering 40 tasks spanning phenotype recovery, classification, follow-up forecasting, and participant ordering. Without cohort-specific fitting, language models made informative predictions on some tasks, but showed task-specific capability gaps, shared failures of scale, and rarely surpassed models fitted on the same fields. PhenoBench turns a multimodal longitudinal cohort into a versioned, auditable evaluation system where new questions, measurements, and models can be added without redefining existing comparisons.
Sources
- AI and the Everything in the Whole Wide World Benchmark
- TabICL: A Tabular Foundation Model for In-Context Learning on Large Data
- Chronos-2: From Univariate to Universal Forecasting
- A Multimodal Dataset of 21,412 Recorded Nights for Sleep and Respiratory Research
- The Last Human-Written Paper: Agent-Native Research Artifacts
- TabSwift: An Efficient Tabular Foundation Model with Row-Wise Attention
- EHRSHOT: An EHR Benchmark for Few-Shot Evaluation of Foundation Models
- Better by Default: Strong Pre-Tuned MLPs and Boosted Trees on Tabular Data
- TabDPT: Scaling Tabular Foundation Models on Real Data
- MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks