BioEVAL: A global, multi-institutional benchmark of large language and multimodal models for bioengineering
cs.AI
Submitted: 2026-09-24
Updated: 2026-09-24
Terminology
Sources
- AI and the Everything in the Whole Wide World Benchmark
- Holistic Evaluation of Language Models
- A Survey of Large Language Models in Medicine: Progress, Application, and Challenge
- LAB-Bench: Measuring Capabilities of Language Models for Biology Research
- BioProBench: A Corpus and Benchmark for Biological Protocol Reasoning in Autonomous Science
- GPQA: A Graduate-Level Google-Proof Q&A Benchmark
- Evaluating Large Language Models in Scientific Discovery
- SciBench: Evaluating College-Level Scientific Problem-Solving Abilities of Large Language Models
- PathVQA: 30000+ Questions for Medical Visual Question Answering
- Measuring Massive Multitask Language Understanding
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection