FrontierChallenge: Evaluating Scientific Workflow Completion
cs.AI, cs.CL, cs.SE
Submitted: 2026-08-25
Updated: 2026-09-09
Code: https://github.com/ApodexAI/FrontierAgent
Terminology
Sources
- Apodex 1.1: Scaling Agentic Intelligence for Complex Work
- BioAgent Bench: An AI Agent Evaluation Suite for Bioinformatics
- Kimi K3: Open Frontier Intelligence
- LAB-Bench: Measuring Capabilities of Language Models for Biology Research
- Benchmarking AI Agents for Addressing Scientific Challenges Across Scales
- The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery
- BixBench: a Comprehensive Benchmark for LLM-based Agents in Computational Biology
- Humanity's Last Exam
- SciAgentGym: Benchmarking Multi-Step Scientific Tool-use in LLM Agents
- CORE-Bench: Fostering the Credibility of Published Research Through a Computational Reproducibility Agent Benchmark
- PaperBench: Evaluating AI's Ability to Replicate AI Research
- Agents' Last Exam
- NatureBench: Can Coding Agents Match the Published SOTA of Nature-Family Papers?
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection