DAYJOB: A Benchmark for Long-Horizon Professional Work
cs.AI, cs.CL
Submitted: 2026-10-01
Updated: 2026-10-01
Code: https://github.com/surge-ai/dayjob
Terminology
Sources
- HealthBench: Evaluating Large Language Models Towards Improved Human Health
- Finance Agent Benchmark: Benchmarking LLMs on Real-world Financial Research Tasks
- GDP.pdf: Benchmarking Grounded Multimodal Reasoning over Professional PDF Documents
- Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains
- CRMArena-Pro: Holistic Assessment of LLM Agents Across Diverse Business Scenarios and Interactions
- FinanceBench: A New Benchmark for Financial Question Answering
- Remote Labor Index: Measuring AI Automation of Remote Work
- ComplexConstraints and Beyond: Expert Rubrics for RLVR
- EnterpriseBench Corecraft: Training Generalizable Agents on High-Fidelity RL Environments
- HANDBOOK.md: A Benchmark for Long-Context Agentic Instruction Following
- GDPval: Evaluating AI Model Performance on Real-World Economically Valuable Tasks
- AgentClinic: a multimodal agent benchmark to evaluate AI in simulated clinical environments
- The AI Productivity Index (APEX)
- The OpenHands Software Agent SDK: A Composable and Extensible Foundation for Production Agents
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection