ContractBench: Can LLM Agents Preserve Observation Contracts?
cs.SE, cs.AI
Submitted: 2026-05-17
Updated: 2026-09-27
Code: https://github.com/laude-institute/harbor
Terminology
Sources
- Medical Large Language Model Benchmarks Should Prioritize Construct Validity
- Measuring what Matters: Construct Validity in Large Language Model Benchmarks
- Agent-R1: A Unified and Modular Framework for Agentic Reinforcement Learning
- Can LLM Agents Solve Collaborative Tasks? A Study on Urgency-Aware Planning and Coordination
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks
- Self-Refine: Iterative Refinement with Self-Feedback
- Qwen2.5 Technical Report
- Reflexion: Language Agents with Verbal Reinforcement Learning
- SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering
- ReAct: Synergizing Reasoning and Acting in Language Models
- NATURAL PLAN: Benchmarking LLMs on Natural Language Planning
Related papers
- Falsification-Based Verification of LLM-Generated Optimization Models: Sound Test Batteries and Their Detection Limits
- GitSkills: A Dataset of Agent Skills on GitHub
- SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces
- PackMonitor: Enabling Zero Package Hallucinations Through Decoding-Time Monitoring
- IntentCoding: Amplifying User Intent in Code Generation
- Incentives and Outcomes in Bug Bounties