SCLATE: a Substrate for Continual-Learning Agent Training and Evaluation
cs.AI, cs.DC, cs.LG
Submitted: 2026-09-26
Updated: 2026-09-26
Code: https://github.com/garrytan/gbrain
Terminology
Sources
- SkyRL-Agent: Efficient RL Training for Multi-turn LLM Agent
- AMemGym: Interactive Memory Benchmarking for Assistants in Long-Horizon Conversations
- Mem0: Building Production-Ready AI Agents with Scalable Long-Term Memory
- Gaia2: Benchmarking LLM Agents on Dynamic and Asynchronous Environments
- MemOps: Benchmarking Lifecycle Memory Operations in Long-Horizon Conversations
- Know Me, Respond to Me: Benchmarking LLMs for Dynamic User Profiling and Personalized Responses at Scale
- SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
- Hindsight is 20/20: Building Agent Memory that Retains, Recalls, and Reflects
- Meta-Harness: End-to-End Optimization of Model Harnesses
- HorizonBench: Long-Horizon Personalization with Evolving Preferences
- Evaluating Very Long-Term Conversational Memory of LLM Agents
- Training Software Engineering Agents and Verifiers with SWE-Gym
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- Mem2ActBench: A Benchmark for Evaluating Long-Term Memory Utilization in Task-Oriented Autonomous Agents
- Self-Distillation Enables Continual Learning
- Harbor Adapters and Harbor-Index: Infrastructure and a Curated Meta-Dataset for Large-Scale Agentic Evaluation
- Learning to (Learn at Test Time): RNNs with Expressive Hidden States
- AppWorld: A Controllable World of Apps and People for Benchmarking Interactive Coding Agents
- Voyager: An Open-Ended Embodied Agent with Large Language Models
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection