Just-In-Time Reinforcement Learning: Continual Learning in LLM Agents Without Gradient Updates
cs.LG, cs.AI
Submitted: 2026-01-26
Updated: 2026-09-27
Code: https://github.com/liushiliushi/JitRL
Terminology
Sources
- A Comprehensive Survey of Self-Evolving AI Agents: A New Paradigm Bridging Foundation Models and Lifelong Agentic Systems
- A Survey of Self-Evolving Agents: What, When, How, and Where to Evolve on the Path to Artificial Super Intelligence
- EvoTest: Evolutionary Test-Time Learning for Self-Improving Agentic Systems
- A Definition of AGI
- Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning
- ToRL: Scaling Tool-Integrated RL
- ConfTuner: Training Large Language Models to Express Their Confidence Verbally
- ToolRL: Reward is All Tool Learning Needs
- An evaluation of LLMs for generating movie reviews: GPT-4o, Gemini-2.0 and DeepSeek-V3
- Proximal Policy Optimization Algorithms
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- R1-Searcher: Incentivizing the Search Capability in LLMs via Reinforcement Learning
- Voyager: An Open-Ended Embodied Agent with Large Language Models
- A-MEM: Agentic Memory for LLM Agents
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks