Optimistic Online LQR via Intrinsic Rewards
Marcell Bartos, Bruce D. Lee, Lenart Treven, Andreas Krause, Florian Dörfler, Melanie N. Zeilinger
eess.SY, cs.LG, cs.SY, math.OC
Submitted: 2026-08-21
Updated: 2026-08-24
Code: https://github.com/lenarttreven/lqr_research
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- Playing Atari with Deep Reinforcement Learning
- Training Agents Inside of Scalable World Models
- $\pi^{*}_{0.6}$: a VLA That Learns From Experience
- Policy Gradient Adaptive Control for the LQR: Indirect and Direct Approaches
- Stability of Certainty-Equivalent Adaptive LQR for Linear Systems with Unknown Time-Varying Parameters
- Harnessing Uncertainty for a Separation Principle in Direct Data-Driven Predictive Control
- Gaussian Process Optimization in the Bandit Setting: No Regret and Experimental Design
Related papers
- One Request, Multiple Experts: LLM Orchestrates Domain Specific Models via Adaptive Task Routing
- A Geometric Decision Procedure for STL Feasibility and Repair
- Submodular Multi-Agent Policy Learning for Online Distributed Task Allocation in Open Multi-Agent Systems
- Policy-Level Recursive Self-Improvement for Embodied AI with a Criticality World Model
- Minimal Experiments for Robust Stabilization: Information, Spectral Geometry, and Duration
- Decentralized Power-Optimal Coordination for Spacecraft Swarms Using Time-Varying Magnetorquer Actuation