From Weak Data to Strong Policy: Q-Targets Enable Provable In-Context Reinforcement Learning

arXiv:2609.30391 · cs.LG · Submitted 2026-09-24 · Read on arXiv

cs.LG

Submitted: 2026-09-24

Updated: 2026-09-24

Terminology

Sources

Related papers