Agentic Reinforcement Learning with Observation-Calibrated Self-Distillation

arXiv:2608.04788 · cs.LG, cs.AI, cs.CL · Submitted 2026-08-05 · Read on arXiv

Yi Yang, Cong Qin, Xiaodan Liu, Chishui Chen, Qing Dong, Yan Zhang, Cao Liu, Zhao Yang, Lu Pan, Jiaye Lin, Yi Feng

cs.LG, cs.AI, cs.CL

Submitted: 2026-08-05

Code: https://github.com/yiy1x/OCSD

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Terminology

Sources

Related papers