CARM: Cancellation-Aware Response Masking for LLM Reinforcement Learning

arXiv:2610.02039 · cs.LG, cs.AI, cs.CL · Submitted 2026-10-01 · Read on arXiv

cs.LG, cs.AI, cs.CL

Submitted: 2026-10-01

Updated: 2026-10-01

Code: https://github.com/huggingface/Math-Verify

Terminology

Sources

Related papers