PB-GRPO: Learning Socially Adaptive LLM Agents from Persona-Driven Simulation with Preference-Batched GRPO

arXiv:2610.04132 · cs.CL, cs.AI · Submitted 2026-10-02 · Read on arXiv

cs.CL, cs.AI

Submitted: 2026-10-02

Updated: 2026-10-02

Terminology

Related papers