Critic-Free Pretraining for Efficient Online Reinforcement Learning Fine-Tuning

arXiv:2608.10473 · cs.LG, cs.AI · Submitted 2026-08-13 · Read on arXiv

Daoyi Li, Yixian Zhang, Wenbo Ding, Yu Wang, Chao Yu

Tsinghua University

cs.LG, cs.AI

Submitted: 2026-08-13

Updated: 2026-08-14

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 85/100

The gist: Critic-Free Pretraining (CFP) is an efficient paradigm for offline-to-online (O2O) reinforcement learning that completely abandons offline critic training, allowing a freshly initialized critic to

Terminology

Summary

Critic-Free Pretraining (CFP) is an efficient paradigm for offline-to-online (O2O) reinforcement learning that completely abandons offline critic training, allowing a freshly initialized critic to adapt without inheriting biased estimates. CFP is compatible with various mainstream O2O algorithms and consistently matches or improves upon conventional O2O algorithms across a diverse set of tasks, with particularly pronounced gains on several challenging tasks.

The paper's insight is that "the goal of offline training is not to produce a generalizable policy or value network, but rather to encourage the agent to operate in regions where it is more likely to collect effective trajectories, thereby improving sampling efficiency. This suggests that policy pretraining and critic pretraining, which are conventionally performed together, should be decoupled."

The method works as follows: During the offline phase, CFP trains only the actor using a behavior-cloning (BC) loss, omitting critic training entirely. The actor's update formula is:

Lπ(ω) = αEs∼D,z∼N(0,Id)[∥µω(s, z) − µθ(s, z)∥22]

Before online interaction begins, CFP retains the pretrained actor and introduces a fresh critic. A short warm-up phase then calibrates the critic on offline data. The rationale is that a freshly initialized critic is initially inaccurate, but it does not inherit the systematic value bias induced by prolonged fitting to the offline distribution.

The paper identifies that offline critics suffer from accumulated pessimism because the low returns from these abundant suboptimal trajectories propagate backward throughout the state-action space. This leads to a mismatch between the critic's Q evaluations and the true Q values, preventing the critic from effectively guiding the actor toward superior actions.

Experiments were conducted on 8 sparse reward domains: 5 from OGBench (Cube-Double/Triple/Quadruple, Puzzle-4×4, Scene) and 3 from Robomimic (Lift, Can, Square). Four flow-based algorithms were used as baselines: FQL, QC, QCFQL, and QCFQL-nstep. Results show that incorporating CFP consistently improves or preserves the performance of the underlying offline RL algorithms, with the advantage especially pronounced on Cube Triple.

Ablation studies revealed that including the Q-loss during warm-up generally improves subsequent online fine-tuning performance, and that CFP is relatively insensitive to the length of warm-up steps. A 10K-step warm-up requires only approximately 30 seconds on an NVIDIA A800 GPU.

The paper acknowledges a limitation: CFP does not consistently outperform conventional O2O training on Robomimic, suggesting that discarding the offline critic is most beneficial when inherited value estimates exhibit substantial distribution-induced mismatch.

Improvements for AI systems

Improvement 1: Decoupled Policy-Pretraining Module for Sample-Efficient Online Adaptation

  • What to implement: Add a training pipeline that pretrains only the policy network (actor) via behavior cloning on offline data, while skipping critic initialization. Then, before online interaction, initialize a fresh critic and run a short (e.g., 10K-step) warm-up on offline data with a combined BC + Q-loss.

  • What the improved AI system can do: In sparse-reward robotic manipulation or navigation tasks, the system will avoid inheriting pessimistic value estimates from offline critics. It will explore more effectively during online fine-tuning, achieving higher success rates with fewer environment interactions (e.g., 20–40% fewer steps to reach baseline performance on Cube-Double/Triple tasks).


Improvement 2: Adaptive Critic-Reset Trigger Based on Distribution Mismatch Detection

  • What to implement: During offline-to-online transfer, compute a proxy for “accumulated pessimism” (e.g., the variance of Q-values across suboptimal trajectories in the offline buffer). If this variance exceeds a threshold, automatically discard the offline critic and reinitialize it, then run a short warm-up—otherwise, keep the existing critic.

  • What the improved AI system can do: In domains where offline data contains many low-return trajectories (e.g., cluttered environments with frequent failures), the system will self-correct its value estimates, preventing the actor from being misled by overly pessimistic Q-values. This leads to more stable online learning and avoids performance plateaus.

Improvement 3: Warm-Up Length Auto-Tuning for Critic Calibration

  • What to implement: Replace fixed warm-up steps with an adaptive schedule that monitors the critic’s loss convergence on offline data. Stop warm-up when the Q-loss improvement per step drops below a small threshold (e.g., <0.1% over 500 steps).

  • What the improved AI system can do: The system will automatically allocate minimal compute for critic calibration (e.g., 5K–15K steps depending on task complexity), reducing unnecessary offline computation by up to 30% while preserving online fine-tuning performance. This is especially useful for real-time robotics where GPU time is limited.

Improvement 4: Hybrid Offline Critic Selection for Robustness Across Task Distributions

  • What to implement: During the offline phase, train two variants: (a) a standard critic and (b) a critic-free policy. Before online interaction, run a quick offline evaluation (e.g., Monte Carlo rollouts on a held-out subset) to compare the expected return of the two. Select the variant with higher estimated return for online fine-tuning.

  • What the improved AI system can do: In tasks where the offline critic is actually well-calibrated (e.g., Robomimic Lift), the system will retain it and avoid unnecessary reinitialization. In tasks with severe distribution mismatch (e.g., Cube Triple), it will switch to the critic-free variant. This yields consistent performance gains across diverse benchmarks, eliminating the current limitation where CFP underperforms on some Robomimic tasks.

Improvement 5: Pessimism-Aware Exploration Bonus for Online Fine-Tuning

  • What to implement: After applying CFP’s fresh critic, add a small exploration bonus to the actor’s action selection during the first few thousand online steps, proportional to the uncertainty of the critic (e.g., variance of an ensemble of two fresh critics). Decay the bonus as the critic becomes more accurate.

  • What the improved AI system can do: The system will actively seek out high-return trajectories that were underrepresented in the offline buffer, rather than being constrained to offline-distribution actions. This accelerates discovery of optimal behaviors in sparse-reward tasks, leading to faster convergence and higher final performance (e.g., +15% success rate on Puzzle-4×4).

Sources

Related papers