Certified Predictive Value-of-Advice Gating for Cost-Aware Language-Model Guidance in Reinforcement Learning
cs.LG
Submitted: 2026-08-26
Updated: 2026-08-26
Terminology
Sources
- Do As I Can, Not As I Say: Grounding Language in Robotic Affordances
- FrugalML: How to Use ML Prediction APIs More Accurately and Cheaply
- FrugalGPT: How to Use Large Language Models While Reducing Cost and Improving Performance
- Guiding Pretraining in Reinforcement Learning with Large Language Models
- Selecting Computations: Theory and Applications
- Enabling Intelligent Interactions between an Agent and an LLM: A Reinforcement Learning Approach
- Reward Design with Language Models
- RouteLLM: Learning to Route LLMs with Preference Data
- Randomized Prior Functions for Deep Reinforcement Learning
- Deep Exploration via Bootstrapped DQN
- (More) Efficient Reinforcement Learning via Posterior Sampling
- Generalization and Exploration via Randomized Value Functions
- A Tutorial on Thompson Sampling
- Voyager: An Open-Ended Embodied Agent with Large Language Models
- ReAct: Synergizing Reasoning and Acting in Language Models
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks