Test-Time Policy Adaptation for Enhanced Multi-Turn Interactions with LLMs
cs.CL
Submitted: 2025-09-27
Updated: 2026-09-27
Terminology
Sources
- A Survey on Evaluation of Large Language Models
- Evaluating Large Language Models Trained on Code
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- Retrieval-Augmented Generation for Large Language Models: A Survey
- SLOT: Sample-specific Language Model Optimization at Test-time
- Rewarding Chatbots for Real-World Engagement with Millions of Users
- Adam: A Method for Stochastic Optimization
- Training Language Models to Self-Correct via Reinforcement Learning
- LLMs Get Lost In Multi-Turn Conversation
- Beyond Single-Turn: A Survey on Multi-Turn Interactions with Large Language Models
- Let's Verify Step by Step
- Qwen2.5 Technical Report
- WildFeedback: Aligning LLMs With In-situ User Interactions And Feedback
- Linguistic Generalizability of Test-Time Scaling in Mathematical Reasoning
- SuperGPQA: Scaling LLM Evaluation across 285 Graduate Disciplines
- PAFT: Prompt-Agnostic Fine-Tuning
- Qwen3 Technical Report
- A Survey on Multi-Turn Interaction Capabilities of Large Language Models
- A Survey on Test-Time Scaling in Large Language Models: What, How, Where, and How Well?
- SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering