Learning to Sell: Reinforcement Learning for Strategic Large Language Model Agents in Multi-Product Markets
cs.AI
Submitted: 2026-09-27
Updated: 2026-09-27
Terminology
Sources
- Training Language Models for Bilateral Trade with Private Information
- How Well Can LLMs Negotiate? NegotiationArena Platform and Analysis
- MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention
- Robust Data-Collection Policy Learning for Low-Variance Online Policy Evaluation
- Fast Rates in $\alpha$-Potential Games via Regularized Mirror Descent
- Pessimism-Free Offline Learning in General-Sum Games via KL Regularization
- Training Verifiers to Solve Math Word Problems
- Evaluating Language Model Agency through Negotiations
- Improving Language Model Negotiation with Self-Play and In-Context Learning from AI Feedback
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- Kimi K2: Open Agentic Intelligence
- Strategic Bargaining in Multi-Buyer Markets: Reinforcement Learning from Verifiable Rewards for LLM Negotiations
- Instructing LLMs to Negotiate using Reinforcement Learning with Verifiable Rewards
- OR-Transformer: Scaling Real-Time Decision-Making to 1,000 Items
- Let's reward step by step: Step-Level reward model as the Navigators for Reasoning
- Qwen3 Technical Report
- Too Nice to Tell the Truth: Quantifying Agreeableness-Driven Sycophancy in Role-Playing Language Models
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- OpenAI GPT-5 System Card
- ReAct: Synergizing Reasoning and Acting in Language Models
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection