Sell More, Play Less: Benchmarking LLM Realistic Selling Skill
cs.CL
Submitted: 2026-04-08
Updated: 2026-08-26
Terminology
Sources
- Persuade Me if You Can: A Framework for Evaluating Persuasion Effectiveness and Susceptibility Among Large Language Models
- Language Models are Few-Shot Learners
- Injecting Salesperson's Dialogue Strategies in Large Language Models with Chain-of-Thought Reasoning
- M3-Embedding: Multi-Linguality, Multi-Functionality, Multi-Granularity Text Embeddings Through Self-Knowledge Distillation
- DeepSeek-V3 Technical Report
- Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators
- ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools
- The Llama 3 Herd of Models
- Richelieu: Self-Evolving LLM-Based Agents for AI Diplomacy
- From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline
- RoBERTa: A Robustly Optimized BERT Pretraining Approach
- AgentBoard: An Analytical Evaluation Board of Multi-turn LLM Agents
- Flipping the Dialogue: Training and Evaluating User Language Models
- GPT-4o System Card
- Qwen2.5 Technical Report
- Direct Preference Optimization: Your Language Model is Secretly a Reward Model
- A Survey of Evaluation Metrics Used for NLG Systems
- Seed1.5-Thinking: Advancing Superb Reasoning Models with Reinforcement Learning
- MiMo-V2-Flash Technical Report
- Gemini: A Family of Highly Capable Multimodal Models
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering