PriceBench: A Diagnostic Benchmark for Price, Quality, and Brand Preferences in LLM Booking Agents
econ.GN, cs.AI, cs.CL, q-fin.EC
Submitted: 2026-09-25
Updated: 2026-09-25
Code: https://github.com/Pashasan/pricebench-emnlp
Terminology
Sources
- Magentic Marketplace: An Open-Source Environment for Studying Agentic Markets
- Algorithmic Collusion by Large Language Models
- EconEvals: Benchmarks and Litmus Tests for Economic Decision-Making by LLM Agents
- The Challenge of Using LLMs to Simulate Human Behavior: A Causal Inference Perspective
- Large Language Models as Simulated Economic Agents: What Can We Learn from Homo Silicus?
- LLM economicus? Mapping the Behavioral Biases of LLMs via Utility Theory
- TravelPlanner: A Benchmark for Real-World Planning with Language Agents
- $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains
- Large Language Models Are Not Robust Multiple Choice Selectors
- Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
- WebArena: A Realistic Web Environment for Building Autonomous Agents
Related papers
- How optimistic inflow forecasts distort dispatch, prices, and contracts in hydro-dominated power systems: evidence from Brazil
- The Economics of AI Inference: Inflation Dynamics, Welfare Costs, and Optimal Monetary Policy under the Inference-Cost Phillips Curve
- AI Economist Agent: An Agentic Framework for Evidence-Based Economic and Financial Analysis with RAG, Knowledge Graphs, and Large Language Models
- The Economics of Model Collapse: Equilibrium, Welfare, and Optimal Provenance Subsidies in Synthetic Data Markets
- Is Decentralized Finance Actually Decentralized? An Interdisciplinary Framework Integrating Network Theory, Agent-Based Simulation, and Longitudinal Evidence from Aave, GHO Issuance, and Cross-Chain Expansion
- Dynamic Resource Allocation with Karma: An Experimental Study