TripScore: Aligning LLMs for Real-World Travel Planning via Expert-Calibrated Reward

arXiv:2510.09011 · cs.AI, cs.CL · Submitted 2025-10-10 · Read on arXiv

cs.AI, cs.CL

Submitted: 2025-10-10

Updated: 2026-09-17

Comments: EMNLP2026 Industry track

License: http://creativecommons.org/licenses/by/4.0/

The gist: In our deployed travel-planning service, most users give minimal inputs or free-form requests rather than the structured constraint checklists assumed by existing benchmarks.

Terminology

Abstract

In our deployed travel-planning service, most users give minimal inputs or free-form requests rather than the structured constraint checklists assumed by existing benchmarks. We therefore present TripScore, a behavior-grounded benchmark and evaluation framework built from real user logs and calibrated against 1,468 pairwise judgments by 203 travel experts. TripScore couples a hierarchical feasibility gate (format and commonsense) with a unified, point-wise reward that aggregates soft quality and preference fulfillment. Using TripScore as both evaluator and reward signal, we benchmark direct prompting, test-time compute, neuro-symbolic solvers, code agents, and fine-tuning. We find that reinforcement learning fine-tuning (e.g., GRPO) provides consistent gains over other approaches under the same base model and practical latency.

Sources

Related papers