Do LLM Recommenders Know When They're Hallucinating? Auditing Confidence Calibration in Catalog Faithfulness
Srijith Ravikumar
cs.IR, cs.CL, cs.LG
Submitted: 2026-08-22
Updated: 2026-08-25
Code: https://github.com/rsrijith/cikm26-catalog-faithfulness
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- Mitigating LLM Hallucinations via Conformal Abstention
- A Bi-Step Grounding Paradigm for Large Language Models in Recommendation Systems
- Limits to scalable evaluation at the frontier: LLM as Judge won't beat twice the data
- Bridging Language and Items for Retrieval and Recommendation: Benchmarking LLMs as Semantic Encoders
- Evaluation on Entity Matching in Recommender Systems
- Language Models (Mostly) Know What They Know
- Why Language Models Hallucinate
- Calibrated Language Models Must Hallucinate
- AbstentionBench: Reasoning LLMs Fail on Unanswerable Questions
- Uncertainty Quantification and Decomposition for LLM-based Recommendation
- Eliminating Out-of-Domain Recommendations in LLM-based Recommender Systems: A Unified View
- Large language models can accurately predict searcher preferences
- Can LLMs Express Their Uncertainty? An Empirical Evaluation of Confidence Elicitation in LLMs
- R-Tuning: Instructing Large Language Models to Say `I Don't Know'
- Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Related papers
- The Price of Isolation: Estimating the Ecosystem Cost of Symmetric Two-Sided A/B Testing
- SCAR: Semantic Continuity-Aware Retrieval for Efficient Context Expansion in RAG
- MixLoRA-DSI: Dynamically Expandable Mixture-of-LoRA Experts for Rehearsal-Free Generative Retrieval over Dynamic Corpora
- RRCM: Ranking-Driven Retrieval over Collaborative and Meta Memories for LLM Recommendation
- Right Family, Wrong Skill: Evaluating Risk Exposure in Agent Skill Retrieval
- UltRAG: a Universal Simple Scalable Recipe for Knowledge Graph RAG