NC-Bench: An LLM Benchmark for Evaluating Conversational Competence
cs.CL, cs.AI
Submitted: 2026-01-10
Updated: 2026-09-25
Code: https://github.com/IBM/nc-bench
Terminology
Sources
- Do Large Language Models Need Intent? Revisiting Response Generation Strategies for Service Assistant
- Benchmarking Large Language Models in Retrieval-Augmented Generation
- A Complete Survey on LLM-based AI Chatbots
- Survey of the State of the Art in Natural Language Generation: Core tasks, applications and evaluation
- The Llama 3 Herd of Models
- MTRAG: A Multi-Turn Conversational Benchmark for Evaluating Retrieval-Augmented Generation Systems
- LLMs Get Lost In Multi-Turn Conversation
- LaRA: Benchmarking Retrieval-Augmented Generation and Long-Context LLMs -- No Silver Bullet for LC or RAG Routing
- DailyDialog: A Manually Labelled Multi-turn Dialogue Dataset
- Zero-shot Conversational Summarization Evaluations with small Large Language Models
- Evaluating LLM Metrics Through Real-World Capabilities
- A Survey of the Evolution of Language Model-Based Dialogue Systems: Data, Task and Models
- Qwen2.5 Technical Report
- Does Model Size Matter? A Comparison of Small and Large Language Models for Requirements Classification
- A Comprehensive Analysis of the Effectiveness of Large Language Models as Automatic Dialogue Evaluators
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering