Coherence-Aware Distributional Evaluation of Open-Ended Text Generation
cs.CL, cs.AI
Submitted: 2026-09-28
Updated: 2026-09-29
Code: https://github.com/MAPS-research/CHORD
Terminology
Sources
- NeoBERT: A Next-Generation BERT
- LangFlow: Continuous Diffusion Rivals Discrete in Language Modeling
- SummEval: Re-evaluating Summarization Evaluation
- Hacking Generative Perplexity: Why Unconditional Text Evaluation Needs Distributional Metrics
- Gemma 2: Improving Open Language Models at a Practical Size
- The Llama 3 Herd of Models
- Continuous Latent Diffusion Language Model
- GANs Trained by a Two Time-Scale Update Rule Converge to a Local Nash Equilibrium
- The Curious Case of Neural Text Degeneration
- PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference
- CTRLEval: An Unsupervised Reference-Free Metric for Evaluating Controlled Text Generation
- Efficient Memory Management for Large Language Model Serving with PagedAttention
- G-Eval: NLG Evaluation using GPT-4 with Better Human Alignment
- Discrete Diffusion Modeling by Estimating the Ratios of the Data Distribution
- Pointer Sentinel Mixture Models
- GPT-4 Technical Report
- Granite Guardian
- PyTorch: An Imperative Style, High-Performance Deep Learning Library
- MAUVE: Measuring the Gap Between Neural Text and Human Text using Divergence Frontiers
- Generative Frontiers: Why Evaluation Matters for Diffusion Language Models
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering