IDEAlign: Comparing Ideas of Large Language Models to Domain Expert
cs.CL, cs.CY
Submitted: 2025-09-02
Updated: 2026-08-26
Comments: EACL 2026
Code: https://github.com/EduNLP/IDEAlign
License: http://creativecommons.org/licenses/by-nc-nd/4.0/
The gist: Large language models (LLMs) are increasingly used to produce open-ended, interpretive annotations, yet there is no validated, scalable measure of idea-level similarity to expert annotations.
Terminology
Abstract
Large language models (LLMs) are increasingly used to produce open-ended, interpretive annotations, yet there is no validated, scalable measure of idea-level similarity to expert annotations. We (i) introduce the content evaluation of LLM annotations as a core, understudied task, (ii) propose IDEAlign for capturing expert similarity judgments via pick-the-odd-one-out tasks, and (iii) benchmark various similarity methods (text embeddings, topic models, and LLM-as-a-judge) against these human ratings. Applying this approach to two real-world educational datasets (e.g., interpreting math reasoning and feedback generation), we find that most metrics fail to capture the nuanced dimensions of similarity meaningful to experts. LLM-as-a-judge performs best (11 18% improvement over other methods) but still falls short of expert alignment, making it useful as a triage tool rather than a substitute for human review. Our work demonstrates the difficulty of evaluating open-ended LLM annotations at scale, and positions IDEAlign as a reusable protocol for benchmarking on this task to help guide responsible deployment of LLMs.
Sources
- SimCSE: Simple Contrastive Learning of Sentence Embeddings
- BERTopic: Neural topic modeling with a class-based TF-IDF procedure
- Evaluating Creative Short Story Generation in Humans and Large Language Models
- Towards General Text Embeddings with Multi-stage Contrastive Learning
- SummEval: Re-evaluating Summarization Evaluation
- TruthfulQA: Measuring How Models Mimic Human Falsehoods
- All-but-the-Top: Simple and Effective Postprocessing for Word Representations
- Text and Code Embeddings by Contrastive Pre-Training
- Large Dual Encoders Are Generalizable Retrievers
- GPT-4 Technical Report
- MAUVE: Measuring the Gap Between Neural Text and Human Text using Divergence Frontiers
- LearnLM: Improving Gemini for Learning
- Text Embeddings by Weakly-Supervised Contrastive Pre-training
- Bridging the Novice-Expert Gap via Models of Decision-Making: A Case Study on Remediating Math Mistakes
- Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks
- MiniLM: Deep Self-Attention Distillation for Task-Agnostic Compression of Pre-Trained Transformers
- Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge
- Smarter, Better, Faster, Longer: A Modern Bidirectional Encoder for Fast, Memory Efficient, and Long Context Finetuning and Inference
- Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates
- BERTScore: Evaluating Text Generation with BERT
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering