Oolong: Evaluating Long Context Reasoning and Aggregation Capabilities
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.
Jane: Today's paper: "Oolong: Evaluating Long Context Reasoning and Aggregation Capabilities".
Tom: The gist: OOLONG introduces a benchmark for long-context reasoning tasks that require analyzing individual chunks of text at an atomic level and aggregating these analyses to answer distributional questions,
Jane: First, who's behind it and why it matters.
Paper summary: Tom: We've been looking at the Oolong: Evaluating Long Context Reasoning and Aggregation Capabilities paper, which sets up this two-part benchmark of synthetic and real tasks #pg7.
Jane: The authors are pointing out that this testing method is valuable because it forces models to perform a multi-step information aggregation process where they have to analyze text at an atomic level before answering distributional questions #pg2.
Lu: The big implication is that simply increasing context length doesn't automatically mean models are better at reasoning over it; the difficulty comes from the aggregation part, not just fitting the data in #pg10.
Meng: It means for us building systems, we need to focus on creating better mechanisms for stitching together those individual pieces of reasoning correctly rather than just feeding them more tokens #pg7.
Lalam: Oolong provides a concrete way to probe those limits and see exactly where models fall short when context gets really long, showing that there's still a long way to go in designing robust long context aggregation capabilities for LLMs #pg10.
Tom: So, in simple terms, this paper is about proving that we need better ways to aggregate information over massive amounts of text if we want models to truly handle those huge context windows effectively #pg7.
Conclusion: Tom: So we've seen how Oolong sets up this whole test for models dealing with long contexts #pg7
Jane: Right, it’s about testing if these large context windows actually mean something or just more space to hold junk #pg5
Lu: The point is that these models have to look at every little chunk of text separately and then figure out how all those tiny pieces fit together for a final answer #pg5
Meng: From an engineering view, it seems the real struggle isn't just fitting the tokens, but actually stitching the correct information from those chunks together #pg7
Lalam: Exactly, it’s about that identification and aggregation of information that’s really the bottleneck here #pg5
Tom: So what does this mean for us right now when we're trying to build these next generation AI systems?
Jane: It means we need to stop just feeding them bigger text and start teaching them how to reason over those massive inputs, step by step #pg4
Lu: It forces us to design architectures that can handle this multi-step problem where you classify things piece by piece and then pool those results for the final answer #pg5
Meng: Practically speaking, it shows that we’re not just looking at the raw context length anymore; we need better internal logic for how to process it #pg7
Lalam: It pushes us to build systems that can handle these complex reasoning tasks over huge amounts of text reliably #pg7
Tom: It sounds like Oolong is really showing us where the current AI is still weak when it comes to deep, long-context understanding #pg5
Jane: And the results show a real drop in performance as the context windows get even longer across both the synthetic and real tests #pg7
Lu: It’s a clear signal that we still have a long way to go in designing robust long context aggregation capabilities for LLMs #pg7
Meng: We gotta keep pushing on those methods because this kind of structured challenge is what tells us if our scaling efforts are actually working or just masking underlying weaknesses #pg7
Amanda Bertsch Adithya Pratapa Teruko Mitamura Graham Neubig Matthew R. Gormley
Carnegie Mellon University
cs.CL, cs.AI
Submitted: 2025-11-04
Updated: 2026-10-05
Importance score: 83/100
The gist: The gist: OOLONG introduces a benchmark for long-context reasoning tasks that require analyzing individual chunks of text at an atomic level and aggregating these analyses to answer distributional
Key concepts
- OOLONG Benchmark Structure
- This benchmark is designed to test long-context reasoning by requiring models to perform multi-step tasks: identify relevant text parts, perform a subtask on each part (like classification), and then pool all those results for a final answer. It is split into synthetic and real data sets.
- OOLONG-synth Design
- This synthetic set creates challenging questions by sampling label classes, dates, and user IDs from existing classification datasets. Questions are categorized as counting statistics about labels, cross-referencing user IDs, or analyzing temporal changes in distributions.
- OOLONG-real Data
- This real data set uses transcripts from Dungeons & Dragons role-playing game shows. These transcripts cover long conversations and can be unscripted. Context windows are very large, ranging from one to 24 episode transcripts, testing aggregation over massive amounts of real dialogue.
Terminology
Summary
The gist: OOLONG introduces a benchmark for long-context reasoning tasks that require analyzing individual chunks of text at an atomic level and aggregating these analyses to answer distributional questions, which is crucial as model context lengths continue to grow.
OOLONG Benchmark Structure
OOLONG is separated into two task sets: OOLONG-synth, which consists of naturalistic synthetic tasks where components can be easily ablated, and OOLONG-real, a downstream setting requiring reasoning over real-world conversational data<ref:2511.02817#pg4>. OOLONG requires models to perform reason over large quantities of examples, to perform both classification and counting in-context, and to reason over temporal and user relations
<ref:2511.02817#pg4>. The benchmark demands that models reason over large quantities of text
through a multi-step problem where they must identify parts of the input relevant to the question, perform a subtask at each relevant part (e.g. a simple classification) and then pool the results from all the subtasks to generate the final answer
<ref:2511.02817#pg5>.
OOLONG-synth Design
OOLONG-synth is constructed by constructing challenging corpus-level questions over existing in-context learning (ICL) datasets
<ref:2511.02817#pg4>. The data collection involves 10 common text classification datasets, and these are split into validation and test tasks<ref:2511.02817#pg4>. Context window construction involves sampling a distribution over label classes so the model cannot use any information about the true distribution over labels, as well as sampling dates and user IDs<ref:2511.02817#pg5>. Questions are constructed in three types: counting questions concern simple statistical properties of the label distribution,
user information questions require additional cross-reference with the user ID field,
and timeline questions ask about changes in distribution before or after a date, between years, or between months across years
<ref:2511.02817#pg5>.
OOLONG-real Data
OOLONG-real poses challenging information aggregation questions over real data, specifically transcripts from Dungeons & Dragons role-playing game shows<ref:2511.02817#pg4>. These transcripts involve several levels of conversation, from out-of-character chitchat to rules discussion to incharacter actions and speech,
and they can be unscripted, involving tangents or side channels<ref:2511.02817#pg4>. The dataset is compiled from the Critical Role Dungeons and Dragons Dataset (CRD3), using episode transcripts with player names labeled<ref:2511.02817#pg6>. Context windows for OOLONG-real range from one to 24 episode transcripts,
covering input lengths of 55K to 1.3M tokens
<ref:2511.02817#pg7>.
Evaluation and Results
The evaluation involves a baseline, parsing answers using specified output formats, and scoring based on exact match for certain answer types or a score of score(ˆy) = 0.75y−yˆ for numerical answers
<ref:2511.02817#pg5>. Frontier models struggle with the task, as GPT-5, Claude-Sonnet-4, and Gemini-2.5-Pro all achieving less than 50% accuracy on both splits at 128K
<ref:2511.02817#pg4>. The study found that identification and aggregation of information is the bottleneck, not labeling
when ablating settings in OOLONG-synth<ref:2511.02817#pg5>. Furthermore, temporal questions are consistently the most challenging for models on OOLONG-synth<ref:2511.02817#pg5>. The results show a significant drop in performance at higher context windows
across both splits<ref:2511.02817#pg7>.
Model Behavior Analysis
The analysis of model behavior reveals that for OOLONG-synth, the apparent strategy of labeling each example before deciding which are relevant results in running out of context tokens
<ref:2511.02817#pg5>. In contrast, traces for OOLONG-real do not show these pathologies<ref:2511.02817#pg7>. DeepSeek R1 shows a discrepancy, performing below the random baseline on OOLONG-synth because the model’s apparent strategy of labeling each example before deciding which are relevant results in running out of context tokens
<ref:2511.02817#pg7>. The paper concludes that OOLONG is a usefully challenging evaluation of long-context reasoning abilities
<ref:2511.02817#pg5>.
Future Directions and Limitations
The authors argue that posing questions over a large block of context represents a realistic use scenario
<ref:2511.02817#pg5>. They also note that the performance differences between these models is not due to differing ability to perform the classification task
when providing labels in-context<ref:2511.02817#pg10>. The work suggests that there is still a long way to go in designing robust long-context aggregation capabilities for LLMs
<ref:2511.02817#pg7>. The study's findings are supported by results from prior work on related benchmarks, such as RULER and HELMET<ref:2511.02817#pg10>. The use of D&D data is noted as a novel approach, as we are the first to use fan annotations of gold labels and to consider the generation of these statistics as a task in its own right
<ref:2511.02817#pg10>. The work acknowledges that some of these incomplete traces do contain an answer
in a small number of cases<ref:2511.02817#pg10>.
Conclusion
OOLONG introduces a challenging long-context information aggregation benchmark in two parts<ref:2511.02817#pg7>. OOLONGsynth uses synthetic aggregation tasks over ICL data to enable finer-grained control of the benchmark settings, while OOLONG-real poses questions over real long-context conversational data and human-annotated labels<ref:2511.02817#pg7>. On both splits, models struggle, with performance dropping with increasing context length even when controlling for the potential compounding of mislabeling errors<ref:2511.02817#pg7>. We see substantial headroom between strong open weights models and API-based models on this task, particularly OOLONG-synth<ref:2511.02817#pg7>.
Improvements for AI systems
-
Improved long-context reasoning capability through multi-step information aggregation: The OOLONG benchmark requires models to
reason over large quantities of examples, to perform both classification and counting in-context, and to reason over temporal and user relations,
enabling models toperform multi-hop reasoning over long inputs
in a single pass. -
Enhanced real-world conversational data reasoning: By including OOLONG-real, the system can reason over
real world conversational data,
specifically "transcripts from live-action Dungeons & Dragons shows," allowing it to handle complex, unscripted context where information is not easily decomposable. -
Improved robustness against context length limitations: The research demonstrates that even when controlling for mislabeling errors, models struggle with aggregation as input grows, suggesting the system can be trained or prompted to employ strategies that mitigate token exhaustion, such as
identifying and aggregation of information is the bottleneck.
Sources
- In-Context Learning with Long-Context Models: An In-Depth Exploration
- SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
- Michelangelo: Long Context Evaluations Beyond Haystacks via Latent Structure Queries
- HELMET: How to Evaluate Long-Context Language Models Effectively and Thoroughly
- Many-Shot In-Context Learning
- SQUINKY! A Corpus of Sentence-level Formality, Informativeness, and Implicature
- MetaICL: Learning to Learn In Context
- Detecting Label Errors by using Pre-Trained Language Models
- Conversational User-AI Intervention: A Study on Prompt Rewriting for Improved LLM Response Generation
- Exploring Prompt Engineering Practices in the Enterprise
- Too Long, Didn't Model: Decomposing LLM Long-Context Understanding With Novels
- A Taxonomy for Data Contamination in Large Language Models
- ZeroSCROLLS: A Zero-Shot Benchmark for Long Text Understanding
- GSM-Infinite: How Do Your LLMs Behave over Infinitely Increasing Context Length and Reasoning Complexity?
- I Cast Detect Thoughts: Learning to Converse and Guide with Intents and Theory-of-Mind in Dungeons and Dragons
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering