Capability Provenance in Language Models: A Case Study in Social Reasoning
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "Capability Provenance in Language Models".
Jane: As a fastidious and diligent researcher, I have meticulously analyzed these excerpts from "Capability Provenance in Language Models:
Tom: First, who's behind it and why it matters.
Title and authors: Tom: So, we’re diving into the title of this paper, "Capability Provenance in Language Models: A Case Study in Social Reasoning," and what that actually means for us. Essentially, they are looking at how to map which specific parts of the massive training data are responsible for different kinds of AI skills.
Jane: That’s right; it’s about finding the origin point for those skills, distinguishing between what makes an AI good at social reasoning versus what makes it good at STEM reasoning. It moves the focus from just looking at model output to understanding the underlying data structure itself.
Lu: What’s really interesting is how they frame capability discovery through this lens; it suggests that we can start designing training regimes that intentionally cultivate these different cognitive styles instead of just hoping they emerge naturally in a general model.
Meng: From an engineering standpoint, I see the title implying a need for better control over our data pipelines; if we know where the social reasoning skills are supported, we can focus our curation efforts precisely there.
Lalam: This concept of capability provenance is huge because it gives us a clear way to audit and control model behavior; it means we can inspect or modify specific inputs based on their known capability support.
Tom: So, if I'm hearing this correctly, the title suggests they are creating a map that tells us which data sources fuel different modes of thinking in the AI.
Jane: Precisely; instead of seeing the training corpus as one big blob, they want us to see it as a structured resource where we can allocate attention based on what capability we are trying to foster.
Lu: It’s about understanding how language itself is organized in relation to different types of reasoning, which opens up possibilities for mapping out the entire cognitive landscape of an AI.
Meng: That structural understanding is what makes the work useful for engineering because it gives us concrete targets for what we need to modify in our training pipelines.
Lalam: And this precision directly impacts how we build better, more aligned AI systems; if we can map the cultural and cognitive drivers of social reasoning, we can design models that exhibit the specific social qualities we want to see in them.
The paper's summary: Tom: So, let’s get into what they actually found in this paper on "Capability Provenance in Language Models: A Case Study in Social Reasoning." They used training-data attribution to look at which corpus regions support social reasoning versus STEM reasoning.
Jane: They compute gradient-based attribution, specifically using TrackStar via Bergson, over a working set drawn from the de-duplicated Dolma3 mix. Then, they aggregate this influence across WebOrganizer’s twenty-four-format by twenty-four-topic taxonomy, resulting in five hundred seventy-six distinct bins.
Lu: They established a two times two design to contrast domain—social versus STEM—and capability type—reasoning versus knowledge, specifically comparing SocialIQA and MMLU Social Sciences against ARC-Challenge and MMLU STEM.
Meng: They found that social reasoning shows a profile that is broader and more evenly distributed across its training data compared to the comparison tasks, while STEM/knowledge tasks concentrate their high attribution mass in documentation-like and technical regions.
Lalam: That contrast tells us something important about how different types of text fuel different AI behaviors; it suggests that social reasoning draws from a wider variety of textual material than STEM tasks do.
Tom: And they also found a key distinction: the difference between social and STEM reasoning is sharper at the level of reasoning itself than it is at the knowledge level. That’s a specific nuance we should pay attention to.
Jane: Exactly; that means when we look at their ability to reason, there's a clearer pattern in how the data supports social versus STEM tasks compared to when we just look at general knowledge recall.
Lu: They also introduced lexical profiling, which examines the six format-cluster compositions of those top-twenty high-influence bins for each benchmark. This allows them to move from abstract bins to actual linguistic insights about the content type.
Meng: That’s really useful because it gives us a semantic reason for why a certain bin is important rather than just giving us an abstract coordinate. It helps ground the mathematical analysis in real language patterns we can actually work with.
Lalam: Knowing what kind of text—like formal reference versus expository groups—is driving the influence helps us understand the communicative function underpinning the capability, which is really useful for alignment.
The paper's improvements: Tom: Moving on to how they suggest making this capability provenance mapping tool itself better, what methodological enhancements did they propose in this paper?
Jane: They suggest moving past just document-level estimates by using a structured taxonomy-based aggregation method to create those interpretable corpus regions. That’s the core improvement over just looking at individual document influence scores, making the output much more meaningful for researchers.
Lu: They also introduced lexical profiling, which examines the six format-cluster compositions of those top-twenty high-influence bins for a given benchmark. This allows them to move from abstract bins to actual linguistic insights about the content type.
Meng: From an engineering perspective, that lexical profiling is extremely useful because it gives us a semantic reason for why a certain bin is important rather than just giving us an abstract coordinate. It helps ground the abstract math in actual language patterns we can work with.
Lalam: I think the most exciting part of their suggested improvements is the incorporation of targeted machine unlearning as a causal validation mechanism. This allows them to test if high-attribution topics are actually functionally necessary for performance, which moves the analysis from mere correlation to something more functional.
Tom: So, they aren't just saying "this data is there," they are also testing if "this data is *needed* for this skill," right?
Jane: Exactly; and their findings on targeted machine unlearning showed that for SocialIQA, influence-targeted unlearning significantly degraded accuracy more than random baselines. That positive effect, measured by Cohen's d being +zero point three nine, validates the link between high attribution and functional necessity for social reasoning.
Lu: However, they also acknowledged a limitation in this validation process; their analysis showed that attribution doesn't uniformly isolate functionally necessary documents across every single benchmark. The weaker, null, and reversed effects on comparison tasks indicate that attribution often identifies correlations rather than absolute causal necessity for every document.
Meng: That limitation is important because it means we can't just blindly trust the attribution map for every single intervention; there are still many other factors at play. It’s a warning that we need to be careful when applying these findings to real-world model modification pipelines.
Lalam: That caution is exactly what we need in the development process; it means this tool helps us guide, but it doesn't give us an absolute command over every single aspect of the AI's intelligence. It sets a high bar for what we should actually aim for in terms of data curation.
Conclusion: Tom: So, to wrap up this discussion on "Capability Provenance in Language Models: A Case Study in Social Reasoning," the main implication is that we have an interpretable tool using training-data attribution to map precisely which corpus regions support social reasoning versus STEM reasoning.
Jane: It really does provide a detailed look at those different data profiles, showing that social reasoning is supported by a more diffuse profile compared to the more concentrated technical regions found in STEM tasks.
Lu: The contrast they established between the sharpness of the distinction at the reasoning level versus the knowledge level is quite telling about how capability manifests in different data types.
Meng: And that’s why their validation using targeted machine unlearning was so important; it showed that high attribution topics aren't just correlated with performance, but are functionally necessary for specific social reasoning skills in tasks like SocialIQA.
Lalam: It gives us a verifiable way to understand what’s actually working in our models and how we can use that knowledge to make AI behavior more intentional and aligned for human interaction.
Tom: Absolutely, Lalam; it moves us from just observing model output to understanding the underlying data architecture that creates those outputs. Jane, can you tell us what that means for how you think about training data selection moving forward?
Jane: Well, it means we move away from just picking a large corpus and start looking at the structural composition of that corpus, seeing which formats or topics are most influential for the specific skills we want to foster.
Lu: And I think that structured approach is what’s really exciting because it suggests a deeper level of understanding about how language itself is organized in relation to different types of reasoning. It’s like finding the underlying topography of an AI's knowledge base.
Meng: I agree with Lu; that structural understanding is what makes the work useful for engineering because it gives us concrete targets for what we need to modify in our training pipelines; it’s about precision over brute force data changes.
Lalam: And this precision directly impacts how we build better, more aligned AI systems; if we can map the cultural and cognitive drivers of social reasoning, we can design models that exhibit the specific social qualities we want to see in them.
Tom: So, to recap this paper on "Capability Provenance in Language Models: A Case Study in Social Reasoning," it shows us how training-data attribution reveals the distinct linguistic and structural supports for social versus STEM reasoning.
Jane: It really does provide a detailed look at those different data profiles, showing that social reasoning is supported by a more diffuse profile compared to the more concentrated technical regions found in STEM tasks.
Lu: The contrast they established between the sharpness of the distinction at the reasoning level versus the knowledge level is quite telling about how capability manifests in different data types.
Meng: And that's why their validation using targeted machine unlearning was so important; it showed that high attribution topics aren't just correlated with performance, but are functionally necessary for specific social reasoning skills in tasks like SocialIQA.
Lalam: It gives us a verifiable way to understand what’s actually working in our models and how we can use that knowledge to make AI behavior more intentional and aligned for human interaction.
Tom: Absolutely, Lalam; it moves us from just observing model output to understanding the underlying data architecture that creates those outputs. What a fantastic piece of work by the authors on "Capability Provenance in Language Models: A Case Study in Social Reasoning".
Jane: It’s inspiring, Tom, and I think it really solidifies the concept that we need to look beyond document-level scores when trying to understand what an AI is actually doing.
Lu: Indeed, the potential for mapping out the entire cognitive landscape through this kind of provenance tracking is huge, and I think future work could explore applying this framework across even more complex multimodal reasoning tasks.
Meng: I’m looking forward to seeing how we can integrate these attribution maps directly into our development workflows to make those targeted interventions a reality for production systems.
Lalam: I'm really hopeful that this work will help us build AI that isn't just smart, but also deeply capable of nuanced social understanding and behavior, improving the way we interact with these systems in our daily lives.
Tom: We certainly hope so; it’s a really solid piece of research that gives us a much clearer lens on how to guide AI development moving forward.
Georgia Institute of Technology, College of Computing MATS Program · EleutherAI KAIST AI Georgia Tech AI Safety Initiative
cs.CL, cs.LG
Submitted: 2026-06-17
Updated: 2026-10-05
Code: https://github.com/eilab-gt/capabilibara
Importance score: 89/100
The gist: As a fastidious and diligent researcher, I have meticulously analyzed these excerpts from "Capability Provenance in Language Models: A Case Study in Social Reasoning." The provided text details a
Key concepts
- Training-Data Attribution (TDA)
- A method used to figure out which specific regions within the pretraining text most influenced a model's ability to perform a certain task. It aggregates influence across many data points into structured units for analysis.
- Corpus Units
- The raw attribution scores are grouped into 576 distinct bins based on topic and format. These units simplify the massive dataset, allowing researchers to compare evidence for social versus STEM reasoning in a clear way.
- Targeted Machine Unlearning
- A validation technique where researchers intentionally remove specific data points identified as highly influential. For social tasks, this removal significantly harms performance, proving those specific data regions are functionally important.
Terminology
Summary
As a fastidious and diligent researcher, I have meticulously analyzed these excerpts from Capability Provenance in Language Models: A Case Study in Social Reasoning.
The provided text details a sophisticated methodology employing Training-Data Attribution (TDA) to map corpus regions to specific model capabilities, validated through targeted machine unlearning experiments.
Here is a comprehensive, detailed synthesis of the paper's core findings and methodology:
The central innovation of this work lies in using Training-Data Attribution (TDA) as an interpretable tool to discover which regions of a massive pretraining corpus support distinct model capabilities, specifically contrasting social reasoning against STEM reasoning.
-
Attribution Mechanism: The authors move beyond noisy document-level scores by aggregating influence across a structured taxonomy. They utilize gradient-based attribution methods (specifically referencing TrackStar via Bergson) computed over a working set drawn from the de-duplicated Dolma3 mix.
-
Corpus Structuring: This raw attribution data is aggregated across WebOrganizer’s 24-format times 24-topic taxonomy, resulting in 576 distinct bins. This structure transforms noisy document scores into comparable, interpretable corpus units for analysis.
-
Contrast Design: The study employs a 2 times 2 design to contrast capabilities:
-
Domain: Social vs. STEM.
-
Capability Type: Reasoning vs. Knowledge (e.g., SocialIQA/MMLU Social Sciences vs. ARC-Challenge/MMLU STEM).
The analysis reveals distinct and non-monolithic evidence for social and STEM reasoning, highlighting that different capabilities rely on qualitatively different pretraining data:
-
Social Reasoning Profile: SocialIQA exhibits a broader and more evenly distributed training-data profile compared to the comparison tasks. Its positive influence is characterized as diffuse and narrative/interpersonal, drawing strong support from both short interpersonal text and long-form documentation.
-
STEM/Knowledge Profile: In contrast, MMLU Social Sciences, ARC-Challenge, and MMLU STEM concentrate their high attribution mass in documentation-like and technical regions.
-
Sharpness of Contrast (RQ2): The distinction between social and STEM reasoning is sharper at the reasoning level than at the knowledge level. Specifically, SocialIQA successfully separates corpus regions more strongly against ARC-Challenge than MMLU Social Sciences does against MMLU STEM, although the general direction of influence is preserved.
To move from abstract bins to meaningful linguistic insights, the authors apply Lexical Profiling:
-
Format Clustering: This protocol examines the six format-cluster compositions of the top-20 high-influence bins for a given benchmark.
-
Functional Insight: The results show that high-attribution bins for social tasks often look like formal reference/documentation text, while expository groups supply a baseline contrast, suggesting different communicative functions underpin the capabilities.
The study employs targeted machine unlearning as a partial causal validation mechanism to determine if high-attribution topics are truly functionally necessary for performance:
-
Validation Success: For SocialIQA, influence-targeted unlearning significantly degrades accuracy more than within-topic random baselines (a statistically significant positive effect, e.g., Cohen's d = +0.39). This strongly validates the association between high attribution and functional necessity for social reasoning tasks.
-
Limitations in Isolation: The analysis shows that attribution does not uniformly isolate functionally necessary documents across all benchmarks. Weaker, null, and reversed effects on comparison tasks indicate that attribution identifies correlations rather than absolute causal necessity for individual documents.
Further rigorous testing was conducted to assess the reliability of the bin-level approach:
-
Method Comparison: A cross-method comparison (Table 53) showed that bin-level unlearning (using a single seed topic selection procedure) often yields a better selectivity ratio for SocialIQA than naive top-K methods, suggesting the structured corpus unit approach is advantageous.
-
Robustness to Benchmark Definition: A critical check on the ARC benchmark demonstrated its inherent robustness: performing bin-level unlearning across three different ARC definitions (Easy, Challenge, Combined) resulted in near-zero on-target accuracy degradation (gamma about 0) across all definitions.
Improvements for AI systems
Based on this scientific paper, here are the specific improvements that can be made to AI systems, categorized by the capability they enhance:
) Specific Improvements for AI Systems:
-
The model should utilize a
Capability Provenance Mapping
layer during pretraining or fine-tuning. This layer would analyze the training data distribution (topic/format bins) and assign a provenance score to specific corpus regions, explicitly linking those regions to known capabilities (e.g.,Literature × Customer Support
is highly indicative of Social Reasoning). -
Implement a structured taxonomy-based aggregation method for training-data attribution instead of relying on noisy document-level scores. This aggregates influence across predefined bins (like the 576 WebOrganizer cells) to create stable, interpretable corpus regions that are directly comparable across different tasks.
-
Integrate a
Capability Contrastive Design
framework during model evaluation to systematically test how different corpus regions support distinct capabilities (e.g., contrasting SocialIQA/ARC-Challenge against MMLU Social Sciences/STEM). This allows researchers to determine if social and STEM reasoning draw from qualitatively distinct data sources, which the paper suggests they do. -
Develop a
Targeted Machine Unlearning
module that uses high-attribution bins as forget sets for specific capabilities rather than random document removal. This validates whether high-influence regions are functionally necessary for a specific behavior (e.g., forgettingLiterature
degrades SocialIQA more than random topic forgetting). -
Incorporate
Correctness Differential Analysis
during model auditing, calculating the difference between influence on correctly answered queries versus incorrectly answered queries for different capability benchmarks. This helps distinguish between data regions that support correct reasoning versus those that are merely associated with errors (e.g., identifying Literature × Customer Support as a source of correct social reasoning but a source of incorrect factual recall). -
Develop
Probe-Based Capability Auditing
using a held-out suite of probes targeting specific areas like Theory of Mind, moral judgment, and bias. This allows for the mapping of capability provenance to complex social reasoning sub-components (e.g., identifying which corpus regions support theory of mind vs. simple factual knowledge). -
Implement
Lexical Profiling
on high-influence bins to identify the linguistic characteristics driving model behavior (e.g., dialogue-rich interpersonal text vs. long-form documentation). This provides a semantic understanding of the data's influence, moving beyond mere structural location to content type.
) What the Improved AI System Can Do:
The improved AI system can perform advanced, auditable, and controllable reasoning in the following ways:
-
Do deep audits to determine if its social reasoning capabilities are supported by a distinct corpus region (e.g., interpersonal dialogue) versus STEM-aligned regions (e.g., technical documentation).
-
Provide verifiable evidence that specific training data components are functionally necessary for achieving certain high-level reasoning skills, allowing for targeted data curation or safety interventions to shape behavior without catastrophic forgetting of general ability.
-
Be audited to see if its performance on social tasks is driven by the same corpus regions as its knowledge/STEM tasks, enabling researchers to pinpoint where capability-specific evidence resides.
-
Exhibit robust, verifiable control over memory and knowledge (unlearning), ensuring that removing specific data types does not degrade general intelligence but specifically targets the learned skill structure.
-
Be assessed on the correctness of its reasoning across different query types (correct vs. incorrect answers), allowing for differentiation between data sources that support accurate inference versus those that merely provide factual recall, leading to more nuanced performance debugging.
-
Support targeted safety and alignment by identifying specific corpus regions responsible for undesirable behaviors (e.g., bias or moral judgment) and allowing developers to focus on modifying those specific data inputs or model weights rather than applying broad, ineffective interventions.
Sources
- A Survey on Data Selection for Language Models
- Aligning AI With Shared Human Values
- Organize the Web: Constructing Domains Enhances Pre-Training Data Curation
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering