Where Animacy Lives in Large Language Models: Tracing the Circuits of the Animacy Concept
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Where Animacy Lives in Large Language Models: Tracing the Circuits of the Animacy Concept".
Jane: The paper was written by Samuele Punzo, Giovanni Cinà and Sandro Pezzelle from University of Amsterdam and Amsterdam University Medical Center.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title: Tom: We're looking at 'Where Animacy Lives in Large Language Models: Tracing the Circuits of the Animacy Concept' by Samuele Punzo, Giovanni Cinà, and Sandro Pezzelle.
Jane: The title alone makes it sound like they're searching for a soul in the machine.
Tom: They're looking for the mathematical fingerprint of life and agency.
Lu: It's a beautiful question to ask in this era of rapid development. If these models can distinguish between a human and a rock, they're closer to understanding our world than we thought. This could lead to AI that truly understands the social context of our lives.
Meng: I'm thinking about the safety side of that, Lu. If we know exactly where the model decides something is a person, we can prevent it from treating people like inanimate objects in dangerous scenarios. It's about building more reliable, predictable systems.
Lalam: That reliability builds a foundation for cultural trust. When an AI understands the distinction between a living being and a tool, it can interact with our traditions and social norms more gracefully. It moves us from mere calculation to a shared sense of meaning.
Jane: So we're moving from just seeing words to seeing the intent behind them.
Tom: Exactly, Jane, and that's what we'll explore when we look at how they actually found these circuits.
Summary: Tom: To get to those circuits, the researchers used a really clever setup involving twenty thousand minimal pairs to test models like GPT-two Llama three point two 3B, Gemma three 4B, and Qwen three 4B.
Jane: Minimal pairs are just sentences that are almost identical, except for one tiny thing that changes the whole meaning.
Tom: Right, like saying "The victim was rescued by the" versus "The victim was crushed by the."
Jane: One sentence expects a person to follow, while the other expects a natural force or an object.
Lu: I love how they used GPT five point four to help generate these semantic frames to ensure they were plausible. It's like using a master architect to design the testing ground for a new building. They're creating a perfect environment to see if the model's internal logic holds up.
Meng: The technical part that really caught my eye was the use of Edge Attribution Patching with Integrated Gradients. They aren't just looking at which neurons fire, but specifically which connections, or edges, are moving the needle. It's a much more precise way to map the actual information flow through the network.
Lalam: It's a very disciplined way to study the emergence of meaning. By focusing on these specific edges, they're showing that animacy isn't just a random byproduct of training, but a structured part of the model's internal world.
Meng: It's definitely more structured than we often give them credit for.
Tom: And that structure is exactly what they managed to uncover in their results.
Findings: Jane: They actually found a circuit, but it wasn't as simple or localized as they might have expected.
Tom: Instead of one single "animacy button," it's a distributed network of components.
Jane: They found that the very first MLP layer, which they call MLP0, acts like a sorting office.
Tom: It seems to prepare the verb information early on, and then later MLPs take that info and align it with the final decision.
Jane: And the attention heads aren't even deciding if something is alive; they're mostly just doing the heavy lifting of moving information around, like a relay race from the verb to the passive marker.
Lu: This suggests that intelligence in these models isn't a single switch, but a fluid, interconnected web. It's much more organic than a simple flowchart. If the animacy concept is spread out like this, it means it's woven into the very fabric of how the model processes language.
Meng: That actually makes my job a bit harder, Lu. If the mechanism is distributed and task-specific, it means we can't just go in and "fix" one spot to change how a model perceives agency. The paper shows that if you change both the data and the targets, the circuit can actually break, even if the model still gets the answer right.
Lalam: That nuance is so important for our understanding of digital cognition. It shows that these models have a "graded" understanding, where concepts aren't just black and white. They're learning the subtle, messy boundaries of the real world, just like we do.
Tom: It's a much more complex picture than anyone realized.
Conclusion: Jane: We've seen that while there isn't one universal "animacy circuit," there is a measurable, causal mechanism that models use to navigate these concepts.
Tom: It's a huge step forward for mechanistic interpretability.
Lu: I think this opens the door to teaching models even deeper layers of human experience. If we can map these circuits, we can guide them toward a more profound understanding of life and agency.
Meng: From a practical standpoint, this gives us a roadmap for more robust testing. We can stop looking for single neurons and start looking at how these functional subgraphs evolve as models get larger.
Lalam: Ultimately, this work helps us bridge the gap between mathematical weights and human values. It brings us closer to technology that feels less like a machine and more like a participant in our shared culture.
Jane: It really has been a fascinating look into the hidden structures of AI.
Tom: We're all walking away with a lot to think about. We'll be back next time with another deep dive.
Jane: Goodbye for now!
University of Amsterdam · Amsterdam University Medical Center
cs.CL
Submitted: 2026-07-23
Updated: 2026-09-13
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 84/100
The gist: This paper investigates whether the animacy-sensitive behavior of Large Language Models (LLMs) can be traced to a "localized set of causally relevant components and connections." By performing
Key concepts
- Minimal pairs
- Sentences that are nearly identical except for one small change that alters the meaning, such as switching between an expected person or an object following a verb. This method helps researchers test how models distinguish specific semantic concepts like animacy.
- Edge Attribution Patching with Integrated Gradients
- A precise technical method used to map information flow through a neural network. Instead of just identifying which neurons fire, this technique identifies the specific connections, or edges, that drive the model's decisions and move information through the system.
- Mechanistic interpretability
- The study of how models process language by uncovering their internal logic and hidden structures. This research helps bridge the gap between mathematical weights and human concepts, moving beyond seeing words to understanding the intent behind them.
Terminology
Summary
This paper investigates whether the animacy-sensitive behavior of Large Language Models (LLMs) can be traced to a localized set of causally relevant components and connections.
By performing circuit discovery on four open-weight models, the researchers aim to determine if a dedicated animacy circuit
exists, which is essential for understanding how models capture complex semantic concepts that are not signaled by visible surface features.
Methodology and Dataset Construction
The researchers constructed a controlled dataset of minimal pairs
designed to isolate animacy-sensitive continuation preferences. Using the template The [patient] was [verb p.p.] by the,
they generated 20,000 minimal pairs by pairing semantic frames with animate and inanimate verbs. To ensure semantic plausibility, they used a zero-shot prompt with Llama 3.3 70B to score prefixes, retaining only those with a plausibility score above 0.90. The task is evaluated using the average logit difference
(avgLD) between two broad target sets:
-
Animate targets, restricted to unambiguous
noun.person
words such as officer or detective. -
Inanimate targets, belonging to categories such as
noun.event,
noun.phenomenon,
ornoun.artifact.
Circuit discovery was performed via Edge Attribution Patching with Integrated Gradients (EAP-IG)
on GPT-2 Small, Llama 3.2 3B, Gemma 3 4B, and Qwen 3 4B.
Circuit Discovery and Robustness
The study confirms that a causal mechanism responsible for handling animacy does exist, identifying circuits that are both sufficient and necessary to recover the behavior.
To ensure the reliability of these findings, the authors performed several validation checks:
-
Random-edge controls: Circuits composed of randomly sampled edges failed to restore meaningful task performance.
-
Stability: Multiple discovery runs showed
very high overlap
between discovered circuits, with IoU values between 0.8–1.0. -
Necessity: Ablating the top-20 highly ranked edges caused faithfulness to drop below 0.1, demonstrating the existence of a
compact necessary core.
However, the researchers note that these circuits are less localized compared to other known ones
and are more distributed
than circuits found for phenomena like Indirect Object Identification (IoI).
Functional Role of Components
Analysis reveals that the animacy circuit is not concentrated in a single localized region. Instead, the most important components are primarily MLPs, spanning from the first layer toward middle/late layers, while attention heads carry a much smaller share of importance.
The functional roles of these components appear to follow a specific pattern:
-
Early components, such as MLP0, participate by
routing or preparing information
rather than directly writing the decision to the logits. -
The verb-level contrast is available in the earliest residual representations, but becomes
logit-aligned only after later transformations.
-
Attention heads mostly support
positional routing,
such as pathways from the verb to the passive marker and the subsequent determiner.
Generalization and Transferability
The researchers conclude that the discovered circuits should not be interpreted as universal animacy circuits,
as their transfer to related settings is uneven across datasets and models.
The study tested three transfer setups to evaluate this:
-
Setup 1 (BLiMP prefixes + original targets): The circuits generalized
extremely well
in most models. -
Setup 2 (Original prefixes + named entities): The circuits generalized well, though faithfulness was
substantially lower.
-
Setup 3 (BLiMP prefixes + named entities): The circuits
do not generalize,
and the mechanism breaks.
This suggests that while the models possess a robust part of the animacy computation,
the circuit remains partly constrained by the specific task and data.
Improvements for AI systems
1. Circuit-Based Semantic Steering (Safety & Alignment)
-
Improvement: Implement surgical intervention protocols using the identified
necessary core
of animacy circuits (specifically targeting the high-attribution edges in early MLPs like MLP0 and subsequent logit-aligning MLPs). Instead of broad, unpredictable prompt engineering or RLHF, use Edge Attribution Patching (EAP) to boost or ablate specific edge activations during inference. -
Improved Capability: The AI system can precisely control agentic behavior. It can be programmed to strictly avoid
hallucinated agency
(assigning intentionality to inanimate objects in safety-critical instructions) or, conversely, ensure highly consistent persona adherence in agentic workflows by reinforcing the specific pathways that drive animate-oriented logit preferences.
2. Semantic-Structure Contrastive Fine-Tuning (Robustness & Generalization)
-
Improvement: Transition from standard next-token prediction training to a
Minimal-Pair Contrastive Objective.
Utilizing the paper’s methodology, training data should be structured intoclean
(animate-favoring) andcorrupt
(inanimate-favoring) minimal pairs. The loss function should specifically penalize the model if the internal activations of the identified animacy circuit do not show a significant logit-difference margin between these pairs. -
Improved Capability: The AI system will exhibit robust semantic generalization. It will move beyond
task-specific
pattern matching (which fails when both prefix and target change) toward a generalized concept of animacy, allowing it to correctly interpret agent/patient roles across diverse syntactic structures, such as active voice, passive voice, and complex relative clauses, without distributional drift.
3. Mechanistic Unit Testing in CI/CD Pipelines (Quality Assurance)
-
Improvement: Integrate an automated
Mechanistic Interpretability Suite
into the model deployment pipeline. This suite would use the paper’s minimal-pair dataset construction and average logit-difference (AVG LD) metric to performsemantic regression testing.
It would specifically monitor whether fine-tuning or quantization has degraded the faithfulness or accuracy of the identified animacy circuits. -
Improved Capability: The AI development lifecycle will gain a high-fidelity diagnostic tool. Engineers can detect
semantic concept drift
immediately after a model update, ensuring that optimizations (like pruning or quantization) do not inadvertently destroy the model's ability to perform complex reasoning regarding agency, causality, or intentionality.
4. Semantic-Importance Pruning & Distillation (Efficiency)
-
Improvement: Apply
Circuit-Aware Model Compression.
During distillation or pruning, use EAP-IG scores to identify and protect thenecessary core
edges and components (the small percentage of edges that provide 85%+ faithfulness). Non-essential edges—those not contributing to these vital semantic circuits—can be aggressively pruned. -
Improved Capability: The system can produce highly efficient, small-parameter models (SLMs) that retain the sophisticated semantic reasoning capabilities of much larger models. This allows for high-reasoning, agentic AI to run on edge devices without sacrificing the ability to distinguish between intentional agents and inanimate causes.
Abstract
Distinguishing animate from inanimate concepts in written language requires more than shallow text processing, as it involves recognizing complex selectional constraints and contextual cues, such as verb-argument interactions. Yet, current large language models (LLMs) appear to be capable of doing it. We investigate whether this animacy-sensitive behavior of LLMs can be traced to a localized set of causally relevant components and connections. To do so, we construct a controlled dataset of minimal pairs and perform circuit discovery on four open-weight models. Through in-depth experiments and ablations, we show that a causal mechanism responsible for handling animacy in these models does exist, thus discovering an animacy circuit. At the same time, this circuit appears to be less localized compared to other known ones and generalizes only partially across models and animacy tasks, confirming the distributed, context-dependent, and somewhat graded nature of the animacy concept.
Sources
- Towards Automated Circuit Discovery for Mechanistic Interpretability
- The Llama 3 Herd of Models
- Have Faith in Faithfulness: Going Beyond Circuit Overlap When Finding Model Mechanisms
- How to use and interpret activation patching
- The Hydra Effect: Emergent Self-repair in Language Model Computations
- Locating and Editing Factual Associations in GPT
- Missed Causes and Ambiguous Effects: Counterfactuals Pose Challenges for Interpreting Neural Networks
- Gemma 3 Technical Report
- Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small
- Qwen3 Technical Report
- Towards Best Practices of Activation Patching in Language Models: Metrics and Methods
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering