Where Animacy Lives in Large Language Models: Tracing the Circuits of the Animacy Concept

arXiv:2607.20995 · cs.CL · Submitted 2026-07-23 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Where Animacy Lives in Large Language Models: Tracing the Circuits of the Animacy Concept".

Jane: The paper was written by Samuele Punzo, Giovanni Cinà and Sandro Pezzelle from University of Amsterdam and Amsterdam University Medical Center.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title: Tom: We're looking at 'Where Animacy Lives in Large Language Models: Tracing the Circuits of the Animacy Concept' by Samuele Punzo, Giovanni Cinà, and Sandro Pezzelle.

Jane: The title alone makes it sound like they're searching for a soul in the machine.

Tom: They're looking for the mathematical fingerprint of life and agency.

Lu: It's a beautiful question to ask in this era of rapid development. If these models can distinguish between a human and a rock, they're closer to understanding our world than we thought. This could lead to AI that truly understands the social context of our lives.

Meng: I'm thinking about the safety side of that, Lu. If we know exactly where the model decides something is a person, we can prevent it from treating people like inanimate objects in dangerous scenarios. It's about building more reliable, predictable systems.

Lalam: That reliability builds a foundation for cultural trust. When an AI understands the distinction between a living being and a tool, it can interact with our traditions and social norms more gracefully. It moves us from mere calculation to a shared sense of meaning.

Jane: So we're moving from just seeing words to seeing the intent behind them.

Tom: Exactly, Jane, and that's what we'll explore when we look at how they actually found these circuits.

Summary: Tom: To get to those circuits, the researchers used a really clever setup involving twenty thousand minimal pairs to test models like GPT-two Llama three point two 3B, Gemma three 4B, and Qwen three 4B.

Jane: Minimal pairs are just sentences that are almost identical, except for one tiny thing that changes the whole meaning.

Tom: Right, like saying "The victim was rescued by the" versus "The victim was crushed by the."

Jane: One sentence expects a person to follow, while the other expects a natural force or an object.

Lu: I love how they used GPT five point four to help generate these semantic frames to ensure they were plausible. It's like using a master architect to design the testing ground for a new building. They're creating a perfect environment to see if the model's internal logic holds up.

Meng: The technical part that really caught my eye was the use of Edge Attribution Patching with Integrated Gradients. They aren't just looking at which neurons fire, but specifically which connections, or edges, are moving the needle. It's a much more precise way to map the actual information flow through the network.

Lalam: It's a very disciplined way to study the emergence of meaning. By focusing on these specific edges, they're showing that animacy isn't just a random byproduct of training, but a structured part of the model's internal world.

Meng: It's definitely more structured than we often give them credit for.

Tom: And that structure is exactly what they managed to uncover in their results.

Findings: Jane: They actually found a circuit, but it wasn't as simple or localized as they might have expected.

Tom: Instead of one single "animacy button," it's a distributed network of components.

Jane: They found that the very first MLP layer, which they call MLP0, acts like a sorting office.

Tom: It seems to prepare the verb information early on, and then later MLPs take that info and align it with the final decision.

Jane: And the attention heads aren't even deciding if something is alive; they're mostly just doing the heavy lifting of moving information around, like a relay race from the verb to the passive marker.

Lu: This suggests that intelligence in these models isn't a single switch, but a fluid, interconnected web. It's much more organic than a simple flowchart. If the animacy concept is spread out like this, it means it's woven into the very fabric of how the model processes language.

Meng: That actually makes my job a bit harder, Lu. If the mechanism is distributed and task-specific, it means we can't just go in and "fix" one spot to change how a model perceives agency. The paper shows that if you change both the data and the targets, the circuit can actually break, even if the model still gets the answer right.

Lalam: That nuance is so important for our understanding of digital cognition. It shows that these models have a "graded" understanding, where concepts aren't just black and white. They're learning the subtle, messy boundaries of the real world, just like we do.

Tom: It's a much more complex picture than anyone realized.

Conclusion: Jane: We've seen that while there isn't one universal "animacy circuit," there is a measurable, causal mechanism that models use to navigate these concepts.

Tom: It's a huge step forward for mechanistic interpretability.

Lu: I think this opens the door to teaching models even deeper layers of human experience. If we can map these circuits, we can guide them toward a more profound understanding of life and agency.

Meng: From a practical standpoint, this gives us a roadmap for more robust testing. We can stop looking for single neurons and start looking at how these functional subgraphs evolve as models get larger.

Lalam: Ultimately, this work helps us bridge the gap between mathematical weights and human values. It brings us closer to technology that feels less like a machine and more like a participant in our shared culture.

Jane: It really has been a fascinating look into the hidden structures of AI.

Tom: We're all walking away with a lot to think about. We'll be back next time with another deep dive.

Jane: Goodbye for now!

University of Amsterdam · Amsterdam University Medical Center

cs.CL

Submitted: 2026-07-23

Updated: 2026-09-13

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 84/100

The gist: This paper investigates whether the animacy-sensitive behavior of Large Language Models (LLMs) can be traced to a "localized set of causally relevant components and connections." By performing

Key concepts

Minimal pairs
Sentences that are nearly identical except for one small change that alters the meaning, such as switching between an expected person or an object following a verb. This method helps researchers test how models distinguish specific semantic concepts like animacy.
Edge Attribution Patching with Integrated Gradients
A precise technical method used to map information flow through a neural network. Instead of just identifying which neurons fire, this technique identifies the specific connections, or edges, that drive the model's decisions and move information through the system.
Mechanistic interpretability
The study of how models process language by uncovering their internal logic and hidden structures. This research helps bridge the gap between mathematical weights and human concepts, moving beyond seeing words to understanding the intent behind them.

Terminology

Summary

This paper investigates whether the animacy-sensitive behavior of Large Language Models (LLMs) can be traced to a localized set of causally relevant components and connections. By performing circuit discovery on four open-weight models, the researchers aim to determine if a dedicated animacy circuit exists, which is essential for understanding how models capture complex semantic concepts that are not signaled by visible surface features.

Methodology and Dataset Construction

The researchers constructed a controlled dataset of minimal pairs designed to isolate animacy-sensitive continuation preferences. Using the template The [patient] was [verb p.p.] by the, they generated 20,000 minimal pairs by pairing semantic frames with animate and inanimate verbs. To ensure semantic plausibility, they used a zero-shot prompt with Llama 3.3 70B to score prefixes, retaining only those with a plausibility score above 0.90. The task is evaluated using the average logit difference (avgLD) between two broad target sets:

  • Animate targets, restricted to unambiguous noun.person words such as officer or detective.

  • Inanimate targets, belonging to categories such as noun.event, noun.phenomenon, or noun.artifact.

Circuit discovery was performed via Edge Attribution Patching with Integrated Gradients (EAP-IG) on GPT-2 Small, Llama 3.2 3B, Gemma 3 4B, and Qwen 3 4B.

Circuit Discovery and Robustness

The study confirms that a causal mechanism responsible for handling animacy does exist, identifying circuits that are both sufficient and necessary to recover the behavior. To ensure the reliability of these findings, the authors performed several validation checks:

  • Random-edge controls: Circuits composed of randomly sampled edges failed to restore meaningful task performance.

  • Stability: Multiple discovery runs showed very high overlap between discovered circuits, with IoU values between 0.8–1.0.

  • Necessity: Ablating the top-20 highly ranked edges caused faithfulness to drop below 0.1, demonstrating the existence of a compact necessary core.

However, the researchers note that these circuits are less localized compared to other known ones and are more distributed than circuits found for phenomena like Indirect Object Identification (IoI).

Functional Role of Components

Analysis reveals that the animacy circuit is not concentrated in a single localized region. Instead, the most important components are primarily MLPs, spanning from the first layer toward middle/late layers, while attention heads carry a much smaller share of importance. The functional roles of these components appear to follow a specific pattern:

  1. Early components, such as MLP0, participate by routing or preparing information rather than directly writing the decision to the logits.

  2. The verb-level contrast is available in the earliest residual representations, but becomes logit-aligned only after later transformations.

  3. Attention heads mostly support positional routing, such as pathways from the verb to the passive marker and the subsequent determiner.

Generalization and Transferability

The researchers conclude that the discovered circuits should not be interpreted as universal animacy circuits, as their transfer to related settings is uneven across datasets and models. The study tested three transfer setups to evaluate this:

  • Setup 1 (BLiMP prefixes + original targets): The circuits generalized extremely well in most models.

  • Setup 2 (Original prefixes + named entities): The circuits generalized well, though faithfulness was substantially lower.

  • Setup 3 (BLiMP prefixes + named entities): The circuits do not generalize, and the mechanism breaks.

This suggests that while the models possess a robust part of the animacy computation, the circuit remains partly constrained by the specific task and data.

Improvements for AI systems

1. Circuit-Based Semantic Steering (Safety & Alignment)

  • Improvement: Implement surgical intervention protocols using the identified necessary core of animacy circuits (specifically targeting the high-attribution edges in early MLPs like MLP0 and subsequent logit-aligning MLPs). Instead of broad, unpredictable prompt engineering or RLHF, use Edge Attribution Patching (EAP) to boost or ablate specific edge activations during inference.

  • Improved Capability: The AI system can precisely control agentic behavior. It can be programmed to strictly avoid hallucinated agency (assigning intentionality to inanimate objects in safety-critical instructions) or, conversely, ensure highly consistent persona adherence in agentic workflows by reinforcing the specific pathways that drive animate-oriented logit preferences.

2. Semantic-Structure Contrastive Fine-Tuning (Robustness & Generalization)

  • Improvement: Transition from standard next-token prediction training to a Minimal-Pair Contrastive Objective. Utilizing the paper’s methodology, training data should be structured into clean (animate-favoring) and corrupt (inanimate-favoring) minimal pairs. The loss function should specifically penalize the model if the internal activations of the identified animacy circuit do not show a significant logit-difference margin between these pairs.

  • Improved Capability: The AI system will exhibit robust semantic generalization. It will move beyond task-specific pattern matching (which fails when both prefix and target change) toward a generalized concept of animacy, allowing it to correctly interpret agent/patient roles across diverse syntactic structures, such as active voice, passive voice, and complex relative clauses, without distributional drift.

3. Mechanistic Unit Testing in CI/CD Pipelines (Quality Assurance)

  • Improvement: Integrate an automated Mechanistic Interpretability Suite into the model deployment pipeline. This suite would use the paper’s minimal-pair dataset construction and average logit-difference (AVG LD) metric to perform semantic regression testing. It would specifically monitor whether fine-tuning or quantization has degraded the faithfulness or accuracy of the identified animacy circuits.

  • Improved Capability: The AI development lifecycle will gain a high-fidelity diagnostic tool. Engineers can detect semantic concept drift immediately after a model update, ensuring that optimizations (like pruning or quantization) do not inadvertently destroy the model's ability to perform complex reasoning regarding agency, causality, or intentionality.

4. Semantic-Importance Pruning & Distillation (Efficiency)

  • Improvement: Apply Circuit-Aware Model Compression. During distillation or pruning, use EAP-IG scores to identify and protect the necessary core edges and components (the small percentage of edges that provide 85%+ faithfulness). Non-essential edges—those not contributing to these vital semantic circuits—can be aggressively pruned.

  • Improved Capability: The system can produce highly efficient, small-parameter models (SLMs) that retain the sophisticated semantic reasoning capabilities of much larger models. This allows for high-reasoning, agentic AI to run on edge devices without sacrificing the ability to distinguish between intentional agents and inanimate causes.

Abstract

Distinguishing animate from inanimate concepts in written language requires more than shallow text processing, as it involves recognizing complex selectional constraints and contextual cues, such as verb-argument interactions. Yet, current large language models (LLMs) appear to be capable of doing it. We investigate whether this animacy-sensitive behavior of LLMs can be traced to a localized set of causally relevant components and connections. To do so, we construct a controlled dataset of minimal pairs and perform circuit discovery on four open-weight models. Through in-depth experiments and ablations, we show that a causal mechanism responsible for handling animacy in these models does exist, thus discovering an animacy circuit. At the same time, this circuit appears to be less localized compared to other known ones and generalizes only partially across models and animacy tasks, confirming the distributed, context-dependent, and somewhat graded nature of the animacy concept.

Sources

Related papers