Transformer See, Transformer Do: Copying as an Intermediate Step in Learning Analogical Reasoning

arXiv:2604.06501 · cs.LG, cs.CL · Submitted 2026-08-19 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "Transformer See, Transformer Do".

Jane: The research investigates how Copying mechanisms can serve as an intermediate step in learning complex analogical reasoning, drawing parallels between model attention patterns and human cognitive processes.

Tom: First, who's behind it and why it matters.

Title and authors: Tom: Okay, so we’re looking at the title "Transformer See, Transformer Do: Copying as an Intermediate Step in Learning Analogical Reasoning," and the authors are Philipp Hellwig, Willem Zuidema, Claire Stevenson, and Martha Lewis. The core idea here is that copying isn't just some side task; it’s a necessary step for the transformer to learn how to solve analogies robustly across different alphabets.

Jane: It’s interesting that they focus on letter-string analogies specifically and then extend the learning to new alphabets and new transformations. That suggests they are testing the model's ability to learn a general method rather than just memorizing specific inputs.

Lu: The authors seem keen on investigating whether meta-learning, using this copying mechanism, can improve the ability of transformers to discover solutions that generalize systematically over different alphabets and patterns. That’s a big question for how we teach these models to think flexibly.

Meng: I'm curious about how they handle those generalization tests; if the model struggles with new transformations, does that mean the copying mechanism isn't strong enough to enforce the required structural mapping? It’s a practical concern regarding deployment.

Lalam: If this work shows that we can use these intermediate steps to guide learning, it gives us a blueprint for developing AI systems that possess more flexible reasoning capabilities. This moves us closer to systems that can handle novel problems in ways we couldn't expect from current models.

The paper's summary: Tom: The paper summarizes their training process using Meta-Learning for Compositionality, or MLC, on letter-string analogies. Essentially, they found that when you include copying tasks in the training data, the models become much better at solving these analogies than if you didn't.

Jane: So, the main finding is that guiding the models to pay attention to those most informative problem elements—which copying tasks induce—is what makes them perform well on new alphabets. It’s a direct link between how the model attends and its ability to generalize.

Lu: They specifically evaluate how well these models learn the training task and then test their generalization to various targets, including new shuffled alphabets, new combinations of seen transformations, or entirely new transformations. That systematic testing is crucial for understanding the limits of this approach.

Meng: The paper points out a clear limitation when it comes to generalization; they found that the models struggle to generalize to new transformations in Section four. That means while copying helps with known patterns, it doesn't automatically solve the problem of applying something entirely novel.

Lalam: That limitation is actually very insightful because it tells us exactly where we need to focus our next efforts; the current mechanism is excellent at learning and adapting within a familiar framework but needs more support for true novelty.

The paper's improvements: Tom: Now, looking at how they suggest improving this, the authors focus on using Meta-Learning to discover general solutions systematically over alphabets and analogical patterns during training. They are trying to build a framework that lets the model learn *how* to learn the transformation itself rather than just memorizing it once.

Jane: The paper suggests that incorporating different permuted alphabets during training helps this meta-learning effect work better for letter-string analogies. It’s about introducing variety to force the model to develop deeper, more transferable rules for reasoning.

Lu: They are using a setup that involves a small encoder-decoder transformer trained on these datasets and then systematically evaluating its learning ability. This systematic evaluation is key to understanding the scope of what this meta-learning approach can actually achieve in practice.

Meng: I wonder if this meta-learning approach translates directly into practical application for more complex, multi-step tasks we see in the real world. We need to know if this learning structure holds up when the analogy requires several sequential operations rather than just one mapping.

Lalam: The improvements suggest that the path forward involves making training more structured around discovering general solutions instead of just solving specific examples. That focus on discovering underlying principles is exactly what we need to push AI toward greater versatility.

Conclusion: Tom: So, wrapping up the discussion on "Transformer See, Transformer Do: Copying as an Intermediate Step in Learning Analogical Reasoning," it seems the main point is that copying tasks act as a bridge that helps transformers learn how to solve analogies and generalize across alphabets.

Jane: That means we are seeing a clear way to guide the model’s attention toward what matters most in an analogy, which is really helpful for understanding how these systems develop intelligence. It connects the internal mechanism directly to human-like reasoning steps.

Lu: The implication is that if we can successfully implement this copying mechanism, we can create transformers that are much more adaptable and capable of handling a wider variety of reasoning problems than they currently are. That potential for adaptability is what keeps the theoretical side very excited.

Meng: From my perspective, if this copying mechanism helps with structure, it means we might be able to design AI that handles multi-step reasoning sequences more reliably. We need to see that structural guidance translate into stable performance in real-world applications before we can really talk about impact on complex tasks.

Lalam: Ultimately, this paper shows how we can use training structure to foster a deeper understanding of knowledge transfer in AI. It’s an important step toward making our AI systems not just reactive but truly capable of reasoning about new situations.

Philipp Hellwig, Willem Zuidema, Claire Stevenson, Martha Lewis

University of Amsterdam

cs.LG, cs.CL

Submitted: 2026-08-19

Updated: 2026-08-21

Importance score: 82/100

The gist: The research investigates how Copying mechanisms can serve as an intermediate step in learning complex analogical reasoning, drawing parallels between model attention patterns and human cognitive

Key concepts

Copying mechanisms
These tasks are included in the training data to help transformers learn how to solve analogies robustly across different alphabets. They guide the model's attention to informative problem elements, which is key for learning.
Meta-Learning for Compositionality (MLC)
This is a training process used where models are trained on letter-string analogies, and copying tasks are included. The goal is to use this method to teach the model how to learn the transformation itself rather than just memorizing specific examples.
Generalization
This refers to a model's ability to solve problems it has not seen before, such as new transformations or shuffled alphabets. The research shows that copying helps with known patterns but struggles with entirely novel transformations.
Attention patterns
These are the internal mechanisms of a transformer that show which parts of the input data are most important when processing information. The paper links how these attention patterns change during training to the model's ability to generalize.

Terminology

Summary

The research investigates how Copying mechanisms can serve as an intermediate step in learning complex analogical reasoning, drawing parallels between model attention patterns and human cognitive processes.

Alignment with Human Analogical Reasoning:

The study establishes a direct comparison between the matching step performed by the models and human analogy solving. The authors note that We identify parallels between the matching step of the models from Section 5 and human analogy solving. Based on this, they describe similarities to human analogical reasoning steps in Figure 12. Specifically, they find that model behavior aligns with two key human steps: "(1) encoding relevant information of the domains and (2) searching for and retrieving relationships and similarities between elements closely align with the attention patterns from head #7 in the second layer of the encoder (Figure 12, left)."

Model Performance Comparison (MLC vs. Frontier LLMs):

The performance of their best MLC model, Copy 200 - trained with copy tasks on 200 alphabets, is compared against frontier Large Language Models (LLMs) using various transformation types (Figure 13).

  • Strengths: The MLC model outperforms all frontier models on trained (compositional) transformations (blue coloured bars). Furthermore, for unseen compositional transformations involving two trained transformations (purple coloured bars with 2 transformations), the model performs at or above the level of the best-performing frontier models (GPT-OSS models).

  • Limitations: The model exhibits limitations when complexity increases. For instance, for the 'Extend and Group' transformation our models lag behind most frontier models. Moreover, for unseen compositions involving three trained transformations, the performance is noted as under par, indicating that our model cannot handle the increased complexity of applying three successive transformations. When dealing with transformations completely novel to the model (orange coloured bars), it is observed that our model does not perform as well as the frontier models.

Mechanism and Implementation (RASP-L):

The researchers detail their technical implementation using RASP-L for predecessor analogy, which involves several distinct computational layers:

  1. Locating Separators: The first layer of the RASP program is designed to locate separators within the input string, as shown in Figure 14.

  2. Locating Relevant Items: The second layer focuses on locating relevant items for the transformation (Figure 15). This involves calculating indices for elements such as x1 fst, x2 fst, and qy fst by aggregating selections based on the initial findings (fstind).

  3. Computing Transformation: The core transformation is computed using a linear combination of these indices, corresponding to an MLP operation: out_ind = x2_fst - x1_fst + qy_fst.

  4. Output Generation: The final layer builds component attention matrices (id, o slct, cpslct) and combines them using logical operations. The output is then aggregated to determine the final result: out = rsp. aggr (o slct, v = tn (s), default =0, reduction = ' mean '). This final attention pattern, as shown in Figure 16, is comparable to the matching head in their trained transformer.

Improvements for AI systems

Based on a meticulous review of this paper—which details advanced mechanisms for compositional reasoning, structured pattern matching, and the alignment between attention patterns and human analogy solving—the core opportunity lies in moving beyond mere attention-based correlation toward explicitly computable, modular reasoning structures that can guide the LLM's internal process.

The existing system is highly successful at pattern recognition within trained domains (Figure 13), but it lacks robust generalization for complex, multi-step, unseen compositions. The RASP implementation provides the blueprint for how to achieve this structure.

Here are the specific architectural and training improvements required:


The primary improvement is integrating a Program Synthesis Module (PSM) that operates alongside and guides the standard self-attention layers of the Transformer. This creates a hybrid architecture, which I call the Structured Reasoning Transformer (SRT).

We must replace or augment standard attention mechanisms with Circuited Attention Modules (CAMs).

  • Mechanism: Instead of allowing attention weights to be purely derived from dot products (softmax(QK T over sqrt d)), the SRT calculates a weighted attention score based on the output of a small, specialized circuit that models logical operations (like rsp.select, rsp.equals, and rsp.aggr).

  • Function: The CAM acts as a filter, forcing the attention head to attend not just to similar tokens, but to tokens that satisfy specific pre-defined relational constraints required by the task (e.g., Token X must appear immediately before Token Y, or The index of element A must be greater than the index of element B ).

  • Implementation Detail: This involves training a meta-controller that learns to generate the sequence of circuit calls (the program) necessary to solve the analogy, and then passes this program's output indices/values back into the attention calculation.

The model needs specialized pre-training to generalize beyond trained transformations.

  • Mechanism: We must train the SRT on a massive corpus of synthetic reasoning problems that require explicit decomposition. These problems are not just analogies, but structured tasks: Given A to B and C, find D such that... where the solution requires multiple, verifiable computational steps (e.g., subtraction, index mapping, grouping).

  • Loss Function: The training objective must include a Program Fidelity Loss (L PF). This loss penalizes the model not just for an incorrect final output, but for generating an internal reasoning path that is logically inconsistent or computationally impossible based on the rules of the task.

Total Loss = L Next Token + lambda times L PF

To handle unseen compositions (the weakness identified in Figure 13), the SRT requires a dedicated memory structure.

  • Mechanism: Implement a differentiable, external Symbolic Memory Store that holds the results of intermediate computations and defined relationships derived from the input prompt. When the model encounters a base domain element (e.g., 'ghi'), it first queries this SMS to retrieve all established facts/mappings regarding that element before proceeding to generate an attention vector.

  • Effect: This prevents the model from treating each step as purely sequential context and allows it to treat relationships as persistent, retrievable variables, similar to a database lookup.

The resulting Structured Reasoning Transformer (SRT) will transform the LLM from a powerful pattern correlator into a Generalized Computational Reasoner. Specifically, it can:

  1. Solve Arbitrary Compositional Transformations: It will move far beyond trained tasks (like reverse or shift). Given any set of N distinct, defined transformations (T 1, T 2,, T N), the SRT can correctly apply them in any sequence and combination (e.g., T new = T 3(T 1(T 2(Input)))), treating the process as a verifiable program execution rather than a statistical guess.

  2. Perform Explicit Relational Indexing: It can solve complex analogical tasks that require precise indexing and arithmetic operations, regardless of token placement or alphabet size (e.g., In this sequence, find the element whose index is A units greater than the index of element B, but only if both indices fall within the range defined by tokens 5 through 10).

  3. Self-Correct and Debug Reasoning: If a prompt contains contradictory or ambiguous information (e.g., Find X, given that X is both prime and divisible by 4), the SRT will not simply guess; it will use its internal circuit logic to identify the contradiction, flag it, and request clarification or state the impossibility of the task based on formal constraints.

  4. Generalize to Multi-Domain Problem Solving: By integrating symbolic memory and program induction, it can solve complex problems that cross domains—for instance, combining a linguistic analogy (like A:B:: C:D) with a mathematical constraint (e.g., the resulting value must be a perfect square) or a physical constraint (e.g., the arrangement must obey conservation of mass).

Sources

Related papers