TANGO: Treating Tokens as Operators
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "TANGO: Treating Tokens as Operators".
Jane: Token-Aggregated Nonlinear Gating Operators (TANGO) and its variant, Windowed Aggregation of Nonlinear Gating Operators (WANGO),
Tom: First, who's behind it and why it matters.
Title and authors: Tom: So we're talking about "TANGO: Treating Tokens as Operators," and it seems the paper opens by immediately setting up this new concept of treating tokens in a Transformer block as active operators. It introduces the idea that we can replace the standard self-attention and feed-forward sublayers with one unified cross-token gated update.
Jane: That title is quite evocative; "Treating Tokens as Operators" suggests they are no longer just passive data points that are fed into a network, but active components doing work on the representation themselves. It implies a much more dynamic relationship between tokens.
Lu: I think the authors are aiming to show that this specific mechanism provides a more direct pathway for source information to influence destination features than traditional self-attention does, which is a very important distinction they want to highlight.
Meng: I’m curious if this means we have less flexibility in how we define the operations compared to just having separate attention and FFN layers? The standard setup gives us more parameters per layer.
Lalam: They are trying to show that by collapsing those two layers into one gated update, you can achieve similar or better results with a different structural philosophy, which is always interesting for efficiency.
Tom: Right, so the paper argues that this single gated update replaces the standard self-attention and position-wise feed-forward network with one cross-token gated residual update. It's about unifying those two functions into one step.
Jane: And what they are emphasizing is that this new mechanism allows a gate computed at one position to modulate the projected features at another based on source token interactions, which is a key feature of the TANGO model.
Lu: This unification means that the traditional separation between interaction and transformation is being dissolved into a single step, which could lead to more integrated representations in sequence modeling.
Meng: From an engineering perspective, integrating things might make debugging harder if something breaks because you can't isolate whether it was the attention part or the feed-forward part that caused an issue.
Lalam: But structurally, it seems cleaner because you have a single control point for feature scaling at each destination position rather than two separate operations happening sequentially.
Tom: So they are proposing this unified cross-token gated residual update as their core mechanism, replacing the standard self-attention and position-wise feed-forward network with one single step.
Jane: And that means every source token generates a SwiGLU gate vector, and these gates are then aggregated to determine how they scale the destination features before adding them to the residual stream.
Lu: That’s the essence of TANGO: source tokens determine feature-wise scaling at a destination rather than supplying value vectors averaged by self-attention.
Meng: It seems like they are proposing a different way of handling feature flow entirely, not just tweaking existing components, which is a significant architectural change.
Lalam: It’s about fundamentally changing the structure so that the source tokens dictate the feature scaling at each destination based on their specific interactions.
The paper's summary: Tom: Moving on to summarizing what "TANGO: Treating Tokens as Operators," it boils down to this new block structure where you replace two standard sublayers with one cross-token gated residual update. This update uses a SwiGLU gates derived from the source tokens.
Jane: So, the summary is that each source token produces its own SwiGLU gate vector and a projected feature vector, and these components are then used to scale the destination's features via an averaged gate before adding them to the residual stream.
Lu: The main point is that query-key similarities determine a weighted average of the source gates available to each destination, which then rescales the destination's projected features before they are added to the residual stream. That’s how it operates.
Meng: So, what I need to grasp is that this weighting step—how we decide which source gates are important for a given destination—is done through query-key similarities rather than a standard self-attention mechanism determining the value vectors.
Lalam: That's right; they compute these weights as softmax weights from scaled cosine similarities between every destination and every causally visible source, and this allows the gate to act as a diagonal scaling operator on the destination features W v,x i.
Tom: So it’s about moving away from having self-attention average value vectors and instead using these computed weights to scale the projected features directly.
Jane: That means that source tokens are actively determining the feature scaling at a destination rather than just supplying values that are averaged by self-attention. It’s a different way of thinking about dependency modeling.
Lu: This mechanism is designed so that source tokens determine feature-wise scaling at a destination, which is the core mechanism described in Equation (one) for this paper.
Meng: So, if I understand correctly, the source tokens are actively controlling the flow of information by scaling features before they enter the residual stream based on their interaction.
Lalam: Precisely; those components—the gates and projected features—are derived from different positions, and then content-dependent weights average those gates to form that gate i for each destination i.
The paper's improvements: Tom: Now, let's talk about the specific improvements they suggest in this paper. They focus on the architectural structure itself and how we compute those weights to give us better control over the model's behavior.
Jane: So one major suggestion is moving from standard self-attention and position-wise feed-forward networks to that single cross-token gated update, which is a big structural change in how information moves through the model.
Lu: They are also proposing two specific aggregation methods: TANGO, which computes weights for every visible source using those softmax weights over scaled cosine similarities between every destination and every causally visible source.
Meng: And then they have WANGO, which is a variant that keeps the same unnormalized exponential similarity scores within a recent window for sources but uses a positive feature map weighting rule for older sources to maintain statistics with running sums.
Lalam: These aggregation methods give us two different ways to compute those source weights, one being quadratic in sequence length for TANGO and the other being linear in sequence length when the window and feature dimensions are fixed.
Tom: So it’s a direct trade-off between the full prefix architecture, TANGO, which is complex but likely yields lower NLL on benchmarks like FineWeb-Edu, and WANGO, which offers better complexity scaling for long contexts.
Jane: The overall improvement they point to is that by offering both options based on these different weighting schemes—TANGO versus Wango—they provide a choice based on the required sequence length characteristics.
Lu: This flexibility in choosing between the two methods allows researchers to tailor the architecture precisely to whether they prioritize maximum performance or computational efficiency for a specific task.
Meng: It’s about giving us tools that address both high-accuracy needs and high-efficiency needs simultaneously, which is something I can actually work with when designing practical systems.
Lalam: And another improvement is the use of source-conditioned SwiGLU gating, ensuring that the gates and projected features come from different positions for each token.
Conclusion: Tom: So to wrap up "TANGO: Treating Tokens as Operators," the main point is this new block structure replaces self-attention and FFN with one cross-token gated residual update where source tokens actively determine feature scaling at a destination based on their interactions.
Jane: It’s a really interesting idea because it shifts the focus from passive averaging to active, context-aware modulation of features using source gates. This gives us a new way to model dependencies that is worth exploring further in the future.
Lu: The paper's contribution lies in showing that this method can yield strong results across different language modeling benchmarks by proving its effectiveness in practice and providing a solid framework for future research into operator-based architectures.
Meng: For me, the practical implication is that if we can successfully implement WANGO with linear complexity, it means we could handle much larger sequences efficiently in deployment without massive resource overhead.
Lalam: I think this work provides a solid foundation for building more efficient and context-aware AI systems that are better suited for real-world applications by giving us explicit control over how features are modulated based on source context.
Tom: So that's our rundown of "TANGO: Treating Tokens as Operators," and it’s a paper that shows a new way to structure the core processing of sequence data. We’ll be taking these ideas off the air right now, and I think we’ve got some really exciting directions to follow.
Jane: It's been great talking about how this new gating mechanism works in plain terms. We’ve got some fascinating stuff coming up next on the air soon, so stay tuned for those updates.
Lu: I’m really looking forward to seeing how this operator-based thinking evolves and what new possibilities it unlocks for sequence modeling.
Meng: I hope we can see practical applications of these efficiency gains in real-world systems soon.
Lalam: This paper on "TANGO: Treating Tokens as Operators" gives us a solid framework for building more efficient and context-aware AI systems that are better suited for real-world applications by giving us explicit control over how features are modulated based on source context.
Department of Informatics · Luddy School of Informatics, Computing, and Engineering · Cognitive Science Program · Indiana University Bloomington
cs.LG, cs.CL
Submitted: 2026-08-22
Updated: 2026-09-30
Importance score: 83/100
The gist: Token-Aggregated Nonlinear Gating Operators (TANGO) and its variant, Windowed Aggregation of Nonlinear Gating Operators (WANGO), introduce a novel block structure that replaces standard
Key concepts
- Cross-Token Gated Residual Update
- This is the core update mechanism that replaces traditional Transformer layers. Instead of standard attention and feed-forward steps, it uses a single operation where a gate derived from one token (source) scales the projected features of another token (destination). This allows information flow based on direct source-destination interactions.
- Tango Model
- The full-prefix architecture that computes weights for every causally visible source. It uses softmax weights derived from scaled cosine similarities between every destination and its visible sources. This results in a quadratic complexity in sequence length but achieves the lowest mean validation NLL across tested benchmarks.
- WANGO Aggregation
- A variant of TANGO designed for better scaling. It retains similarity scores within a recent window but uses running sums derived from a positive feature map to weight older sources. This allows WANGO to achieve linear complexity with respect to sequence length when the window and feature dimensions are fixed.
- SwiGLU Gate
- The Swish-gated linear unit used by source tokens. It produces both a gate vector and a projected feature vector. The gate controls how much of the projected features is passed through, effectively acting as a content-weighted scaling operator for the destination features.
Terminology
Summary
Token-Aggregated Nonlinear Gating Operators (TANGO) and its variant, Windowed Aggregation of Nonlinear Gating Operators (WANGO), introduce a novel block structure that replaces standard self-attention and position-wise feed-forward networks with a single cross-token gated residual update, demonstrating superior performance across various language modeling benchmarks. The core innovation lies in allowing a gate computed at one position to modulate the projected features at another based on source token interactions.
The gist: TANGO computes a separate weight for every causally visible source, using query–key similarities to determine a weighted average of the source gates available to each destination, rescaling projected features before adding them to the residual stream.
How it works
The fundamental update mechanism replaces the standard Transformer sublayers with a single cross-token gated residual update. Each source token produces a Swish-gated linear unit (SwiGLU) gate vector and a projected feature vector. For each destination, these source gates are aggregated to form an average gate, which then scales the destination's projected features before they are added to the residual stream:
For source representation xj and destination representation xi, the corresponding update is f(xi; xj) = Wo[SiLU(Wgxj) ⊙ Wvxi].
(1)
Token Aggregation and Weighting
The model computes weights for source gates based on similarities between every destination and every causally visible source. The Tango model specifically calculates these as softmax weights from scaled cosine similarities between every destination and every causally visible source.
This process allows the gate to act as a diagonal scaling operator on the destination features Wvxi,
meaning it controls feature-wise scaling at a destination rather than supplying value vectors averaged by self-attention.
TANGO vs. WANGO Aggregation
The paper introduces two aggregation methods:
-
Tango (Full-Prefix Architecture): This model computes weights for every visible source, resulting in a
quadratic in sequence length
complexity because it computes softmax weights from scaled cosine similarities between every destination and every causally visible source. It is described as thefull-prefix architecture.
-
WANGO (Windowed Aggregation): This variant retains the same unnormalized exponential similarity scores within a recent window, but for older sources, it uses a
positive feature map weighting rule whose required statistics can be maintained with running sums.
This results in a computation that islinear in sequence length for fixed window and feature dimensions.
Architectural Comparison and Scaling
The models are compared against several baselines, including Recurrent Transformer++, Untied Transformer++, the full-attention Gated Attention Unit (GAU), and FLASH. The comparison focuses on parameter matching:
We adjust each architecture’s internal hidden dimensions to match approximately 44.3 million distinct parameters outside the tied token embedding/output matrix.
The proposed models (Tango and Wango) apply one shared block four times, while Untied Transformer++, GAU, and FLASH use four independently parameterized blocks. The comparison shows that among architectures with computation linear in sequence length, the Wango model has the lowest mean validation negative log-likelihood (NLL).
Performance Results
Across benchmarks like FineWeb-Edu, Lean, and DeepMind Mathematics:
The Tango model obtains the lowest mean validation NLL on FineWeb-Edu, Lean, and DeepMind Mathematics.
On FineWeb-Edu, Wango has the lowest NLL among architectures with computation linear in sequence length
and lower than Recurrent Transformer++ at a similar analytical forward-pass multiply–accumulate count. On Lean, Tango obtains the lowest joint validation NLL at 1.719 ± 0.026.
The analytical operation counts highlight the complexity difference: Tango is quadratic in sequence length,
whereas Wango is linear when the window and feature dimensions are fixed.
Contributions
The main contributions are:
"We introduce the Tango model, which replaces self-attention and the position-wise feed-forward network with one gated update. Source tokens produce SwiGLU gates, and each destination applies their content-weighted average to its own projected features."
The development of WANGO, which utilizes direct pairwise weighting within a recent window and running sums derived from a positive feature map for older sources,
achieving linear complexity for fixed dimensions.
The comparison of six architectures under matched nonembedding parameter counts, training objectives, data order, and training budgets. Among those with linear sequence-length computation, the Wango model exhibits the lowest mean FineWeb-Edu validation NLL. The full-prefix model (Tango) achieves the lowest mean validation NLL overall on all three benchmarks.
Limitations
The paper notes limitations:
**"We do not match forward-pass operation counts, and these experiments are not sufficient to establish scaling laws.
Improvements for AI systems
Here are the specific improvements that can be made to AI systems based on this research, and what these improved systems can achieve:
) 1. Implement Cross-Token Gated Residual Updates (TANGO/WANGO):
Instead of using a standard self-attention layer followed by an independent feed-forward network, replace the entire block with a single cross-token gated residual update.
The improved system can model long sequences more efficiently and capture complex interactions between source tokens and destination features directly within one step, potentially leading to better performance in tasks requiring deep contextual dependencies.
) 2. Utilize Source-Conditioned SwiGLU Gating:
Each source token should produce a Swish-gated linear unit (SwiGLU) gate vector, which then modulates the projected features of the destination token based on query-key similarities.
This allows for feature-wise scaling at a destination that is dynamically determined by the content and context of all causally visible source tokens, enabling more nuanced feature modulation than independent position-wise transformations.
) 3. Employ Full-Prefix Aggregation (TANGO) for State-of-the-Art Performance:
For maximum performance on benchmarks like FineWeb-Edu, Lean, and DeepMind Mathematics, use the full prefix aggregation rule where the model computes softmax weights over all causally visible source logits (temperature-scaled cosine similarities).
This architecture is capable of achieving the lowest validation Negative Log-Likelihood (NLL) across a broad range of natural language and formal reasoning tasks.
) 4. Implement Windowed Aggregation (WANGO) for Efficiency in Long Contexts:
For applications where sequence length scaling is critical, use the Wango model, which retains unnormalized pairwise scores within a recent window and uses a positive feature-map weighting rule for older sources, resulting in linear complexity with respect to sequence length.
This enables the AI system to process extremely long documents or proofs (like those in Lean) while maintaining computational efficiency that scales linearly rather than quadratically with context length, making long-context language modeling feasible.
) 5. Achieve Parameter Efficiency Through Block Sharing:
Implement the architecture by reusing a single parameterized block across multiple sequential applications (e.g., four times for the proposed models).
This allows for scaling up model depth (number of layers) without incurring a proportional increase in non-embedding parameters, leading to deeper, more complex models with comparable parameter counts to less-deep but more independent architectures.
) 6. Enhance Reasoning and Proof Completion Capabilities:
Specifically target formal reasoning tasks by utilizing the Tango architecture on Lean source code modeling and proof completion targets.
The improved system can perform higher-quality source code analysis, generate more coherent proofs, and achieve superior performance in mathematical problem-solving benchmarks compared to standard Transformer baselines.
) 7. Optimize for Task-Specific Supervision:
For mathematical tasks (DeepMind Mathematics), ensure the loss function is applied only to the answer characters and the terminal EOS symbol, using context length 256.
This allows the AI system to be highly effective at generating precise mathematical answers within a fixed, manageable context window.
In summary, these improvements transform a standard Transformer block into a more context-aware, efficient gated mechanism capable of achieving state-of-the-art results in both natural language understanding and formal mathematical reasoning by intelligently aggregating source information.
Sources
- Longformer: The Long-Document Transformer
- Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach
- From Growing to Looping: A Unified View of Iterative Computation in LLMs
- Generative Language Modeling for Automated Theorem Proving
- GLU Variants Improve Transformer
- RoFormer: Enhanced Transformer with Rotary Position Embedding
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks