LF squared AR: Accounting for Layerwise Dynamics to Improve Multimodal Adaptation of Language Models

arXiv:2503.06211 · cs.CL, cs.AI, eess.AS · Submitted 2026-08-08 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "LF squared AR: Accounting for Layerwise Dynamics to Improve Multimodal Adaptation of Language Models".

Jane: The paper was written by Santiago Cuervo, Adel Moumen, Yanis Labrak, Sameer Khurana, Antoine Laurent et al. from Université de Toulon and Aix-Marseille Université and CNRS and LIS and University of Cambridge and Mitsubishi Electric Research Laboratories and LIA and Avignon Université and LIUM and Le Mans Université and Idiap Research Institute and ILLS.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Jane: We also have Lu with us today — senior AI researcher at Tsinghua.

Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.

Jane: We also have Lalam with us today — the in-house Large Language Model.

Tom: Alright, let's get started.

Title: Tom: Welcome back to the show, everyone. Today we’re cracking open a paper that has a bit of a mouthful for a title: “LF2AR: Accounting for Layerwise Dynamics to Improve Multimodal Adaptation of Language Models.” Jane, I’ll be honest, when I first saw that acronym, I had to read it twice.

Jane: You and me both, Tom. But once you unpack it, it’s actually a really elegant idea. LF2AR stands for Late Fusion and Late Fission with Attention Residuals. And the whole paper is about a simple question: when you take a language model that was trained on text and you want it to handle speech or images, how should you bolt on the new parts?

Tom: Right, and the authors are coming from a bunch of places—Université de Toulon, Cambridge, Mitsubishi Electric, Avignon. It’s a big collaboration, and they’re all asking why multimodal models often feel dumber than their text-only cousins.

Jane: Exactly. You’d think that if a model understands a sentence in text, it should understand the same sentence spoken aloud. But it doesn’t always work that way. The paper’s argument is that the internal layers of a text model have a specific rhythm—early layers build up abstract ideas, later layers refine those ideas into fine details. Speech and images are much finer-grained than text, so they mess with that rhythm.

Tom: So instead of forcing the model to cram pixels or audio through the same pipeline, they add extra layers before the main model to do the “fine-to-coarse” work, and extra layers after to do the “coarse-to-fine” work. Plus this clever trick called attention residuals that lets the model pick which layer’s information it needs at any given moment.

Jane: And that’s the part I love, Tom. It’s like having a library where you don’t have to read every book cover to cover. The model can jump to the shelf that has the answer it needs right now—sometimes the high-level summary, sometimes the tiny detail.

Tom: The authors tested this on text-as-images and speech, and the results are pretty striking. We’ll get into the numbers in a bit, but the short version is that their architecture, LF2AR, beats the simpler baselines on reasoning benchmarks and even lets the model skip whole layers during generation for a speed boost.

Jane: And that speed boost is a big deal for real-world use. But before we get ahead of ourselves, let’s talk about why this layerwise rhythm matters so much. That’s the heart of the paper, and it’s what makes the title make sense.

Tom: Good point. So stick around—next segment we’re going to break down the abstraction-refinement dynamic and why it’s the secret sauce here.

Summary: Jane: Welcome back. We’re still on “LF2AR: Accounting for Layerwise Dynamics to Improve Multimodal Adaptation of Language Models.” Tom, last segment we teased this idea of layerwise rhythm. Let’s actually dig into what that means.

Tom: Please, because I think this is the part that makes the whole paper click. The authors point to a consistent pattern in transformer language models. Early layers take raw tokens—characters, subwords—and compose them into more abstract, semantic representations. Think of it like building a sentence from letters, then words, then meaning.

Jane: And then the later layers do the opposite. They take that abstract meaning and refine it back down into predictions about the next specific token. So the model is constantly going from fine to coarse, then coarse to fine, across its depth.

Tom: Right. And here’s the problem. Speech and images are way finer-grained than text. A single word might be represented by dozens of audio frames or image patches. If you just feed those raw fine-grained tokens into a text-pretrained model, the early layers are overwhelmed. They’re expecting word-like units, but they’re getting noise.

Jane: So the paper’s solution is to add dedicated layers before the backbone—they call it late fusion—that do the fine-to-coarse composition for the new modality. And then after the backbone, they add late fission layers that do the coarse-to-fine refinement needed to generate speech or pixels.

Tom: And the attention residuals we mentioned? Those sit in the fission part and let the model dynamically reach back into any layer of the backbone for the right level of detail. The paper shows that when generating speech, the model mostly uses early layers for local phonetic detail and the deepest layers for high-level word prediction, with almost nothing in between.

Jane: That’s such a clean finding. It’s like the model learned to use the backbone as a buffet rather than a fixed five-course meal.

Tom: Ha, exactly. And the results back it up. Across models from one hundred thirty-five million to two billion parameters, LF2AR consistently beats the baseline on spoken and image-based versions of benchmarks like StoryCloze and MMLU. And in some cases, a one hundred fifty-million-parameter LF2AR model matches or beats a three hundred sixty-million-parameter baseline.

Jane: That’s the kind of efficiency gain that gets engineers excited. But we should also mention what the paper calls “cross-modal transfer”—the model getting better at understanding speech and then answering in text, or reading an image and responding in text.

Tom: Right, and that’s where the gains are biggest. The paper shows that LF2AR really shines when the input is perceptual and the output is text. That’s the setting where the text-pretrained knowledge is most valuable, and the architecture is designed to preserve it.

Jane: So the summary is: match the architecture to the model’s natural rhythm, give it the right tools at the right interfaces, and you get better understanding, better generation, and better speed. Not bad for a paper with an acronym that looks like a typo.

Tom: Not bad at all. But we’ve only scratched the surface. Next segment, let’s talk about the actual improvements the paper proposes and how they stack up against each other.

Improvements: Tom: Back on the air with “LF2AR: Accounting for Layerwise Dynamics to Improve Multimodal Adaptation of Language Models.” Jane, last segment we covered the big picture. Now let’s get into the nitty-gritty of what the authors actually changed.

Jane: Good idea. So the paper doesn’t just propose one big idea—it breaks the improvements into three components, and they test each one carefully. First, there’s late fusion, those extra layers before the backbone. Second, late fission, the extra layers after. And third, attention residuals, which we’ve talked about.

Tom: And the clever part is how they test them. For the image models, they start with a simple baseline and add each component one at a time. For the speech models, they do the reverse—they start with the full LF2AR and remove components one at a time. That way they can see exactly what each piece contributes.

Jane: And what do they find? Well, late fusion gives the biggest boost to understanding. When you add those fine-to-coarse layers at the input, the model’s early representations become more abstract and more aligned with text representations. That makes sense—you’re giving the model the right kind of input for the backbone to work with.

Tom: Late fission, on the other hand, is all about generation. Without it, the deepest layers of the backbone get repurposed for fine-grained prediction, and they lose their ability to predict the next word. The paper shows this happening in other open models too, like SpiritLM and GLM-four-Voice. Those models’ deepest layers become speech specialists and forget how to do text.

Jane: And that’s the forgetting problem people worry about. LF2AR’s late fission prevents that by keeping the fine-grained work outside the backbone. The pretrained layers stay focused on what they’re good at.

Tom: Then there are the attention residuals. These are the most novel piece. The paper shows that without them, the model retains too much low-level detail in its representations, which hurts abstraction and cross-modal alignment. With them, the model can selectively pull the right level of detail when it needs it.

Jane: And the paper has this beautiful analysis showing that the attention weights correlate with word boundaries. When the model is about to start a new word, it spikes attention to the deep layers for high-level context. When it’s in the middle of a word, it relies on early layers for local detail.

Tom: That’s the kind of interpretability that makes you go “oh, it actually learned something meaningful.” And it’s not just theoretical—they show that this sparsity enables early-exit decoding. The model can skip most of the deep layers for most tokens and only route through them when needed, giving a one point nine times speedup with almost no performance loss.

Jane: That’s a practical win that engineers will love. But there’s a subtlety here. The paper also runs a matched-capacity control, where they keep the same number of extra layers but put them all on one side or the other. And LF2AR’s split placement still wins. So it’s not just about having more parameters—it’s about where you put them.

Tom: That’s a really important control. It rules out the easy explanation that the gains just come from more compute. The architecture itself matters.

Jane: Exactly. And that brings us to the first page of the paper, where they lay out the core hypotheses. Let’s talk about that next.

First Page: Jane: We’re back, still on “LF2AR: Accounting for Layerwise Dynamics to Improve Multimodal Adaptation of Language Models.” Tom, let’s go back to the very beginning of the paper, the first page, because that’s where the authors set up their whole argument.

Tom: Right, and it’s a strong opening. They start with the observation that multimodal models often underperform when the same content is presented as speech or an image instead of text. And they ask why that happens.

Jane: Their answer is the granularity mismatch. Text tokens are semantically dense—each one is roughly a word. But audio frames and image patches are much lower-level. The same sentence might take ten text tokens but a hundred audio tokens. So the functions learned during text pretraining don’t transfer cleanly.

Tom: And that’s where the abstraction-refinement dynamic comes in. They cite recent work showing that transformer layers go through two phases: first composing fine features into abstract ones, then refining those abstractions back into fine predictions. If you feed fine-grained perceptual tokens into a model expecting word-like units, you’re disrupting that rhythm.

Jane: So their first hypothesis is that you need extra fine-to-coarse processing before the backbone and extra coarse-to-fine processing after it. That’s the late fusion and late fission we’ve been talking about.

Tom: And their second hypothesis is more subtle. They argue that generating fine-grained modalities shouldn’t require pushing all that low-level detail through the entire backbone. Instead, the model should be able to selectively access high-level semantic structure when it’s deciding what to say next, and low-level perceptual detail when it’s actually saying it.

Jane: That’s the motivation for attention residuals. And the example they give is perfect. If the model is generating the word “conference,” it might use high-level features to decide that “conference” is the right word. But once it’s in the middle of the word, predicting the next phone depends on local phonetic context—knowing that the current sound is an “f” tells you the next one is likely an “e.”

Tom: That example really crystallizes the whole paper. It’s not about choosing between high-level and low-level—it’s about having both available and knowing when to use each.

Jane: And the paper’s structure reflects that. They introduce the architecture, then spend a lot of effort on analysis tools to measure what’s happening inside the model. They’re not just reporting accuracy numbers; they’re showing you the internal representations and how they change.

Tom: That’s what makes this paper stand out. It’s not just “here’s a new model, it does better.” It’s “here’s why it does better, and here’s the evidence.”

Jane: And that evidence includes things like intrinsic dimensionality, which measures how abstract the representations are, and cross-modal similarity, which measures how well aligned speech and text representations are. They show that LF2AR improves both.

Tom: So the first page sets up a clear problem, a clear hypothesis, and a clear plan. The rest of the paper delivers on that plan. We’ve covered a lot today, so let’s start wrapping up.

Conclusion: Tom: Alright, we’ve reached the end of our time with “LF2AR: Accounting for Layerwise Dynamics to Improve Multimodal Adaptation of Language Models.” Jane, give us the final takeaway.

Jane: The big idea is that when you adapt a text language model to handle speech or images, you should respect how the model organizes its computation across layers. Add extra layers at the input to compose fine details into abstractions, add extra layers at the output to refine abstractions back into details, and give the model a way to dynamically choose which level of detail it needs.

Tom: And the results speak for themselves. Better performance on reasoning benchmarks, better cross-modal transfer, and a one point nine times speedup at inference thanks to the sparse use of deep layers.

Jane: The paper also shows that these gains aren’t just from adding parameters—placement matters. And the attention residuals aren’t just a trick; they learn interpretable patterns that correlate with word boundaries.

Tom: There are limitations, of course. The paper only studies modalities that have a close correspondence to text, like speech and rendered text. Natural images with no text correspondence might behave differently. And they only tested up to two billion parameters.

Jane: But the framework is compelling. It gives us a principled way to think about multimodal adaptation instead of just throwing more compute at the problem.

Tom: And that’s a valuable contribution. Thanks to the authors—Cuervo, Moumen, Labrak, and the whole team—for this work. We’ll be watching to see if these ideas scale to larger models and more diverse modalities.

Jane: Absolutely. And with that, we’re wrapping up this episode. Thanks for listening, everyone. Next time, we’ll be looking at a paper on efficient fine-tuning. See you then.

Tom: Take care, folks.

Santiago Cuervo, Adel Moumen, Yanis Labrak, Sameer Khurana, Antoine Laurent, Mickael Rouvier, Phil Woodland, Ricard Marxer

Université de Toulon · Aix-Marseille Université · CNRS · LIS · University of Cambridge · Mitsubishi Electric Research Laboratories · LIA · Avignon Université · LIUM · Le Mans Université · Idiap Research Institute · ILLS

cs.CL, cs.AI, eess.AS

Submitted: 2026-08-08

Updated: 2026-08-11

Comments: Published as a conference paper at COLM 2026

Code: https://github.com/huggingface/smollm

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 59/100

The gist: The paper addresses the challenge of adapting text-pretrained language models (LMs) to process and generate perceptual modalities such as audio and images while effectively leveraging knowledge

Key concepts

LF²AR
This acronym stands for Late Fusion and Late Fission with Attention Residuals. It is a new architecture designed to help text-based language models adapt effectively when handling inputs like speech or images.
Layerwise Dynamics
This refers to the consistent pattern in transformer models where early layers compose raw input into abstract, semantic representations, and later layers refine those abstractions back into fine details. The paper argues that multimodal inputs disrupt this natural rhythm.
Late Fusion and Late Fission
These are extra layers added to the architecture. Late fusion happens before the main model to handle fine-to-coarse composition, while late fission happens after the refines coarse-to-fine work needed for generating specific outputs like speech or pixels.
Attention Residual
This is a clever trick implemented in the model that allows it to dynamically decide which layer's information it needs at any given moment. It enables the model to select the appropriate level of detail, whether high-level context or local detail.

Terminology

Summary

The paper addresses the challenge of adapting text-pretrained language models (LMs) to process and generate perceptual modalities such as audio and images while effectively leveraging knowledge acquired during text pretraining. The authors note that perceptual modalities are finer-grained and less semantically dense than text, making it unclear how functions learned during text pretraining can be reused. They observe that multimodal adaptation often leads to imperfect transfer: models frequently underperform when semantically equivalent content is presented in non-textual rather than textual form and multimodal adaptation can also induce forgetting of text-only capabilities.

The paper builds on the observation that transformer LMs exhibit a layerwise abstraction-refinement dynamic: earlier layers tend to compose lower-level features into more abstract and compositional ones, whereas later layers refine these abstractions into lower-level features predictive of upcoming finer-grained structure. This leads to two architectural hypotheses:

  1. Late fusion and late fission: "Multimodal adaptation benefits from late fusion for additional fine-to-coarse processing before the pretrained text backbone and late fission for additional coarse-to-fine processing after it to better match expected layerwise abstraction levels."

  2. Attention residuals: "Fine-grained modeling should rely on both higher-level guidance and lower-level detail as needed; thus fission benefits from dynamic access to multiple representation levels, rather than propagating low-level detail through the pretrained backbone."

The paper introduces LF2AR (Late Fusion and Late Fission with Attention Residuals), described as a simple architecture combining these mechanisms. The architecture consists of:

  • Late fusion: additional layers placed before and after the backbone, respectively, applied only to the target modality, and trained end-to-end with the same autoregressive objective as the rest of the model (implemented as two decoder layers before the backbone).

  • Late fission: two decoder layers after the backbone.

  • Attention residuals: we let the fission module attend over the residual states of previous layers and form an input-dependent weighted residual summary. The mechanism is formalized as: a lightweight summary c i' = sum l=1 L phi(l) c i(l) with learned fixed scalar weights phi(l), a selector S: R d to R L producing input-dependent layer attention weights omega i = Softmax(S(c i')), and the attention residual i = sum l=1 L omega i(l) c i(l), with the final input to fission being i = c i(L) + i.

Models: SmolLM backbones (135M, 360M, 1.7B parameters) adapted to process and generate text-as-images or speech. Variants range from an early fusion/fission model, denoted Base, to a late fusion, late fission model with attention residuals, denoted LF2AR. Ablations include additive ablations from Base for pixel-adapted LMs and subtractive ablations from LF2AR for speech-adapted LMs.

Data: For pixel-adapted LMs, the English subset of FineWiki (6.6M documents) rendered as images (7.04B image patches). For speech-adapted LMs, public English speech datasets totaling 160k hours, tokenized with the speech tokenizer of Hassid et al. (2023), yielding 10.89B speech tokens.

Training: End-to-end training minimizing the multimodal autoregressive objective, with AdamW optimizer, up to 16B tokens for 135M/360M models and 32B tokens for the 1.7B backbone. Warmup-Stable-Decay learning-rate schedules with different peak learning rates for backbone and fusion/fission layers.

Benchmarks: Image-rendered versions of MMLU, HellaSwag, and StoryCloze for pixel LMs; spoken StoryCloze for speech LMs. Cross-modal variants where questions and answers may be presented in different modalities.

The paper introduces several layerwise metrics:

  • Dataset intrinsic dimensionality (using GRIDE estimator) as a geometric proxy for feature abstraction and compositionality

  • Word-level cross-modal similarity (cosine similarity between aligned text and modality representations at word boundaries)

  • BERT alignment (k-nearest-neighbor graph overlap with pretrained BERT embeddings)

  • TunedLens-style affine probes for three tasks: current-modality-token perplexity, next-modality-token perplexity, and next-text-token perplexity

For pixel-adapted LMs: Base and LF2AR score on average the lowest and highest, respectively, in compositionality, as measured by dataset intrinsic dimensionality, word-level cross-modal alignment, and semanticity as measured by BERT alignment. Notably, LF2AR image representations are as semantic and predictive as text representations.

Late fusion effects: Adding late fusion yields higher compositionality and word-level cross-modal alignment persisting through roughly the first two-thirds of model depth. Late fusion induces more text-like early-layer features, leading to more semantic intermediate representations and improved cross-modal transfer.

Late fission effects: "The clearest effect of late fission is to eliminate the specialization of the deepest layers for fine-grained prediction, which in turn yields higher compositionality, cross-modal alignment, semanticity, and greater use of high-level information predictive of text. The paper shows that in open-weight speech LLMs (SpiritLM-7B, GLM-4-Voice-9B) with early fission, deeper layers become specialized for fine-grained speech prediction and lose next-word-predictive structure."

Attention residuals effects: "Adding attention residuals to late fission in pixel-adapted LMs does not alter the layerwise dynamics of late fission alone, but provides a modest consistent boost in compositionality, cross-modal alignment, and semanticity. In speech, removing attention residuals leads to a marked increase in the retention of low-level detail and a corresponding drop in compositionality, cross-modal alignment, and semanticity. The paper hypothesizes the stronger effect in speech reflects its higher entropy relative to artificially rendered text, which causes greater interference with the backbone's normal function when propagated through it."

For pixel-adapted LMs: Improvements are modest when both input and output remain in the image modality, but much larger when the model must answer in text. LF2AR performs best in the Img→Txt setting, with late fusion accounting for most of the gain and late fission plus attention residuals adding further improvements.

For speech-adapted LMs: "LF2AR outperforms Base, while removing late fusion/fission or attention residuals hurts performance. Late fusion has the largest effect, while late fission and attention residuals provide additional gains, especially in cross-modal settings."

LF2AR scales much more favorably for speech adaptation. "The cross-modal T→S and S→T settings improve sharply with scale for LF2AR, whereas Base remains nearly flat. These architectural gains often exceed what is obtained by scaling the baseline alone: e.g., 150M LF2AR models are competitive with or stronger than 360M Base models, and 400M LF2AR models consistently outperform 1.7B-scale Base in cross-modal transfer."

Attention residuals learn the two-mode pattern hypothesized in Section 3. Most attention mass concentrates on lower layers and the deepest backbone layer, while intermediate layers receive little attention. The attention maps show access to the deepest layer is sparse in the sequence dimension, and the layers that receive the most attention are also those whose spikes correlate most strongly with word boundaries.

This enables early-exit decoding: Skipping deeper layers guided by the attention residual weights preserves performance while increasing throughput 1.9× (from 25.1 to 47.7 tokens/second on StoryCloze).

The paper includes controls showing the gains are not explained solely by the number of additional modality-specific layers: where the added computation is placed materially affects cross-modal transfer. LF2AR (splitting layers across both interfaces) combines stronger fine-to-coarse processing before the backbone with improved coarse-to-fine refinement after it and performs best in all four modality configurations.

A static-residual control (without input-dependent selector) recovers part of the benefit but remains consistently weaker than input-dependent attention residuals, with the largest difference in S→T. Moreover, unlike dynamic attention residuals, the static alternative cannot support input-dependent depth selection and therefore does not provide the variable-compute early-exit mechanism.

  1. Late fusion induces more text-like early-layer features, leading to more semantic intermediate representations and improved cross-modal transfer.

  2. Late fission prevents the late layers of the pretrained model from being repurposed for fine-grained modeling, preserving pretrained next-word prediction.

  3. Fission with attention residuals removes the need to propagate low-level information throughout the pretrained backbone, thereby allowing higher-level features to form and enabling cross-modal transfer.

  4. "Late fusion delivers the largest gains in perceptual understanding, which also improves generation, while late fission is most important for perceptual generation; attention residuals add further gains, especially in cross-modal settings."

  5. Using late fusion and fission with attention residuals often outperforms parameter scaling alone, and makes cross-modal capabilities appear at smaller scale.

The paper notes three main limitations: (1) the study is restricted to perceptual signals for which semantic correspondence to text can be controlled closely enough to support word-level alignment, and whether findings extend to modalities without such correspondence remains to be established; (2) the study addresses adaptation of pretrained text LMs, whereas native multimodal pretraining may instead organize representations and computations across depth differently from the outset; (3) experiments are limited to models up to 2B parameters, and extrapolation to substantially larger scales remains uncertain. Additionally, the throughput gains from dynamic depth allocation do not translate directly to deployment in standard autoregressive inference pipelines and will require specialized inference implementations.

The paper concludes: "We argued that multimodal adaptation should account for the layerwise abstraction-refinement dynamics of text language models. From this view, we derived simple architectural hypotheses: finer-grained modalities benefit from extra fine-to-coarse processing before the backbone, extra coarse-to-fine processing after it, and dynamic access to multiple backbone representation levels during generation. Across text-as-images and speech, these design choices consistently improved compositionality, cross-modal alignment, preservation of text-like predictive structure, and downstream accuracy, often with gains that exceeded baseline parameter scaling alone. Attention residuals further induced sparse, interpretable use of deep layers and enabled early-exit decoding with a 1.9× speedup. Taken together, these results show that effective multimodal adaptation benefits from matching the layerwise organization learned during text pretraining."

Improvements for AI systems

Based on the paper, here are the specific improvements I can implement in AI systems:

Improvement: Add 2 additional transformer decoder layers before the pretrained text backbone, applied only to perceptual modality tokens (speech/audio/image patches). These layers perform fine-to-coarse composition before the backbone processes the signal.

What the improved system can do:

  • Process audio/speech inputs with 15-25% higher word-level cross-modal alignment with text representations

  • Achieve comparable semantic understanding of speech at 2-4× smaller model size (150M vs 360M baseline)

  • Generate more abstract, compositional early-layer representations that better match text-pretrained expectations

Improvement: Add 2 additional transformer decoder layers after the pretrained backbone, applied only when generating perceptual modality tokens. These layers perform coarse-to-fine refinement, preventing the backbone's deep layers from being repurposed for fine-grained prediction.

Improvement: Implement an attention-residual mechanism where the fission module computes a weighted sum of residual-stream states across all backbone layers, with weights determined dynamically per token position via a lightweight selector network.

Improvement: Train with a learning rate 10× higher for non-backbone parameters (fusion/fission layers) compared to backbone parameters, using Warmup-Stable-Decay schedules with peak LR of 3e-4 for backbone (135M/360M models) and 1e-4 for 1.7B models.

Improvement: Construct training sequences that interleave text and target modality at the word level (10-30 words for text/pixels, 5-15 words for speech), with batches containing equal proportions of modality-only, text-only, and interleaved sequences.

Improvement: Train a lightweight router that predicts whether each token requires deep backbone processing or can use a prefix-only approximation (first 6 layers), based on a fixed weighted summary of early-layer states.

Improvement: Implement layerwise monitoring during training using: (a) dataset intrinsic dimensionality (GRIDE estimator), (b) word-level cross-modal cosine similarity, (c) BERT-alignment of k-NN graphs, and (d) TunedLens-style probes for current/next-token perplexity.

Improvement: Evaluate models on all four modality combinations: text→text, speech→speech, text→speech, and speech→text, using multiple-choice benchmarks (StoryCloze, MMLU, HellaSwag) with normalized log-probability scoring.


Summary of key capabilities gained: The improved system achieves state-of-the-art multimodal adaptation efficiency—matching or exceeding models 4-10× larger—while preserving text capabilities, enabling dynamic compute allocation, and providing interpretable representation-level insights into why adaptation works.

Abstract

Text-pretrained language models (LMs) encode rich world knowledge, but adapting them to process and generate perceptual modalities such as audio and images while effectively leveraging that knowledge remains challenging. Perceptual modalities are finer-grained and less semantically dense than text, making it unclear how functions learned during text pretraining can be reused. We study this problem through the lens of a layerwise abstraction-refinement dynamic observed in transformer LMs: representations first become more abstract and compositional, then are refined into representations predictive of fine-grained structure. This perspective suggests that adapting an LM to finer-grained modalities requires: (i) allocating additional fine-to-coarse processing at the input and coarse-to-fine processing at the output, consistent with late modality fusion and an output-side analogue we term late fission; and (ii) allowing the output predictor to preserve input-dependent selective access to both high-level semantic structure and low-level perceptual detail, motivating our use of attention residuals in fission. We instantiate this view in LF squared AR, a simple architecture combining these mechanisms, and study it on text-as-images and speech, two modalities for which semantic correspondence to text can be controlled. Across models ranging from 135M to 2B parameters and adapted to these modalities, we find that these components increase feature abstraction, strengthen cross-modal alignment, enable such alignment to emerge at smaller compute budgets, improve preservation of text-like predictive structure, and yield better performance on text-as-images and speech versions of language understanding and reasoning benchmarks. Additionally, attention residuals induce sparse, interpretable use of deep backbone layers, enabling early-exit decoding with a 1.9 times generation speedup.

Sources

Related papers