dQwen3.5: Hybrid-Attention Diffusion Language Models
cs.CL, cs.LG
Submitted: 2026-09-17
Updated: 2026-09-17
Code: https://github.com/QwenLM/Qwen3.5
License: http://creativecommons.org/licenses/by/4.0/
The gist: Adapting a pretrained autoregressive (AR) model is a cost-efficient route to a diffusion language model (DLM).
Terminology
Abstract
Adapting a pretrained autoregressive (AR) model is a cost-efficient route to a diffusion language model (DLM). While nearly all such adaptations start from a full-attention transformer, AR modeling has shifted toward hybrid architectures that interleave attention and RNN layers. This creates an obstacle for adaptation: unlike attention, RNNs are structurally causal and nontrivial to bidirectionalize. Despite this mismatch, we investigate whether such backbones can become effective DLMs by adapting Qwen3.5 at 0.8B, 2B, 4B, and 9B scales, yielding the dQwen3.5 family. We find that hybrid backbones can be efficient starting points for adaptation: against a full-attention control, the hybrid reaches a given training loss in about half the tokens. Across scales, dQwen3.5 resembles full-attention DLMs in any-order decoding behavior and performs strongly under parallel decoding.
Sources
- Program Synthesis with Large Language Models
- LLaDA2.0: Scaling Up Diffusion Language Models to 100B
- CoDA: Coding LM via Diffusion Adaptation
- Evaluating Large Language Models Trained on Code
- Training Verifiers to Solve Math Word Problems
- Nemotron-Labs-Diffusion: A Tri-Mode Language Model Unifying Autoregressive, Diffusion, and Self-Speculation Decoding
- Continual Pre-Training of Large Language Models: How to (re)warm your model?
- Qwen2.5-Coder Technical Report
- Mercury: Ultra-Fast Language Models Based on Diffusion
- Jamba: A Hybrid Transformer-Mamba Language Model
- MiniMax-01: Scaling Foundation Models with Lightning Attention
- Improved Large Language Diffusion Models
- Nemotron 3 Nano: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic Reasoning
- NVIDIA Nemotron Nano 2: An Accurate and Efficient Hybrid Mamba-Transformer Reasoning Model
- Qwen2.5 Technical Report
- Demystifying Diffusion Objectives: Reweighted Losses are Better Variational Bounds
- Dream-Coder 7B: An Open Diffusion Language Model for Code
- Qwen3 Technical Report
- Dream 7B: Diffusion Large Language Models
- FLARE: Diffusion for Hybrid Language Model
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering