E-MoE: Enhanced Mixture-of-Experts for Non-Factorized Diffusion Language Models
cs.CL
Submitted: 2026-09-29
Updated: 2026-09-29
Code: https://github.com/addtt/variational-diffusion-models
Terminology
Sources
- Forward-Learned Discrete Diffusion: Learning how to noise to denoise faster
- Beyond Autoregression: Fast LLMs via Self-Distillation Through Time
- TUBE: Tangent Upper Bound on Evidence for Discrete Diffusion Language Models
- Categorical Reparameterization with Gumbel-Softmax
- Auto-Encoding Variational Bayes
- Variational Diffusion Models
- IDLM: Inverse-distilled Diffusion Language Models
- Breaking the Factorization Barrier in Diffusion Language Models
- Discrete Copula Diffusion
- Discrete Diffusion Modeling by Estimating the Ratios of the Data Distribution
- Large Language Diffusion Models
- Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer
- Self-conditioned Flow Map Language Models via Fixed-point Flows
- LLaDA-MoE: A Sparse MoE Diffusion Language Model
- Di$\mathtt{[M]}$O: Distilling Masked Diffusion Models into One-step Generator
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering