Self-conditioned Flow Map Language Models via Fixed-point Flows
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.
Jane: Today's paper: "Self-conditioned Flow Map Language Models via Fixed-point Flows".
Tom: Self-conditioned flow language models implicitly learn a fixed-point iteration that refines its own denoising estimate, leading to novel flow map language models capable of one- and few-step generation.
Jane: First, who's behind it and why it matters.
Title and authors: Tom: We’re talking about "Self-conditioned Flow Map Language Models via Fixed-point Flows" today, and the authors are Jaehoon Yoo, Wonjung Kim, Floor Eijkelboom, Chanhyuk Lee, Nicholas M. Boffi, and Seunghoon Hong. Their title itself tells us they are bridging the gap between self-conditioning and flow maps using fixed-point dynamics.
Jane: It sounds complicated at first because of all those technical terms like "fixed-point flows," but essentially what they’re doing is showing that the iterative refinement inside self-conditioning can be viewed as a second dimension alongside the main flow process.
Lu: That's the key conceptual move; they are taking a behavior observed in self-conditioned models and giving it a formal mathematical framework, which is super creative because it turns an empirical observation into a definable structure.
Meng: So, if I understand right, instead of just training these massive models end-to-end on denoising loss, they are focusing on modeling the flow *and* that inner iteration separately to see how they interact?
Lalam: Precisely; the paper formalizes this by proposing fixed-point flows as a two-dimensional class of self-conditioned flows where one dimension is the actual flow process and the other dimension is that fixed-point iteration running at each step.
The paper's summary: Tom: So, summarizing what the paper actually presents, they show that self-conditioned flow language models inherently solve a fixed-point iteration that refines their denoising estimate, which is important because it explains the performance boost they see empirically.
Jane: They show this behavior emerges from self-conditioned training when you assume a contractivity condition holds, meaning the iteration naturally converges toward an ideal denoiser prediction. This convergence is what gives the model its improved ability to handle generation tasks.
Lu: The mathematical proof underpinning this convergence under that contractivity assumption is where it gets really interesting; it moves the behavior from just an observation to a theoretically derived property of self-conditioning itself.
Meng: I’m interested in the formula they show, t i+one = t i + (t i+one - t i) t i, which describes how the state updates and shares information across timesteps via z on top of the flow state x. That looks like a very structured way to handle temporal dependencies.
Lalam: That information sharing mechanism, where the update depends on both the flow state and this auxiliary variable z, is what they are leveraging to create this new fixed-point view of self-conditioning for their fixed-point flows.
The paper's improvements: Tom: Now for the improvements they suggest, it’s that by using this fixed-point view, we can define "fixed-point flows," which are valid flow maps that we can actually learn by compressing both the flow and the fixed-point iterations.
Jane: They propose replacing the complex self-conditioning state with its fixed point to define an ordinary flow, leading to a "fixed-point velocity" defined as b t(x):= t(x) - x one - t. This turns the process into a deterministic ODE, which is much cleaner for generation.
Lu: That ODE definition is powerful because it allows us to define a unique flow map operator X s,t that satisfies the composition law X s,t = X u,t X s,u, which confirms these fixed-point flows are mathematically sound and consistent across time steps.
Meng: So the practical improvement here is that instead of running a complex iterative denoising process during inference, we can use this derived fixed-point flow map to define a standard Euler scheme for generation. That sounds like it would significantly reduce the computational burden during text creation.
Lalam: And they show two distillation routes: one where you distill into a self-conditioning-free model, and another where you distill into a Flow Map Language Model FMLM that enables few-step generation by learning the two-time denoiser delta s,t.
Conclusion: Tom: So to wrap up on "Self-conditioned Flow Map Language Models via Fixed-point Flows," the paper successfully shows how self-conditioning implies a fixed-point iteration, which we can then use to define fixed-point flows and flow maps that are learnable through distillation.
Jane: The main implication is that we can distill those powerful, iterative self-conditioned models into deterministic flow maps, allowing us to achieve one- or few-step generation with competitive results on benchmarks like OpenWebText.
Lu: This gives us a clear path forward for understanding how to efficiently deploy these complex generative architectures by turning their implicit iterative mechanisms into explicit mathematical tools.
Meng: For me, the practical impact is that if we can distill these models down to flow maps, we gain a lot of speed and efficiency during deployment, which is crucial when you’re dealing with large-scale text generation pipelines.
Lalam: I think the biggest cultural implication is that this shows us how to systematically analyze and simplify complex AI behaviors by finding underlying mathematical structures, which helps build more interpretable systems overall.
KAIST University of Amsterdam Carnegie Mellon University
cs.CL, cs.AI
Submitted: 2026-07-01
Updated: 2026-10-01
Code: https://github.com/Ugness/self-conditioned-fmlm
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 83/100
The gist: Self-conditioned flow language models implicitly learn a fixed-point iteration that refines its own denoising estimate, leading to novel flow map language models capable of one- and few-step
Key concepts
- Self-Conditioned Flows
- These are language models that use self-conditioning during training. They solve a fixed-point iteration where the model refines its own denoising prediction based on previous steps. This process bootstraps performance, effectively learning how to improve its own output iteratively during generation.
- Fixed-Point Iteration
- Self-conditioning causes the denoiser to learn an iteration, $z_{j+1} = D(x, z_j)$, that converges exponentially to a unique fixed point $z^*$. This fixed point represents a self-corrected approximation of the optimal prediction. It is mathematically proven to emerge under specific training conditions.
- Fixed-Point Flow Maps
- These are ordinary flow maps derived by replacing the self-conditioning state with its learned fixed point. This results in a deterministic ODE defining a 'fixed-point velocity' $b^ ext{s}_t(x)$, which yields a flow map $X^ ext{s},t$. This map satisfies standard composition rules, enabling efficient one- and few-step generation.
Terminology
Summary
Self-conditioned flow language models implicitly learn a fixed-point iteration that refines its own denoising estimate, leading to novel flow map language models capable of one- and few-step generation. This research introduces fixed-point flows,
a mathematical framework that formalizes self-conditioned flows by modeling them as two dimensions: the flow process and the fixed-point iteration, which allows for the distillation of complex self-conditioned models into efficient, deterministic flow maps.
The Gist
Flow language models with self-conditioning solve a fixed-point iteration that bootstraps the performance of the learned denoiser, which can be characterized as a two-dimensional class of flows where the first dimension represents the original flow and the second represents fixed-point iterations.
Fixed-Point View of Self-Conditioning
The paper shows that self-conditioned flow language models implicitly learn a fixed-point iteration that refines the denoising estimate. This behavior is theoretically shown to emerge from self-conditioned training under a contractivity assumption. The generative process under self-conditioning involves updating the denoising estimate as well as the flow state:
xˆti+1 = xˆti + (ti+1 − ti) ˆbti, where ˆbti = zˆti+1 − xˆti 1 − ti.
This creates an information sharing across flow timesteps via zˆ on top of the flow state x.
Fixed-Point Flows and Their Flow Maps
The authors introduce fixed-point flows as a formalization of self-conditioned flows, where the second dimension represents fixed-point iterations running at each flow timestep. They show that these fixed-point flows define valid flow maps, which can be learned by compressing both the flow and the fixed-point iterations using respective few-step distillation objectives. The key observation is that this structure allows for distilling self-conditioned models into flow maps by compressing both components.
Self-Conditioning Induces a Fixed-Point Iteration
A self-conditioned denoiser, trained with a specific loss function (Equation 7), learns to improve its own predictions, implementing an iteration that approximately converges toward the ideal denoiser. This is formalized by defining an iteration:
zj+1 = Dˆt(x, zj).
The paper proves that under a contractivity assumption (Proposition 3.3), this iteration converges exponentially to a unique fixed point z⋆, which is viewed as a self-corrected approximation of the Bayes-optimal prediction Dt(x).
Fixed-Point Flow Maps for Few-Step Generation
The authors propose replacing the self-conditioning state with its fixed point to define an ordinary flow. This leads to the fixed-point velocity
defined as:
b⋆t(x):= D⋆t(x) − x 1 − t.
This results in an ordinary time-dependent ODE, defining a fixed-point flow,
whose solution operator X⋆s,t is the fixed-point flow map. This map satisfies the composition law:
X⋆s,t = X⋆u,t ◦ X⋆s,u.
Distillation Methods
The paper details two main distillation routes:
-
Distilling self-conditioned flow language models into self-conditioning-free models by learning the fixed-point denoiser D⋆ via a fixed-point distillation loss (Equation 19). This yields a
self-conditioning-free model ELF⋆
whose velocity is autonomous, enabling standard Euler scheme generation. -
Distilling self-conditioned flow language models into self-conditioned flow map language models FMLM⋆ by learning the two-time denoiser δs,t via a semigroup distillation objective (Equation 27). This yields a
flow map language model FMLM⋆
that enables few-step generation, achieving state-of-the-art performance on OpenWebText.
Empirical Results
Experiments on OpenWebText demonstrate that the resulting flow map language model, FMLM⋆, outperforms state-of-the-art self-conditioned models and few-step models in one- and few-step generation. Table 3 shows that online distillation saturates around 9 fixed-point iterations, achieving competitive quality at approximately 0.3× the training cost of offline distillation. FMLM⋆ attains the best gPPL among baselines that preserve data-level entropy (5.44 nats) in the one- and few-step regimes under matched sampling budgets.
Conclusion
The work demonstrates that self-conditioning induces a fixed-point iteration, which can be leveraged to define fixed-point flows and flow maps.
Improvements for AI systems
Here are the specific improvements to AI systems derived from this research, along with what those improved systems can achieve:
)Improved System Capability: State-of-the-Art Few-Step Generation
The primary improvement is the ability to generate high-quality text using only one or a few inference steps, rather than requiring many sequential steps.
-
Flow Map Language Models (FMLM⋆): The model can be distilled into a Flow Map Language Model (FMLM⋆). This model defines a deterministic evolution driven by an autonomous velocity field, allowing for text generation in one-step or few-step inference.
-
Efficiency Gains: Reduced Inference Latency and Cost
-
Distillation Capabilities: Transferring Performance
)Specific System Improvements & What They Can Do
The research introduces three main pathways for improvement, all aiming to distill the complex knowledge of large self-conditioned models into faster, more efficient flow map architectures.
-
Self-Conditioning Removal (Q3): Creating Self-Conditioning-Free Flows
-
Fixed-Point Flow Generation (Q4): Enabling Few-Step Distillation
-
Flow Map Language Model Training (Q4): Learning Complex Text Distributions Efficiently
)Detailed Mechanisms and Outcomes of the Improvements
Here is a breakdown of the specific technical improvements:
-
Improvement 1: Removing Self-Conditioning via Fixed-Point Distillation (Self-Conditioning Removal)
-
Mechanism: Fixed-Point Denoiser Distillation (L(D⋆))
-
What the Improved System Can Do: The system can be distilled from a highly effective, but computationally intensive, self-conditioned model (like ELF) into a standard flow map language model (FMLM⋆). This new FMLM⋆ is self-conditioning-free but retains the performance frontier of the original model. This allows for faster inference because it no longer requires the iterative self-correction mechanism during generation.
-
Improvement 2: Learning Flow Maps via Fixed-Point Flows (Few-Step Distillation)
-
Mechanism: Fixed-Point Flow Map Distillation (L(δ))
-
What the Improved System Can Do: The system can be distilled into a FMLM⋆ that is parameterized by a
two-time denoiser
(δs,t). This two-time denoiser directly learns the flow map operator, allowing for one-step generation and competitive performance against state-of-the-art few-step models. -
Improvement 3: Utilizing Fixed Points for Efficient Sampling (Initialization Heuristics)
-
Mechanism: Fixed-Point Iteration Convergence (Proposition 3.3 & B.2)
-
What the Improved System Can Do: The system can leverage the mathematical property that self-conditioned models implicitly solve a fixed-point iteration to improve sampling efficiency during training and inference initialization (warm/cold-start). Specifically, it allows for:
-
Warm-Start Sampling Heuristics: Using the previous timestep's fixed-point estimate as an initialization for the current step significantly improves convergence speed when finding the fixed point.
-
What the Improved System Can Do: The system can reach high generation quality (low gPPL) much faster during training or inference, effectively reducing the required number of flow iterations needed to achieve a target performance level.
)Summary of Overall System Advancement
The resulting AI systems are characterized by:
-
Speed: Capable of generating text in one or few steps (one-step generation).
-
Efficiency: Achieved through distillation into flow maps, reducing the complexity from iterative denoising to solving a deterministic ODE (the fixed-point velocity field).
-
Robustness: The system can maintain state consistency across flow timesteps via the learned two-time denoiser structure, ensuring that generation steps are mathematically coherent.
Abstract
Self-conditioning is a core technique that enhances continuous flow-based language models, where the model learns to denoise generated text by conditioning on its own denoising estimate. While empirically successful, its performance improvements are poorly understood. Moreover, there is growing interest in the use of few-step generators based on flow maps, for which how to leverage self-conditioning is unclear. Here, we show that flow language models with self-conditioning perform a fixed-point iteration that improves generation through iterative refinement. We use this viewpoint to formulate fixed-point flows, a two-dimensional class of self-conditioned flows, where the first dimension represents the flow process and the second represents the fixed-point iteration. We show that fixed-point flows define valid flow maps, and show that they can be distilled from self-conditioned flow models by compressing both fixed-point iterations and the flow process, the former with fixed-point distillation and the latter with flow map distillation. Our resulting flow map language model, FMLM, outperforms state-of-the-art self-conditioned models and few-step models in one- and few-step generation on OpenWebText. Code is available at https://github.com/Ugness/self-conditioned-fmlm.
Sources
- CoBit: Language Modeling with Bitstream Diffusion
- Spherical Flows for Sampling Categorical Data
- Analog Bits: Generating Discrete Data using Diffusion Models with Self-Conditioning
- LangFlow: Continuous Diffusion Rivals Discrete in Language Modeling
- Beyond Autoregression: Fast LLMs via Self-Distillation Through Time
- Language Modeling with Hyperspherical Flows
- Continuous diffusion for categorical data
- Distillation of Discrete Diffusion through Dimensional Correlations
- ELF: Embedded Language Flows
- Numerical Methods for Mean Field Games and Mean Field Type Control
- Flow Map Language Models: One-step Language Modeling via Continuous Denoising
- Consistency Deep Equilibrium Models
- Flow Matching for Generative Modeling
- One-step Latent-free Image Generation with Pixel Mean Flows
- How to Train Your Latent Diffusion Language Model Jointly With the Latent Space
- Fixed-Point Reasoners: Stable and Adaptive Deep Looped Transformers
- Discrete Flow Maps
- CANDI: Hybrid Discrete-Continuous Diffusion Models
- Categorical Flow Maps
- The Diffusion Duality
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering