zgba 站群
How to build a diffusion language model

How to build a diffusion language model

An introduction to diffusion language models and the research advances that underlie today’s diffusion LLMs. We describe the building blocks of recent open-source models, starting from simple masking diffusion, and including techniques for iterative refinement, post-training, and variable-length generation. Material is adapted from workshop talks and lectures at ICLR 2026 and MLSS 2026.

Two families of generative AI algorithms are widely used today. For continuous data such as images or video, the state-of-the-art approach is based on diffusion models. For discrete data such as text or code, the standard approach is instead autoregressive models. This article explores an alternative for discrete data, one built on the modern paradigm of diffusion.

Mainstream language models are autoregressive: they generate tokens left-to-right, one at a time, each conditioned on the tokens before it. This approach is powerful, but it also has inherent limitations:

Diffusion models take a different approach. Rather than producing text one token at a time, they generate the whole sequence at once, starting from an initial guess and iteratively refining it over a number of steps. This unlocks several advantages: generation can trade off speed and quality by using fewer or more steps, mistakes can be corrected along the way, and every step attends to bidirectional context.

Applying diffusion to language had long been an open problem. In 2024 the field reached a turning point, as diffusion models became competitive with autoregressive models on quality. By 2026, diffusion LLMs are a reality, with releases from leading industry labs — Mercury 2 (Inception Labs) , Gemma Diffusion (Google) , and Nemotron Diffusion (NVIDIA) . This article traces the ideas and papers that underlie these modern models.

Before introducing diffusion for language, we start with a brief overview of Gaussian diffusion for image generation. We will then build up discrete diffusion by analogy.

The central concept underlying diffusion models is denoising. Instead of painting an image in one shot, a diffusion model produces images step by step, starting from pure random noise and removing a little of it at every step until a coherent image emerges. Generating an image through many small steps turns out to be far simpler than producing it all at once, and this is what makes diffusion models so effective.

How does a model learn to denoise? The trick is to teach it by showing examples of noise being gradually transformed into an image. Diffusion achieves this via two complementary processes. First, a forward process takes a clean source image and turns it into pure noise, one step at a time. Second, a reverse process learns to invert this transformation, turning pure noise back into an image; it is trained on the image-to-noise trajectories produced by the forward process.

The forward process takes a clean training image and produces a sequence of increasingly noisy images that trace a path from clean data to pure noise. It does this by mixing in a growing amount of random Gaussian noise at each step, until the image dissolves into pure static. This step requires no learning at all — we are simply adding noise — yet it is enormously useful, because it manufactures an endless supply of training data: examples of images being transformed into noise, and vice versa.

The reverse process is where the actual learning happens. We train a model to transform noise into images by following the steps produced by the forward process in reverse.

Concretely, given a noisy image, we train a machine learning model to separate the noise from the underlying image or, equivalently, to predict either the noise that was added or the clean image itself, since given the noisy input, knowing one determines the other. Once the model can do this, generation is simple: start from pure noise, ask the model to estimate and strip away a bit of it, and repeat. Each pass nudges the sample a little closer to something that looks like real data, until a clean image remains.

This forward/reverse recipe — corrupt data with noise, then learn to reverse the corruption one step at a time — is the blueprint for every diffusion model.

The main obstacle in bringing diffusion to language is deciding what “noise” should mean for discrete tokens. For example, the noise used in classical diffusion is Gaussian, and adding continuous Gaussian noise to categorical variables is not well-defined. Below we introduce one simple yet effective approach that defines noise via masking. Our group popularized this approach, and it now forms the basis of most open-source diffusion language models.

The easiest way to understand masked diffusion is as an unmasking transformer. We train the model by taking clean sequences, masking a random fraction of their tokens, and asking a bidirectional transformer to fill in the blanks. If you know BERT, this is essentially BERT with a randomized masking rate — but unlike BERT, the resulting model is generative. You can think of masked diffusion as a generative BERT.

Once we trained the unmasking transformer, we can generate text by starting from a fully masked sequence and repeating two steps many times:

Each round leaves fewer positions masked, until the sequence converges to a clean sample from the model. Generation thus amounts to starting from a sequence full of blanks and gradually filling in words in an arbitrary order.

We can also understand a bit better why this process works by framing it as an analog of the Gaussian diffusion model we saw earlier. Just like Gaussian diffusion, masked diffusion can be described as a model consisting of a forward and a reverse process.

The goal of the forward process is to generate training data for the reverse process. Its output is a trajectory that starts from a datapoint and ends at a sequence of pure noise; the reverse process will then be trained to produce this trajectory in reverse.

The key challenge is deciding what “noisy” should mean. In Gaussian diffusion, we added varying amounts of white noise to an image. In masked diffusion, we instead randomly mask a fraction of the tokens in a discrete sequence. The amount of masking is governed by a schedule alpha_t — the probability that a given token remains unmasked — which plays the role of the signal-to-noise ratio in Gaussian diffusion. It starts at 1 when t = 0 (a clean sequence) and decreases to 0 when t = 1 (a fully masked sequence). The time variable t indexes a path from clean to noisy data, and at time t a partially masked sequence z_t has, in expectation, a fraction alpha_t of its tokens unmasked.

We implement this process as a Markov chain over a sequence of variables z_t indexed by t, with z_0 being the clean, unmasked sequence. For s < t, the chain defines q(z_t mid z_s) by masking each still-unmasked token of z_s with probability (alpha_s - alpha_t)/alpha_s. Running this Markov chain for a number of steps produces a trajectory going from clean data to fully masked noise.

Next, as in Gaussian diffusion, we train the reverse process to walk the sequence of increasingly masked latents in reverse — starting from a fully masked sequence and ultimately generating outputs similar to clean data.

Using Bayes’ rule, we can derive the mathematically optimal reverse process q(z_s mid z_t, x) when the clean sequence x is known . This optimal process has two steps: (1) given a partially masked z_t, we peek at x to find the true clean tokens; (2) form z_s by replacing each masked position of z_t with its value in x with probability (alpha_s - alpha_t) / (1 - alpha_t), and otherwise leaving it masked.

In practice, the final output x is obviously unknown when we generate it. We therefore train a model x_theta(z_t) to predict the final clean sequence given the current state z_t and apply the ideal reverse process q(z_s mid z_t, x) using the estimate x_theta(z_t) in place of the real x. More formally, we define the reverse process as a probability p(z_s mid z_t) = qbig(z_s mid z_t, x_theta(z_t)big). This definition recovers the sampling algorithm we described earlier: at each step, we use the model x_theta(z_t) to fill in the blanks of z_t, and we keep a subset of these filled-in tokens in z_s.

Putting these pieces together gives us the mathematical definition of a masked diffusion language model (MDLM). The forward process q(z_t mid z_s) produces a trajectory from clean to fully masked data, and the reverse process p(z_s mid z_t) learns to undo it. Moreover, the reverse process defines a latent variable model p(x, z_1, dots, z_T) in which T intermediate partially masked samples z_1,…,z_T are latent variables. Generating from the reverse process p(z_s mid z_t) is the same as performing ancestral sampling from this model.

We can also look at the likelihood log p(x) of the model p to assess its quality. In latent variable models this is intractable, so we resort to approximations via variational inference. For a masked diffusion language model, the evidence lower bound (ELBO) used to approximate the likelihood has a surprisingly simple form (assuming for simplicity alpha_t = 1-t) :

Let’s unpack this formula. The inner term log p_theta(x mid z_t) is the likelihood of a clean sequence x given a partially masked sequence z_t sampled from the forward process. In other words, it is the cross-entropy loss between the predictions of our unmasking transformer and the true tokens — this is exactly the BERT loss!

Differently from BERT, this loss is averaged over all t, and hence over all possible masking rates, rather than a single fixed one. It is also normalized by t, the expected fraction of tokens that are masked (since alpha_t = 1-t); this factor ensures that each BERT loss is normalized for the number of tokens over which the loss is taken.

In summary, MDLM is very similar to BERT, with two key differences:

Most interestingly, the evidence lower bound enables a principled comparison between autoregressive and diffusion language models using log-likelihood (or, equivalently, perplexity) — the standard metric for evaluating language models. While for a long time there was a substantial gap in perplexity between diffusion and autoregressive language models, simplified masked diffusion models were among the first to close much of this gap .

As defined above, masked diffusion models are helpful for building intuition, but they are not production-ready: they generate only fixed-length sequences, they do not support iterative refinement (error correction) out of the box, and they are not especially fast without additional post-training. The rest of this article explores extensions that address these limitations, in the context of modern open-weights diffusion models.

The first issue that arises with standard MDLMs is their limitation to generating fixed-length sequences. Block diffusion addresses this limitation by performing diffusion over blocks, conditioned on previously generated tokens . These blocks can be of arbitrary size, ranging from a dozen to thousands of tokens, and should ideally depend on the application domain.

For example, in biological applications, we might have prior knowledge about the length of the interactions we want to capture, and set the block size to the minimum length needed to capture them. In language modeling, we may instead be interested in maximizing GPU utilization; in that case we would choose the block size so that the arithmetic intensity of our forward pass (which also depends on the batch size) matches that of the underlying hardware.

Additionally, block diffusion naturally supports KV caching, a technique that accelerates sequence generation in autoregressive models. Once a block has been generated using a transformer architecture, its keys and values can be cached and reused when generating future blocks.

Other approaches to variable-length generation rely on connections between masked diffusion models and any-order autoregressive models. For instance, Set Diffusion extends block diffusion to operate over arbitrary sets of positions rather than left-to-right blocks. Other approaches, such as Edit Flows or FlexMDM instead model generation as a sequence of insertion, deletion, and substitution operations, which lets the model grow or shrink the sequence as it refines it.

Standard masked diffusion models are effectively encoder-only (like BERT), in contrast to decoder-only autoregressive models (like GPT). Using an encoder-only architecture requires sampling algorithms that invoke the full network at every denoising step, which can incur a relatively high computational cost.

A key insight is that diffusion performs two kinds of computation: (1) computing a representation of the tokens that have been generated so far, and (2) denoising the corrupted

View original article