Skip to content

Latest commit

 

History

6 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 

Repository files navigation

Hierarchical Continuous Diffusion Language Models

Project page arXiv Hugging Face Paper Artifacts

Augmenting Continuous Diffusion Language Models with Discrete Token Guidance

Hui Ren1, Zihan Li1, Chang Liu1, Huidong Liu2, Alexander Schwing1
1University of Illinois Urbana-Champaign    2Amazon.com, Inc.

HC-DLM augments continuous diffusion language models with discrete token guidance:
a shared latent plans every token jointly, while tokens read out at each step keep it anchored to valid text.

One simulated HC-DLM reverse trajectory on a sentence: the continuous latent is denoised step by step; at every step a token draft is read out of it and fed back to guide the next latent update, so early words can still be revised.

The latent plans, the tokens guide, at every step (simulated example). Interactive version on the project page.

πŸ’‘ Idea

Each family of diffusion language models has a blind spot:

  • Discrete diffusion decodes tokens in parallel, but samples each one independently from its marginal.
  • Continuous diffusion plans all tokens in one shared latent, but never checks that plan against real tokens until the very end.

HC-DLM keeps both strengths: tokens are planned jointly in the latent, and the latent is guided by tokens at every step. The continuous latent is the only persistent generative state; at every step the model reads a token draft $k_t$ out of the latent $x_t$, and that draft guides the next latent update.

$$ p_{\theta,\phi}(k_{0:T},x_{0:T})=p(x_T),p_\theta(k_0\mid x_0)\prod_{t=1}^{T}\underbrace{p_\theta(k_t\mid x_t)}_{\text{read out tokens}};\underbrace{p_\phi(x_{t-1}\mid x_t,k_t)}_{\text{token-guided latent denoising}} $$

Four reverse-step diagrams. (a) Discrete diffusion: tokens k_t go to k_{t-1} through a learned predictor, with no latent. (b) Continuous diffusion: the latent x_t goes to x_{t-1}; tokens are decoded only at t = 0. (c) Hybrid diffusion: a discrete chain and a continuous chain run side by side and condition each other. (d) HC-DLM: tokens are read out of x_t and condition the latent update p_phi(x_{t-1} | x_t, k_t).

One reverse step, four designs. Discrete diffusion updates tokens directly, one marginal at a time. Continuous diffusion denoises a latent that is blind to tokens and decodes only at the end. Hybrid models attach a continuous signal to a self-contained discrete chain. In HC-DLM the two levels talk at every step: tokens are read out of the latent, then guide its next update.

Discrete diffusion
e.g. MDM, LLaDA
Continuous diffusion
e.g. Diffusion-LM, Plaid
Hybrid discrete–continuous
e.g. CADD, CCDD
HC-DLM (ours)
Token dependence within a step βœ•
independent marginals
βœ“ ◐
via conditioning only
βœ“
Tied to tokens at every step βœ“ βœ•
only at t = 0
βœ“ βœ“
Tokens revisable at every step ◐
uniform kernel only
βœ“ ◐
uniform kernel only
βœ“
readout from xt

Noise corrupts tokens and latent independently, so training stays simple. A single variational bound on the token likelihood splits into three terms and trains everything end to end: reconstruction for the token readout, token-guided flow matching for the latent denoiser, and an entropy term that keeps the encoder from collapsing.

Training pipeline: an encoder maps clean tokens to a latent, independent forward kernels produce the noisy pair, the token-conditioned denoiser predicts the clean latent, and the token predictor decodes the latent back to tokens.

One principled objective, trained end to end. An encoder maps clean tokens to a latent, independent forward kernels produce the noisy pair, the token-conditioned denoiser predicts the clean latent, and the token predictor decodes it back to tokens.

πŸ“Š Results

From Sudoku and Countdown to open-domain text, HC-DLM leads discrete, continuous and hybrid diffusion baselines of the same size on Hard Sudoku, Countdown and LM1B. Parameter counts exclude token embeddings.

Hard Sudoku accuracy 72.41% at 6M parameters, 1.68 over CCDD and 22.53 over masked diffusion; Countdown CD5 accuracy 37.52% at 6M, 12.17 over CCDD; LM1B generative perplexity 75.5, 1.8 lower than Plaid and 16.7 lower than LangFlow.

Two line charts. Left: Sudoku accuracy of the decoded intermediate prediction along the trajectory rises earlier and higher for HC-DLM than for a latent diffusion model without token guidance. Right: Hard Sudoku accuracy against the number of denoising steps stays high for HC-DLM while MDM degrades as steps are reduced.

Left: solutions emerge early: accuracy of the intermediate prediction decoded at each step, vs. latent diffusion without token guidance. Right: holds up with fewer steps: Hard Sudoku accuracy against step count; MDM degrades sharply as steps shrink.

πŸ“‹ Full tables (accuracy %; generative perplexity, lower is better)

Sudoku & Countdown (6M parameters, accuracy %)

Method Sudoku Easy Sudoku Hard CD4 CD5
MDM (top-prob. margin) 89.49 49.88 50.8 21.3
CCDD 94.65 70.73 81.18 25.35
HC-DLM 94.21 72.41 84.41 37.52

LM1B (generative perplexity)

Method Params Gen. PPL ↓
MDM 116M 103.9
Duo 116M 97.6
LangFlow 117M 92.2
Plaid 109M 77.3
HC-DLM 118M 75.5

Ablation: neither level works alone. Latent DM removes token guidance from the denoiser; MDM removes the latent altogether (Sudoku accuracy %).

Method Cont. latent Token guidance Easy Hard
MDM (top-prob. margin) βœ• βœ• 89.49 49.88
Latent DM (w/o token guidance) βœ“ βœ• 50.46 24.74
HC-DLM βœ“ βœ“ 94.21 72.41

πŸ“¦ Release

The code and the artifacts are coming soon. Stay tuned!

πŸ“– Citation

If you find this work useful, please consider citing:

@misc{ren2026hcdlm,
      title={Hierarchical Continuous Diffusion Language Models}, 
      author={Hui Ren and Zihan Li and Chang Liu and Huidong Liu and Alexander Schwing},
      year={2026},
      eprint={2610.02193},
      archivePrefix={arXiv},
      primaryClass={cs.CL},
      url={https://arxiv.org/abs/2610.02193}, 
}

About

Official implementation of "Hierarchical Continuous Diffusion Language Models"

Topics

Resources

Stars

56 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors