Accepted to EMNLP 2026 (Main Conference).
Official code for "Follow the Latent Roadmap: Navigating Revocable Decoding for Diffusion LLMs with Anchor Tokens" (arXiv:2606.16847).
ASRD (Anchor Supervised Revocable Decoding) is a training-free decoding procedure for masked diffusion LLMs. Existing revocable decoders verify and generate inside a single mixed-quality context, which causes two failures:
- Error propagation — freshly generated tokens attend to unconfirmed neighbours and absorb their errors;
- Local error reinforcement — wrong tokens agree with each other, forming a locally self-consistent but globally incorrect cluster that evades verification.
ASRD separates that context into trusted anchors and uncertain candidates, then uses the anchors to supervise both halves of every step. All intervention happens in the embedding space, so FlashAttention still applies and no attention score is ever modified.
| Component | Paper | What it does |
|---|---|---|
| Anchor Tokens Cache (ATC) | §4.1 | A fixed-size FIFO of positions whose top-1 prediction stayed identical for k consecutive steps after being committed. |
| Anchor-Guided Generation | §4.2 | Replaces each [MASK] embedding with an entropy-weighted blend toward the anchor centroid, e = (1-γ)·e_mask + γ·c̄ with γ = α·Ē. Counters error propagation. |
| Anchor-Perturbed Verification | §4.3 | Adds an orthogonal anchor-derived probe to committed-but-pending embeddings, e ← e + β·d, and remasks those whose top-1 flips. Breaks local error reinforcement. |
Drafting and verification share one forward pass per step, so the step count falls far below the sequence length — that is where the speedup comes from.
Reported in the paper (Table 1), sequence length 512, semi-AR block size 32. Acc is accuracy / pass@1, Steps is forward passes per sequence, Speed is wall-clock speedup over the standard decoder.
| Model | Method | HumanEval | MBPP | GSM8K | MATH500 |
|---|---|---|---|---|---|
| LLaDA-Ins-8B | baseline | 43.9 / 512 / 1.0× | 38.2 / 512 / 1.0× | 80.3 / 512 / 1.0× | 37.8 / 512 / 1.0× |
| WINO | 43.9 / 133.4 / 2.8× | 37.4 / 133.3 / 2.4× | 78.7 / 83.1 / 3.6× | 35.6 / 120.6 / 2.6× | |
| Saber | 44.5 / 198.3 / 2.3× | 37.8 / 188.9 / 2.6× | 79.2 / 141.7 / 3.2× | 34.8 / 192.5 / 2.5× | |
| ASRD | 48.8 / 127.8 / 3.6× | 39.4 / 110.4 / 2.9× | 81.1 / 88.5 / 4.3× | 38.8 / 119.3 / 3.8× | |
| Dream-Ins-7B | baseline | 56.1 / 512 / 1.0× | 56.2 / 512 / 1.0× | 80.2 / 512 / 1.0× | 38.6 / 512 / 1.0× |
| WINO | 51.3 / 102.5 / 3.9× | 55.4 / 86.1 / 5.2× | 79.5 / 92.7 / 5.5× | 39.8 / 139.9 / 2.8× | |
| Saber | 53.1 / 264.6 / 1.9× | 55.8 / 388.2 / 1.3× | 79.9 / 314.5 / 1.6× | 41.4 / 250.7 / 2.0× | |
| ASRD | 56.7 / 85.3 / 5.5× | 57.2 / 74.5 / 7.2× | 81.1 / 78.2 / 6.8× | 45.0 / 94.8 / 5.4× |
Base and RL-tuned backbones (Dream-Base-7B, LLaDA-Base-8B, LLaDA-1.5-8B) are in Appendix B.1 of the paper.
Four self-contained sub-projects, mirroring the paper's experiments:
| Folder | Models | Paper |
|---|---|---|
lladainstruct_1p5/ |
LLaDA-8B-Instruct, LLaDA-1.5-8B | Table 1, Table 6 |
lladabase/ |
LLaDA-8B-Base | Table 6 |
dreaminstruct/ |
Dream-7B-Instruct | Table 1 |
dreambase/ |
Dream-Base-7B | Table 6 |
Each ships its own copy of the evaluation harness (a fork of lm-eval) plus the
task definitions and scoring scripts for the four benchmarks. The decoding
algorithm lives in a single module that is kept byte-identical across all
four sub-projects:
lladainstruct_1p5/eval/models/asrd_decoding.py <- reference copy
lladabase/eval/generate_asrd.py
dream{base,instruct}/need2copy/asrd_decoding.py
Dream additionally needs need2copy/generation_utils_asrd.py, which reuses those
helpers and adds the Dream-specific loop (logits are shifted by one position and
the forward pass takes position_ids).
The code targets the standard LLaDA / Dream inference stack:
pip install torch transformers accelerate datasets tqdmThen download the checkpoints you want to evaluate (LLaDA-8B-Instruct, LLaDA-8B-Base, LLaDA-1.5, Dream-v0-Instruct-7B, Dream-v0-Base-7B). All experiments in the paper ran on 8× NVIDIA A100 80GB.
cd lladainstruct_1p5/eval # or lladabase/eval
# edit MODEL_PATH / BASE_OUTPUT_PATH at the top of eval.sh
bash eval.shDream loads its decoder from the checkpoint directory, so the ASRD files have to be installed there first:
cp dreaminstruct/need2copy/{modeling_dream.py,generation_utils_asrd.py,asrd_decoding.py} \
/path/to/Dream-v0-Instruct-7B/
cd dreaminstruct/eval # or dreambase/eval
# edit `model` at the top of eval.sh
bash eval.shEach run prints the average number of forward passes per sequence, which is the Steps column of the tables above.
All of them are passed through --gen_kwargs (LLaDA-Instruct) or --model_args
(everything else) — there are no environment variables to set.
| Name | Meaning | Paper range (App. A.1) | Default here |
|---|---|---|---|
threshold |
τ, unmasking threshold | {0.6, 0.7, 0.8} |
0.7 |
atc_size |
m, ATC capacity | {12, 16, 20} |
16 |
alpha |
α, mask-side guidance strength | [0.05, 0.2] |
0.1 |
beta |
β, pending-side probe strength | [0.1, 0.4] |
0.2 |
consistency_k |
k, temporal-consistency window | 2 |
2 |
block_length |
semi-AR block size | 32 (instruct); full sequence (base) |
32 / unset |
decoding |
asrd or baseline |
— | asrd |
Temperature is 0 throughout for reproducibility. Base models use the standard
full-sequence diffusion sampling strategy; instruct models use semi-AR decoding
with block size 32.
Note. The defaults shipped in
eval.share the mid-points of the ranges published in Appendix A.1, not the per-benchmark values used for the reported numbers. Expect to tune τ, m, α and β per benchmark to match the tables.
decoding=baseline selects the standard decoder, i.e. the 1.0× reference row.
The paper also compares against WINO (arXiv:2507.18578)
and Saber (arXiv:2510.18165). Those are not
redistributed here; use their official repositories with the same decoding
configuration (block size 32, matching τ, temperature=0) for a like-for-like
comparison.
@inproceedings{yao2026asrd,
title = {Follow the Latent Roadmap: Navigating Revocable Decoding for Diffusion LLMs with Anchor Tokens},
author = {Yao, Yizhen and Zhu, Qinglin and Zhao, Runcong and Dai, Xiangxiang and Xiang, Yanzheng and He, Yulan and Gui, Lin},
booktitle = {Proceedings of the 2026 Conference on Empirical Methods in Natural Language Processing (EMNLP)},
year = {2026}
}MIT, see LICENSE.