Skip to content

Latest commit

 

History

3 Commits

Folders and files

Repository files navigation

ASRD

Accepted to EMNLP 2026 (Main Conference).

Official code for "Follow the Latent Roadmap: Navigating Revocable Decoding for Diffusion LLMs with Anchor Tokens" (arXiv:2606.16847).

ASRD (Anchor Supervised Revocable Decoding) is a training-free decoding procedure for masked diffusion LLMs. Existing revocable decoders verify and generate inside a single mixed-quality context, which causes two failures:

  • Error propagation — freshly generated tokens attend to unconfirmed neighbours and absorb their errors;
  • Local error reinforcement — wrong tokens agree with each other, forming a locally self-consistent but globally incorrect cluster that evades verification.

ASRD separates that context into trusted anchors and uncertain candidates, then uses the anchors to supervise both halves of every step. All intervention happens in the embedding space, so FlashAttention still applies and no attention score is ever modified.

Method

Component Paper What it does
Anchor Tokens Cache (ATC) §4.1 A fixed-size FIFO of positions whose top-1 prediction stayed identical for k consecutive steps after being committed.
Anchor-Guided Generation §4.2 Replaces each [MASK] embedding with an entropy-weighted blend toward the anchor centroid, e = (1-γ)·e_mask + γ·c̄ with γ = α·Ē. Counters error propagation.
Anchor-Perturbed Verification §4.3 Adds an orthogonal anchor-derived probe to committed-but-pending embeddings, e ← e + β·d, and remasks those whose top-1 flips. Breaks local error reinforcement.

Drafting and verification share one forward pass per step, so the step count falls far below the sequence length — that is where the speedup comes from.

Results

Reported in the paper (Table 1), sequence length 512, semi-AR block size 32. Acc is accuracy / pass@1, Steps is forward passes per sequence, Speed is wall-clock speedup over the standard decoder.

Model Method HumanEval MBPP GSM8K MATH500
LLaDA-Ins-8B baseline 43.9 / 512 / 1.0× 38.2 / 512 / 1.0× 80.3 / 512 / 1.0× 37.8 / 512 / 1.0×
WINO 43.9 / 133.4 / 2.8× 37.4 / 133.3 / 2.4× 78.7 / 83.1 / 3.6× 35.6 / 120.6 / 2.6×
Saber 44.5 / 198.3 / 2.3× 37.8 / 188.9 / 2.6× 79.2 / 141.7 / 3.2× 34.8 / 192.5 / 2.5×
ASRD 48.8 / 127.8 / 3.6× 39.4 / 110.4 / 2.9× 81.1 / 88.5 / 4.3× 38.8 / 119.3 / 3.8×
Dream-Ins-7B baseline 56.1 / 512 / 1.0× 56.2 / 512 / 1.0× 80.2 / 512 / 1.0× 38.6 / 512 / 1.0×
WINO 51.3 / 102.5 / 3.9× 55.4 / 86.1 / 5.2× 79.5 / 92.7 / 5.5× 39.8 / 139.9 / 2.8×
Saber 53.1 / 264.6 / 1.9× 55.8 / 388.2 / 1.3× 79.9 / 314.5 / 1.6× 41.4 / 250.7 / 2.0×
ASRD 56.7 / 85.3 / 5.5× 57.2 / 74.5 / 7.2× 81.1 / 78.2 / 6.8× 45.0 / 94.8 / 5.4×

Base and RL-tuned backbones (Dream-Base-7B, LLaDA-Base-8B, LLaDA-1.5-8B) are in Appendix B.1 of the paper.

Code structure

Four self-contained sub-projects, mirroring the paper's experiments:

Folder Models Paper
lladainstruct_1p5/ LLaDA-8B-Instruct, LLaDA-1.5-8B Table 1, Table 6
lladabase/ LLaDA-8B-Base Table 6
dreaminstruct/ Dream-7B-Instruct Table 1
dreambase/ Dream-Base-7B Table 6

Each ships its own copy of the evaluation harness (a fork of lm-eval) plus the task definitions and scoring scripts for the four benchmarks. The decoding algorithm lives in a single module that is kept byte-identical across all four sub-projects:

lladainstruct_1p5/eval/models/asrd_decoding.py   <- reference copy
lladabase/eval/generate_asrd.py
dream{base,instruct}/need2copy/asrd_decoding.py

Dream additionally needs need2copy/generation_utils_asrd.py, which reuses those helpers and adds the Dream-specific loop (logits are shifted by one position and the forward pass takes position_ids).

Setup

The code targets the standard LLaDA / Dream inference stack:

pip install torch transformers accelerate datasets tqdm

Then download the checkpoints you want to evaluate (LLaDA-8B-Instruct, LLaDA-8B-Base, LLaDA-1.5, Dream-v0-Instruct-7B, Dream-v0-Base-7B). All experiments in the paper ran on 8× NVIDIA A100 80GB.

Running

LLaDA (lladainstruct_1p5, lladabase)

cd lladainstruct_1p5/eval    # or lladabase/eval
# edit MODEL_PATH / BASE_OUTPUT_PATH at the top of eval.sh
bash eval.sh

Dream (dreaminstruct, dreambase)

Dream loads its decoder from the checkpoint directory, so the ASRD files have to be installed there first:

cp dreaminstruct/need2copy/{modeling_dream.py,generation_utils_asrd.py,asrd_decoding.py} \
   /path/to/Dream-v0-Instruct-7B/

cd dreaminstruct/eval        # or dreambase/eval
# edit `model` at the top of eval.sh
bash eval.sh

Each run prints the average number of forward passes per sequence, which is the Steps column of the tables above.

Hyper-parameters

All of them are passed through --gen_kwargs (LLaDA-Instruct) or --model_args (everything else) — there are no environment variables to set.

Name Meaning Paper range (App. A.1) Default here
threshold τ, unmasking threshold {0.6, 0.7, 0.8} 0.7
atc_size m, ATC capacity {12, 16, 20} 16
alpha α, mask-side guidance strength [0.05, 0.2] 0.1
beta β, pending-side probe strength [0.1, 0.4] 0.2
consistency_k k, temporal-consistency window 2 2
block_length semi-AR block size 32 (instruct); full sequence (base) 32 / unset
decoding asrd or baseline — asrd

Temperature is 0 throughout for reproducibility. Base models use the standard full-sequence diffusion sampling strategy; instruct models use semi-AR decoding with block size 32.

Note. The defaults shipped in eval.sh are the mid-points of the ranges published in Appendix A.1, not the per-benchmark values used for the reported numbers. Expect to tune τ, m, α and β per benchmark to match the tables.

Baselines

decoding=baseline selects the standard decoder, i.e. the 1.0× reference row.

The paper also compares against WINO (arXiv:2507.18578) and Saber (arXiv:2510.18165). Those are not redistributed here; use their official repositories with the same decoding configuration (block size 32, matching τ, temperature=0) for a like-for-like comparison.

Citation

@inproceedings{yao2026asrd,
  title     = {Follow the Latent Roadmap: Navigating Revocable Decoding for Diffusion LLMs with Anchor Tokens},
  author    = {Yao, Yizhen and Zhu, Qinglin and Zhao, Runcong and Dai, Xiangxiang and Xiang, Yanzheng and He, Yulan and Gui, Lin},
  booktitle = {Proceedings of the 2026 Conference on Empirical Methods in Natural Language Processing (EMNLP)},
  year      = {2026}
}

License

MIT, see LICENSE.

About

Follow the Latent Roadmap: Navigating Revocable Decoding for Diffusion LLMs with Anchor Tokens

Resources

Stars

2 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages