EB-Decode: Accelerating Diffusion LLMs with Learnable Block Sizes and Parallel Sampling

LLaDA-8B-Instruct answering a GSM8K question, decoded two ways on an NVIDIA H100. Vanilla decoding needs 512 steps and 19.05 s; EB-Decode reaches the same answer in 33 steps and 1.34 s, a 14.2× speedup.

Diffusion large language models (dLLMs) predict many tokens in a single forward pass, but those predictions do not all become final output at the same time. Under the common block-wise decoding recipe, the model walks through fixed-length blocks, commits the predictions inside the active block that satisfy a confidence threshold, and leaves the rest masked for further refinement. Because both the block boundaries and the threshold are fixed in advance, a large number of predictions that have already stabilized still sit through additional denoising steps before they are accepted.

We propose Early-Bird Decoding (EB-Decode), which uses two lightweight routers to learn the two decisions that block-wise decoding makes by hand: which tokens should be processed together, and which predictions can be finalized early. A learnable block size module (LBS) groups tokens of similar difficulty, and a position-aware learnable parallel sampling module (LPS) decides when those tokens can be accepted. Both modules reuse the forward pass of the base model and require no change to the pretrained weights. Full details are in our preprint, EB-Decode (arXiv:2609.16450).

Conceptual comparison between fixed block-wise decoding and EB-Decode

Fig.1 - Conceptual comparison between fixed block-wise decoding and EB-Decode. Fixed blocks may mix tokens of different difficulty and prevent correct predictions from being accepted across block boundaries. EB-Decode reduces this waiting through adaptive grouping and early acceptance.

Observation 1: token difficulty follows a staircase pattern. We start by measuring prediction entropy across positions, which reflects how uncertain the model is among the candidate tokens at each masked position. Along the denoising trajectory of a GSM8K sample, low-entropy positions form long plateaus with a few high-entropy positions scattered between them. These plateaus typically correspond to content that is easy to predict, such as the routine connective phrasing in a math solution or the structural scaffolding in code.

Fixed boundaries do not respect this distribution. A stretch of easy content can be split across several blocks and lose the opportunity to be accepted jointly, while a handful of hard tokens can stall the active block so that easier content behind them cannot advance. Allowing non-contiguous grouping lets the model settle the context surrounding a hard position first, then resolve that position with the additional right-side information it has gained.

Prediction entropy heatmaps under fixed and learnable block sizes

Fig.2 - Prediction entropy heatmaps under a fixed block size and under our non-contiguous learnable block size. Gray regions denote tokens that have already been committed. Learnable blocks organize the active positions according to the difficulty pattern, which reduces the stalls caused by fixed boundaries.

Observation 2: predictions stabilize well before confidence catches up. Even with sensible block boundaries, a second kind of waiting happens inside the block. Tracking the per-step prediction at each position shows that many tokens already hold the correct value yet keep participating in denoising because their confidence has not reached the threshold. A single confidence score cannot separate two situations: a prediction that is already correct but held with low confidence, and a prediction that genuinely still needs revision.

The figure below illustrates the gap between the step at which a prediction becomes correct and the step at which it is finally accepted. On GSM8K, vanilla decoding and confidence-based decoding wait an average of 13.4 and 2.3 denoising steps respectively, while EB-Decode shortens the gap to 0.7 steps. This is the early-bird phenomenon we want to exploit: identify predictions that have stabilized ahead of schedule and remove the computation that follows them.

Per-token prediction trajectories showing acceptance delay

Fig.3 - Per-token prediction trajectories across denoising steps. Several positions already match the ground truth many steps before the decoder unmasks them, so the computation spent between those two moments is wasted.

LBS: learning which tokens should be decoded together. To match the staircase structure of token difficulty, Learnable Block Size (LBS) predicts for each masked position whether it belongs in the current block, so the block size follows from the selected positions rather than being set in advance. The router takes two inputs: the normalized prediction entropy, and the embedding of the current top-1 token. Entropy supplies the uncertainty signal, and the semantic embedding helps distinguish different sources of that uncertainty, for instance a choice among interchangeable function words versus a content word whose identity depends on context that has not been resolved.

During training, LBS uses the per-token cross-entropy of the frozen model to learn to prefer positions that are currently easy to predict, and a regularization term that rewards larger selections keeps it from collapsing to an empty block. At inference, the candidate set passes through three steps: EOS predictions are removed so that generation is not cut short; the maximum gap between neighboring selected positions is capped to preserve local coherence; and when too few positions survive, the router falls back to a small left-to-right block. The resulting block may contain non-contiguous positions, but it will not reach arbitrarily far across regions that remain unresolved.

LPS: learning which predictions can be finalized early. LBS decides which positions are active, and Learnable Parallel Sampling (LPS) decides which of them can be accepted. LPS uses a small Transformer that processes the positions in a block jointly, with five inputs per position: the top-1 confidence, the prediction entropy, the probability gap between the top-1 and top-2 candidates, the block-wise mask ratio, and the relative position inside the block. These describe the uncertainty of a prediction, the competition among candidates, the current decoding progress, and positional relationships, which together let the acceptance decision adapt to blocks of different lengths.

LPS is trained by generate-then-replay and needs no human annotation. The model first produces a complete sequence, and the label at each intermediate step is whether the current prediction already matches that final sequence. What LPS learns, therefore, is when a prediction has converged to the model’s own output, not whether that output is correct in an absolute sense. Because a false acceptance locks in a wrong token irreversibly, training assigns higher weight to negative samples.

At inference, LPS keeps the original confidence path intact: a token is accepted if it reaches the base model’s confidence threshold or if its LPS score exceeds a threshold calibrated on a validation split. Once the active block is resolved, LBS selects the next one. If no position qualifies at a given step, the highest-confidence candidate is accepted so that decoding does not stall.

EB-Decode inference pipeline and training procedure

Fig.4 - The EB-Decode inference pipeline and training procedure. LBS first selects and filters the active block, and LPS then decides which predictions inside it can be accepted. The two routers are trained separately, and the base dLLM stays frozen throughout.

Results: block selection and early acceptance contribute in complementary ways. We evaluate EB-Decode on LLaDA-8B-Instruct, Dream-v0-Instruct-7B, and LLaDA-1.5, covering mathematical reasoning on GSM8K and MATH500 and code generation on HumanEval and MBPP. The main experiments use four H100 GPUs, the Hugging Face Transformers backend, batch size 1, and generation length 512. No method uses a KV cache, so the comparison reflects the decoding strategy itself.

Under these settings, full EB-Decode reaches 3.53× to 18.76× the throughput of the vanilla decoder and up to 1.58× that of Fast-dLLM, with task accuracy close to the vanilla decoder throughout. The largest speedup comes from Dream on MBPP, where throughput rises from 2.23 to 41.83 TPS and accuracy moves from 53.40% to 53.60%.

Comparison of EB-Decode and baseline methods across four benchmarks and three base models

Table 1 - Comparison of EB-Decode and baseline methods across four benchmarks and three base models, reporting tokens per second per GPU (TPS), speedup over vanilla decoding (Sp.up), and task accuracy. The best throughput and speedup in each column are in bold. "—" marks the settings where Learn2PD has no open-source implementation for LLaDA-1.5.

Taking LLaDA-8B-Instruct on GSM8K-CoT as an example, LBS alone already improves the efficiency of block-wise decoding, lifting throughput from the vanilla decoder’s 17.04 TPS to 86.80 TPS at 79.68% accuracy, ahead of Fast-dLLM on both counts. Adding LPS raises throughput further to 104.55 TPS, a 6.14× speedup over vanilla, with accuracy at 78.24%.

Deployment: lightweight routers that compose with caching. LBS and LPS contain roughly 0.6M and 0.03M trainable parameters. Measured on an A100, their combined routing overhead is about 6.21% of a single base-model forward pass. The speedup therefore comes from taking fewer denoising steps rather than from making each forward pass cheaper.

EB-Decode also composes with KV caching. For non-contiguous blocks, the cache reuses keys and values outside the leftmost and rightmost selected positions. In a separate experiment on four A100 GPUs, dual caching raises EB-Decode throughput from 65.27 to 78.43 TPS, an additional 1.20×, and EB-Decode still reaches about 1.38× the throughput of Fast-dLLM under the same dual-cache setting.

Our current evaluation covers generation lengths of 256 and 512 tokens and inference at batch size 1. Longer outputs, continuous batching, and system-level optimization remain to be verified.

Summary. EB-Decode starts from two properties of the denoising process: how prediction difficulty is distributed across positions, and how early predictions stabilize in time. LBS groups positions of similar difficulty, and LPS removes the repeated deferral of predictions that have already settled. With these two lightweight modules, a frozen diffusion language model finishes generation in fewer denoising steps and converts more of its latent parallelism into measured throughput.