TSRGet The Daily Report
The Singularity Report

TSR Desk · mathematics · 1 October 2026, 01:00 UTC

Draft in Parallel, Condition Through Depth: Adjacent Causal Injection for Speculative Decoding

What
Draft in Parallel, Condition Through Depth: Adjacent Causal Injection for Speculative Decoding
Who
arxiv.org
When
30 September 2026, 04:00 UTC
Category
Mathematics
Primary source
https://arxiv.org/abs/2609.36173
What is not known
This brief does not claim independent replication. Claims that appear only on X and not in the primary source stay unknown.

At temperature zero on Qwen3-8B, it raises the seven-benchmark mean from DFlash's 3.77 to 4.82 (+27.8%); in SGLang serving tests, it delivers 23.3% higher throughput than DFlash on average. It comes from a paper posted to arXiv on 30 September 2026. Parallel speculative drafting generates multiple candidates in one backbone pass, but independent token selection can produce inconsistent continuations that shorten the accepted prefix. Existing methods mostly leave conditional decoding to a lightweight module after the backbone, which limits the flow of predecessor information to successors. Our analysis of DFlash shows that early positions already form recoverable predictions in shallow layers, and that accurate adjacent predecessors help successors more when they enter earlier. We therefore propose DSpine, a drafter with causal conditioning injection throughout the backbone: at every layer, gated adjacent injection writes each predecessor's predicted feature into its successor, so the causal conditioning chain unfolds over network depth while all positions update in parallel. A unified transfer space built from the target model's output embeddings unifies layer-wise injection with predecessor-conditioned decoding, and layer-wise output-embedding supervision promotes the formation of predicted features in shallow layers. Fused kernels and a transition cache execute both efficiently in parallel within SGLang. Across seven math, code, and chat benchmarks, DSpine achieves the longest acceptance length at both temperatures on Qwen3-4B and Qwen3-8B.

Why it counts

At temperature zero on Qwen3-8B, it raises the seven-benchmark mean from DFlash's 3.77 to 4.82 (+27.8%); in SGLang serving tests, it delivers 23.3% higher throughput than DFlash on average.

Sources

Primary source: primary source

What is not known

This brief does not claim independent replication. Claims that appear only on X and not in the primary source stay unknown.

No clip. The article still stands.