NVIDIA published the third post in its AI model co-design series, explaining how speculative decoding accelerates LLM inference while preserving output accuracy. The post defines draft length and acceptance length, presents a speedup formula, and offers guidelines for selecting draft length, such as setting it based on group size when attention dominates decode time. It also notes that larger draft lengths help linear layers reach compute-bound performance at lower effective batch sizes.
No score is assigned. Sources and their independence are shown in the citation chain below.