Microsoft introduced Differential Transformer V2, an improved version of Differential Transformer, focusing on inference efficiency, training stability for production-level LLMs, and architectural elegance. Authors Tianzhu Ye, Li Dong, Yutao Sun, and Furu Wei detail key changes including faster decoding without custom kernels and a modified differential operation. Differential Transformer V2 doubles query heads while maintaining key-value heads, achieving decoding speeds comparable to standard Transformer. Pretraining experiments on production-scale models are ongoing.
No score is assigned. Sources and their independence are shown in the citation chain below.