← Back to the wire

DeepSeek-V4: a million-token context that agents can actually use

AnnouncementModelApr 24, 2026

DeepSeek released two Mixture-of-Experts checkpoints, DeepSeek-V4-Pro and DeepSeek-V4-Flash, both featuring a 1M-token context window. The models use hybrid attention mechanisms—Compressed Sparse Attention and Heavily Compressed Attention—to reduce KV cache memory to roughly 2% of standard architectures. DeepSeek-V4-Pro requires 27% of single-token inference FLOPs compared to its predecessor. Writing for Hugging Face, ben burtenshaw notes the design targets long-running agentic workloads, preserving reasoning traces across tool calls and user message boundaries.

Receipt № 5041 source · awaiting confirmation ◐

Evidence

1source· awaiting independent confirmation

No score is assigned. Sources and their independence are shown in the citation chain below.

Citation chain · 1 source

01highPRIMARY
Hugging FaceCompanyDeepSeekCompanyben burtenshawPersonDeepSeek-V4-ProModelDeepSeek-V4-FlashModel
Canonical: https://huggingface.co/blog/deepseekv4