← Back to the wire

New Deepseek model V4.1-Flash cuts memory needs for AI agents

AnnouncementModelSep 10, 2026

Deepseek has released V4.1-Flash, an open-weights model that cuts the KV cache in fast GPU memory to roughly a quarter of the space used by predecessor Deepseek-V4-Flash. The 552-billion-parameter model activates 8 billion parameters per token on input versus 16 billion on output, nearly halving input compute. It matches leading closed models from OpenAI and Anthropic on some coding benchmarks but lags on complex scientific tasks and image analysis.

Receipt № 18431 source · awaiting confirmation ◐

Evidence

1source· awaiting independent confirmation

No score is assigned. Sources and their independence are shown in the citation chain below.

Citation chain · 1 source

OpenAICompanyAnthropicCompanyDeepseekCompanyV4.1-FlashModelDeepseek-V4-FlashModel
Canonical: https://the-decoder.com/new-deepseek-model-v4-1-flash-cuts-memory-needs-for-ai-agents/