Deepseek has released V4.1-Flash, an open-weights model that cuts the KV cache in fast GPU memory to roughly a quarter of the space used by predecessor Deepseek-V4-Flash. The 552-billion-parameter model activates 8 billion parameters per token on input versus 16 billion on output, nearly halving input compute. It matches leading closed models from OpenAI and Anthropic on some coding benchmarks but lags on complex scientific tasks and image analysis.
No score is assigned. Sources and their independence are shown in the citation chain below.