THE FUTURELESS

RADAR ·

DeepSeek's V4.1-Flash cuts GPU context cache to a quarter of its predecessor

DeepSeek has released V4.1-Flash, a multimodal model with 552 billion parameters. According to the technical report, the context cache the model keeps in fast GPU memory takes about a quarter of the space its predecessor V4-Flash needed; the portion offloaded to disk shrinks to roughly an eighth. The model activates 8 billion parameters per token while reading input and 16 billion while generating text, with a context limit of one million tokens. The cache is stored in FP4 rather than FP8. On the DeepSWE v1.1 software benchmark it scores 74.2 percent, narrowly ahead of Anthropic's Opus 5 and OpenAI's GPT-5.6 Sol, while trailing on hard scientific tasks and image reading. The files are on Hugging Face under the MIT license, and the model is offered through an API at the same prices as its predecessor.

Source: The Decoder

← Back to the radar