DeepSeek released DeepSeek V4.1-Flash on 10 September, a new open-weight model in the company's Flash line. The announcement went out on the company's X account and the weights appeared on Hugging Face the same morning. By the end of the day the release had drawn close to a thousand points on Hacker News.
The model is a multimodal mixture-of-experts system with 552 billion backbone parameters and a context window of one million tokens. It handles images and text natively instead of through a bolt-on adapter. DeepSeek says it trained the model from scratch on 45 trillion multimodal tokens.
The main engineering change is a Causal Encoder-Decoder architecture. A 40 layer transformer is split into a 20 layer causal encoder and a 20 layer decoder, and the decoder projects its KV cache from the encoder's final hidden states instead of building one at every layer. DeepSeek says this activates only 8 billion parameters per token during prefill and 16 billion during decode.
The payoff is a much smaller memory footprint. DeepSeek reports a global KV cache of 890 bytes per token, about a quarter of what V4-Flash needed, and a persistent cache that shrinks to roughly one eighth. Lower memory use translates into cheaper serving for input-heavy work, which is where long agent runs spend most of their budget. The full specification is in the
on Hugging Face.
DeepSeek also reports strong agentic results. V4.1-Flash scores 90.6 on Terminal-Bench 2.1, ahead of Opus 5 at 89.1 and GPT-5.6 Sol at 88.8, and it leads on DeepSWE v1.1, CyberGym and AutomationBench. Opus 5 still wins the harder Terminal-Bench 3.0 and 4.0 runs. These are vendor reported numbers and have not been independently verified.
The weights are MIT licensed, and OpenRouter lists the model at $0.15 per million input tokens and $0.60 per million output tokens. For founders building agents, the practical point is the combination of a permissive licence, a one million token context and a smaller serving footprint. A model that fits cheaper hardware changes what a fixed token budget can buy, before any benchmark gain is taken into account.