DeepSeek released DeepSeek V4.1-Flash on 10 September, a new open-weight model in the company's Flash line. The announcement went out on the company's X account and the weights appeared on Hugging Face the same morning. By the end of the day the release had drawn close to a thousand points on Hacker News.

The model is a multimodal mixture-of-experts system with 552 billion backbone parameters and a context window of one million tokens. It handles images and text natively instead of through a bolt-on adapter. DeepSeek says it trained the model from scratch on 45 trillion multimodal tokens.

The main engineering change is a Causal Encoder-Decoder architecture. A 40 layer transformer is split into a 20 layer causal encoder and a 20 layer decoder, and the decoder projects its KV cache from the encoder's final hidden states instead of building one at every layer. DeepSeek says this activates only 8 billion parameters per token during prefill and 16 billion during decode.

The payoff is a much smaller memory footprint. DeepSeek reports a global KV cache of 890 bytes per token, about a quarter of what V4-Flash needed, and a persistent cache that shrinks to roughly one eighth. Lower memory use translates into cheaper serving for input-heavy work, which is where long agent runs spend most of their budget. The full specification is in the

model card

on Hugging Face.

DeepSeek also reports strong agentic results. V4.1-Flash scores 90.6 on Terminal-Bench 2.1, ahead of Opus 5 at 89.1 and GPT-5.6 Sol at 88.8, and it leads on DeepSWE v1.1, CyberGym and AutomationBench. Opus 5 still wins the harder Terminal-Bench 3.0 and 4.0 runs. These are vendor reported numbers and have not been independently verified.

The weights are MIT licensed, and OpenRouter lists the model at $0.15 per million input tokens and $0.60 per million output tokens. For founders building agents, the practical point is the combination of a permissive licence, a one million token context and a smaller serving footprint. A model that fits cheaper hardware changes what a fixed token budget can buy, before any benchmark gain is taken into account.