DeepSeek has launched a limited internal beta version of the DeepSeek V4.1 Flash model. This intermediate checkpoint is built on a new architecture that natively supports multimodality.
DeepSeek positions this release as faster, more powerful, and more cost-effective compared to previous Flash versions. During a strictly limited testing period, the pricing remains at the Flash level. Access is available until the model identifier expires.
Developers can use the preview without changing the base URL of their existing API. The model identifier is deepseek-v4.1-flash-expires-on-0910. Payments during the beta test correspond to DeepSeek V4 Flash, but each account is limited to 20 concurrent requests, which is significantly lower than the production throughput of Flash. Thus, this period is clearly intended for evaluation, not active traffic.
The date specified in the model name signals that this identifier will stop working around September 10, 2026, leaving about two days of access from the launch on September 8.
From an architectural standpoint, V4.1 Flash differs from the early experimental V4 Flash Vision, which attached a separate extension package for image processing to the text model. DeepSeek asserts that multimodal input is now an integral part of the redesigned architecture. This means that text and images, as well as other related modalities, are processed within a single stack instead of passing through an external encoder path followed by data insertion.
For the latency-optimized product line, this change is less significant as a statement of peak visual performance and more as an attempt to avoid the additional latency often introduced by plugin-style vision when images must be converted before entering the language sequence.
Early developer reports gathered during the beta test indicate stable generation in the range of approximately 300 to 500 tokens per second at low parallelism. Generation is described as relatively stable for iterative coding and UI prototyping cycles, where frequent retries are more important than one long completed sequence. It should be noted that these figures are community observations, not official benchmarks; DeepSeek has not published the number of parameters, context length, or formal evaluation tables for this interim build.
A question was included in the feedback form accompanying the test asking users whether they believe the model can replace DeepSeek V4 Pro in an online environment. This suggests that the company is exploring the possibility of achieving Pro-level utility at Flash speed and cost, without abandoning the Flash cost structure.
Collectively, the time-limited identifier, the parallelism cap, and the query about replacing the Pro version present V4.1 Flash as a temporary multimodal preview of the new Flash architecture, rather than a permanent SKU. Developers who require this functionality should test the model before its identifier expires; those planning workloads will require a subsequent release with higher parallelism, documented limitations, and a fixed name before considering the new architecture standard infrastructure.

