DeepSeek Adds Vision to V4 Flash Test Model
  • News
  • Asia

DeepSeek Adds Vision to V4 Flash Test Model

Experimental model brings image and screenshot analysis to DeepSeek’s latest AI series

8/24/2026
Ghita Khalfaoui
Back to News

Chinese AI firm DeepSeek has announced the launch of an experimental new model capable of understanding visual inputs. Named DeepSeek-V4-Flash-Vision-Exp, this tool adds multimodal capabilities to its flagship text-based model, positioning it as a direct competitor to leading US technologies. The release underscores the intensifying global rivalry in the artificial intelligence sector, with performance benchmarks nearing those of established players like Anthropic.


Introducing Advanced Multimodal Capabilities

The new release is an experimental extension of DeepSeek’s acclaimed V4 Flash model, which was previously limited to text-only interactions. By integrating multimodal functions, the AI can now analyze and interpret visual prompts, including images and screenshots from users. This enhancement significantly broadens the model's potential applications, moving beyond simple text generation to more complex, visually-grounded tasks.

While DeepSeek has previously developed visual models like its DeepSeek-VL family, this new offering integrates these capabilities into its most advanced series. The company has made DeepSeek-V4-Flash-Vision-Exp accessible to developers and researchers through its dedicated application programming interface (API). This move encourages widespread testing and integration, accelerating its adoption and refinement within the developer community.

Benchmarking Against Industry Leaders

DeepSeek has made bold claims about its new model's performance, stating it is "close to" Anthropic PBC's sophisticated Opus 4.8 model. This comparison is particularly focused on the model's agentic capabilities when handling multimodal tasks. Such benchmarking is crucial for establishing credibility and competitiveness in a market dominated by a few major American developers.

Agentic capabilities refer to an AI's ability to take action and complete objectives with minimal human intervention or continuous prompting. This level of autonomy is a key frontier in AI development, enabling models to function more like independent assistants. By nearing the performance of top-tier models in this area, DeepSeek demonstrates a significant leap in its technological prowess.

The Competitive AI Landscape

This launch occurs amidst a heated technological rivalry between Chinese and US firms at the forefront of artificial intelligence. Chinese developers are increasingly challenging their Western counterparts, often by providing models with similar performance for a fraction of the cost. DeepSeek's strategy appears to follow this trend, aiming to capture market share by offering powerful yet affordable AI solutions.

The company has already made waves in the industry with its previous releases, particularly its flagship text-only model earlier this year. The DeepSeek-V4-Flash reset expectations for what low-cost, open-weight models could achieve in terms of performance and efficiency. This new experimental version builds on that reputation, further solidifying the company's position as a significant innovator.


The unveiling of DeepSeek-V4-Flash-Vision-Exp marks a notable advancement for the Hangzhou-based company and the broader Chinese AI ecosystem. By developing a multimodal model that challenges established industry leaders, DeepSeek is contributing to a more competitive and diverse global market. This development signals a continuing trend of rapid innovation and suggests that high-performance, multimodal AI will become increasingly accessible.

Source: Bloomberg