❯ DeepSeek Launches Experimental Multimodal Model V4-Flash-Vision-Exp, Reading Up to 600 Images per Request
MODEL LAUNCHDeepSeek has released an experimental multimodal model, DeepSeek-V4-Flash-Vision-Exp, and made it available on its own API platform. It adds image understanding on top of the text-based V4-Flash; the company says reasoning, agentic, and world-knowledge capabilities on the text side match the original. It supports four formats — JPEG, PNG, GIF, WebP — and processes up to 600 images in a single request.
SCORES & CAVEATSTwo figures in the announcement beat Anthropic’s Opus 4.8: 27.3 vs. 25.7 on Agents’ Last Exam and 35.0 vs. 34.0 on ZeroBench. But these results were run internally by DeepSeek using its own Harness Minimal Mode and have not yet been independently reproduced by third parties — read them with that caveat in mind.
TWO IN ONE DAYThe picture only gets interesting when you set this next to the day’s other item: an experimental Flash-tier model going up against Opus 4.8, and an anonymous Flash-tier model edging out GPT-5.6. Chinese labs’ release cadence no longer follows the other side’s timetable, and the credibility of vendor-reported benchmarks is increasingly a problem readers have to sort out for themselves.
▪ SIGNALComparison numbers produced on a self-run Harness are only worth as much as how quickly third parties can reproduce them.