❯ DeepSeek open-sources V4-Flash-Vision-Exp weights, 305B parameters under an MIT license
weights outDeepSeek released the full weights for V4-Flash-Vision-Exp on Hugging Face — 305B parameters under an MIT license. It is the first multimodal model in the V4 family; the API went live on August 21, with weights following ten days later after a round of live validation.
capabilitiesThe model adds a vision module to the V4-Flash architecture, accepting JPEG, PNG, GIF and WebP, and can describe images, read text in screenshots, parse charts and run agent tasks with tools. DeepSeek says pure-text performance matches the production V4-Flash and that multimodal agent capability approaches Opus-4.8. The release also includes minimal inference implementations for the vision encoder, Aligner, DFlash Attention and MoE modules. Until now V4 shipped text-only, trailing Moonshot and Z.ai on vision within China’s open-weight camp — this fills that gap.
licenseMIT means commercial use with no strings. For teams embedding vision agents into their own products, the release moves the bar for running closed-model-class capability off the API invoice and onto their own GPUs: a one-time hardware cost instead of per-call billing, in exchange for owning deployment and operations. Watch how quickly inference providers cut prices in response.
▪ SIGNAL305B weights under MIT put the first squeeze not on OpenAI but on every domestic multimodal API billing by the call.