Models
DeepSeek's New Vision Model Caps Any Image at 384 Tokens
DeepSeek's experimental V4-Flash-Vision-Exp prices every image at 384 tokens max and, on the company's own agent benchmarks, approaches Anthropic's Opus 4.8.

384 tokens. That is the most DeepSeek's new vision model will charge for a single image, no matter how large the original file. For anyone running invoices, shelf photos, or scanned documents through an AI pipeline by the thousand, that one number matters more than any benchmark in the announcement.
The model is DeepSeek-V4-Flash-Vision-Exp, released on August 21 as an experimental multimodal variant of the budget V4-Flash line. Pricing follows V4-Flash, whose input rate of $0.14 per million tokens already made it the price leader in its segment. Combine the two figures and the arithmetic gets striking: at 384 tokens per image, a million input tokens covers roughly 2,600 images, so the input cost of processing them runs well under a dollar. An optional "detail" field downscales images to 512x512 pixels when fine detail doesn't matter, trimming the bill further.
The Opus 4.8 claim, and who measured it
DeepSeek's headline claim is bolder than the pricing: on the company's own multimodal agent benchmarks, the model approaches and sometimes beats Anthropic's Opus 4.8. Two caveats belong next to that sentence. The numbers are DeepSeek's internal measurements, with no independent evaluation published yet, and the "Exp" suffix marks the model as explicitly experimental. Past open-model launches, Kimi K3 among them, showed vendor scoreboards and third-party tests can diverge; treat the comparison as a hypothesis until outside results land.
Frictionless to try, deliberate to adopt
Distribution is the quietly clever part. Beyond DeepSeek's own API, the model answers OpenAI-compatible Chat Completions and Responses calls and even Anthropic's Messages format, so a codebase built for GPT or Claude can point at it with a few lines changed. That near-zero switching cost is exactly what the 480-point Hacker News thread fixated on, and it is clearly the point: DeepSeek is making trial effortless.
Adoption is a different decision from trial. An experimental model that shipped days ago has no place in a production pipeline yet, whatever the price. The move that does make sense now is a controlled bake-off: take a few hundred documents or product photos from your real workload, run them through your current vision model and this one, and compare accuracy against cost on your own data. If the vendor's benchmark claims survive contact with your documents, the per-image token cap turns into a budget you can actually plan around; if they don't, you've spent pocket change finding out.
Sources: The Decoder, DeepSeek API Documentation

Written by
Muhammet Fatih Batman
Founder & Editor
Founder of YZ Uzman, with 20+ years of experience in web design and software development.