Back
DeepSeek Introduces Vision: Chat Reads Images, No API Yet
SiTech AI Team3 წთ. საკითხავი

DeepSeek Introduces Vision: Chat Reads Images, No API Yet

DeepSeek has added a Vision mode to its chat app, letting its consumer chatbot read images for the first time. The rollout is a limited beta, and no public vision API sits behind it yet.

DeepSeek has added a Vision (image-recognition) mode to its chat app — the first time its consumer chatbot can work with images rather than text alone. The feature surfaced on Hacker News on June 17, 2026 and rolled out on June 18 as a limited beta to selected users on DeepSeek's website and mobile app.

What shipped

Vision is a third operating mode that sits alongside the existing Fast (Instant) and Expert options: users switch to it the way they switch between text modes, attach a picture and ask questions about it. According to Chen Xiaokang, who leads DeepSeek's multimodal team and is one of the authors of the DeepSeek-VL model series, the function went first to a subset of users for beta testing. The release follows the V4 flagship launch — V4-Pro and V4-Flash, open-weight under the MIT license, shipped in April 2026.

Thinking with visual primitives

The technical idea behind the mode is called "Thinking with Visual Primitives". Instead of describing an image in plain text, the model marks points and boxes directly on the picture as part of its reasoning chain — similar to a person running a finger along a line while counting. The developers argue that text alone is too vague to pinpoint a specific object in a cluttered scene, and that vagueness is what pushes models into confusion and hallucination.

Vision is built on top of DeepSeek-V4-Flash. To keep image processing from consuming too much compute, the team compresses memory: every four visual tokens collapse into a single entry, which yields a noticeably lower cost per image than typical multimodal models. In practice that means more accurate answers for complex diagrams, interface screenshots and detail-heavy photos, because the model can point to the part of the image its conclusion rests on.

What it means

According to the developers, the model matches GPT-5.4, Claude Sonnet 4.6 and Gemini 3 Flash on object-counting and spatial-reasoning tasks — a narrow benchmark set tailored to this specific work, not a general capability claim. The model weights have not been released. Just as important for builders: there is no public vision API model ID yet. The feature lives inside the chat product rather than as a developer endpoint, so anyone planning to call DeepSeek's vision from code still has to wait.

DeepSeek is betting on efficiency rather than raw size — fewer tokens per image means cheaper inference for the end user — while competition in multimodal models stays intense, with new releases from OpenAI, Google and Chinese developers arriving almost weekly.

SSiTech

SiTech — AI-powered web development

We build fast, modern websites and bring AI into real business workflows. Have a project or a question? We'd love to help.