Run E2E with Deepseek multi-modal + Midscene, the speed is almost human... (The video is not accelerated)
DeepSeek just fixed one of the biggest weaknesses in AI agents
They can finally SEE what they’re working on. DeepSeek V4 Flash Vision EXP can inspect screenshots, charts, images, and pages. That opens up some very practical agent workflows.
Demo 2: GitHub Sign-up
Below is a demo of Midscene + DeepSeek signing up for a GitHub account. The model plans and executes the task entirely on its own. The video is neither accelerated nor edited, and completing the entire form takes only about 50 seconds.
Demo 1: LinkGame Tile-Matching Game
Below is a demo of Midscene + DeepSeek playing a LinkGame tile-matching game. The entire process relies purely on visual localization, with no speed-up or editing. Its operating speed is nearly human.
You have reached the end of the archive
All of deepseek