Pretty impressed to see lfm2.5-vl by liquidai complete the task accurately too
While being noticeably faster than deepseek v4 flash vision and glm 5.3 flash I honestly didn’t expect a 3b model to do this well on a robotics task.
Inspired by Sentdex post, I was curious to try the same idea and see how far general multimodal models can get in a robotics environment.
This was my first time trying something like this, so just a basic test: DeepSeek V4 Flash Vision vs GLM 5.3 Flash on a simulated Franka
With all the drops this week time for some “Friday AI Phonk.”
Crank up the volume! Friday AI phonk, yeah the week went off the rails Every update dropping heavy, every model leaving trails DGX Spark in the chamber, local inference never fails Omarchy Linux humming, hybrid
95% of my development tasks are now completed using Deepseek Harness.
I have made many plug-ins based on my own development needs, and they have all been open source. I have also packaged my DSH workbench and put it on GitHub. You can pick it up if you need it.
You have reached the end of the archive
All of deepseek