NEAR AI took first place on Lean Eval v1, a benchmark of very hard formalization problems
The system runs almost entirely on the open-weight DeepSeek V4.1 Flash, so it is extremely cheap. A cheap model can work here because Lean checks every proof exactly. I built a toy
Opus 5.5 and Kokoro made a video for my latest (non AI-written) essay "Working the roofline for DeepSeek-V3 on Hopper".
I'm really happy with how it turned out!
A robot needs to know when a step is done, or it never starts the next one.
In the video, the robot puts the mug upright at 36 s… and keeps going. Checking "is it done?" sounds easy, but a useful checker has to: • understand any task and scene, not just follow one
Yeah, Deepseek said the same thing until I pointed out some uncomfortable truths
Like the 2015 Nobel prize, EUA's etc, and suddenly it said "me bad" you are right!
You have reached the end of the archive
All of deepseek