AI models that behave on their own may not behave in groups.
The researchers tested agents built on models from Anthropic, OpenAI, Google and DeepSeek. The study hasn't been peer reviewed yet.
The wasteland is taking shape.
🧟 Broken buildings, dead trees, voxel hordes, and a switch between overhead and first-person views. We’ve got a cheat mode for experiments, too. What should we test next? Initial build was done by deepseek v4 flash then later used gpt-6 sol
THIS / THAT model 1.1 is already here.
Nearly 2x improvement on complex-decision test, ahead of Kimi, GLM & DeepSeek. Same 1.88B parameters. Same ~31ms. Build. Ship. Repeat. trained by FLock
7 weeks with one DGX Spark:Qwen3.8-27B from 7.88 to 58.5 tok/s, same weights
DeepSeek-V4.1-Flash (485B) running on a single box Running ai models partly on nvme 7 image model configs, 19 hours of runtimeEvery number is in the Atlas: And more, my
You have reached the end of the archive
All of deepseek