UC Berkeley open-sourced FreeToken.
(2–4x faster local LLM inference than Ollama) the results are wild: - Qwen3.6-35B on an 8GB GPU at 39.3 tokens/s - DeepSeek-V4-Flash 284B on a 32GB GPU at 22 tokens/s - GLM-5.2 753B on a 96GB GPU at 14.9 tokens/s a 35B model at 16-bit
Same prompt. Four AI models.
Wildly different results — and one number that changes the entire conversation. The task: build a single-file HTML traffic intersection simulation. Dark environment, glowing headlights, realistic traffic light logic, smooth 30-second loop. Here's
Gave the same prompt to 4 models and asked them to build a pseudo-3D racing game in one shot.
DeepSeek V4.1 Flash MiMo V2.6 Flash Qwen3.8 Flash GPT-6 Luna Pretty different takes
I’m building Little Chittagong 🌆
A browser-based 3D version of Chattogram where you can walk, drive, explore the Port, Patenga, Agrabad, GEC, Shah Amanat Bridge and more. Inspired from masterofnone - Built with Three.js + OpenStreetMap, with
You have reached the end of the archive
All of deepseek