Openai took its first AI chip from design to tapeout in nine months
Now it says the chip beats systems built on NVIDIA’s GB200 and GB300 on the curve agents care about. Today OpenAI published the first measured results from Jalapeño, its custom inference chip. Across GPT-OSS
$0.04. That's what DeepSeek-V4-Flash-Vision-Exp cost to watch the bot Grok Bot library video.
DeepSeek V4 Flash ran the same prompt through ego lite, blind to every frame: 27.9 minutes, $0.14. 3.3x the cost, for a model spending the whole run trying to see.
LLM Model Selection in Loupe for Tableau v1.0.6
Switch between Claude, Deepseek, OpenAI, and xAI (Grok) on the fly. ⚡ One click. Model switched. 🔗
Calculate the recent popular models and hardware data: M5 Max Qwen 3.8 27B MLX 4bit (dense model)
Basic short contextPP about 900-925 tok/sTPS 32-33 tok/s ~32k TPS and 28-29 tok/s In most cases, there is a significant improvement after turning on MTP or DFlash. The actual measurement of MTPLX is mostly 55-65 tok/sDFlash2, some people have reached about 70 tok/s M5 Ultra Today…
You have reached the end of the archive
All of deepseek