Every single rtx 5060 ti 16gb vram owner
Bonsai2 27b dense flies on your gpu with the full 262k context window, the mtp head and vision all loaded at once, and it runs at 67 tok/s. 262k context + mtp head + vision on one card: 15.2 of 15.9 gb 67 tok/s with the mtp head, 71
My new local Qwen 3.8 27B setup running on a new M5 max Mac studio.
Powerful machine, and you can run uncensored models. However, is it worth running local AI models? Spoiler: Only for 1% of use cases (I show one in the video), the rest of 99% time, you're good enough with a
Created by GLM 5.3
I might just go fully GLM 5.3 and Qwen 3.8 27B
Qoder is giving Qwen 3.8 Flash away for free
I asked it to build a live stock market dashboard from scratch. It planned the task, built the interface, opened it in a browser, and verified the result. Here’s what happened ↓
You have reached the end of the archive
All of qwen38