FRAMEWIREIndonesiaUpdated Sep 21Live wire
0:00 / 0:00

Talking about Local Benchmaxxing - here is my 2 year old Intel i9 64 GB machine with Nvidia RTX 4090 running Qwen 3.8 Flash next at 30 t/s.

Happily using opencode at my Macbook Pro using my Older Intel Machine as an API endpoint. Quant : AD-3.84bpw-IQ4_XS-M64

Akash GoswamiSep 21
0:00 / 0:00

Found a sweet spot between quality and speed for Qwen 3.8 Flash on 2x3090s @ 3.5bpw, running some tests

Backend is exllamav3, running pretty sweet @ around 105 tk/s average (92 tk/s prose/reasoning, 105 tk/s coding and 120 tk/s file editing) and 1635 tk/s prefill

VieirowskiSep 2137
0:00 / 0:00

RTX 3060でQwen 3.8 Flash Next

net58264Sep 21
0:00 / 0:00

DeepSeek V4 Flash, Qwen 3.8 27B, GLM 5.2, Nemotron 3 Ultra, Laguna S 2.1, and Jev for agent routing.

All FREE: OpenRouter: DeepSeek paths, Qwen 3.8 27B, GLM 5.2, Ling 3.0 Flash VL, Laguna S 2.1, Nemotron 3.5 Lightning / 3 Ultra: OpenCode Zen:

ZEFISep 2173