FRAMEWIREIndonesiaUpdated Aug 17Live wire
0:00 / 0:00

Gemini 3.7 Flash vs Qwen 3.8 vs

DeepSeek V4 Pro vs GLM 5.3

Fabiano FirmoAug 171
0:00 / 0:00

This is how Qwen 3.8 27B looks at around 30 tok/s in mtplx

I will try a smaller quant cause it's 8bit and it grows to around 35 gb in ram usage

Chris WAug 171
0:00 / 0:00

Stop paying API visual context taxes to the cloud cartel.

I just benched Qwen3.8 27B with vision on a single NVIDIA RTX 4090 (24GB). I fed it a massive 30,000 token payload (28.5k text + 1 hi res image) to see where the physical VRAM and throughput ceilings are with and without

AlokAug 1715
0:00 / 0:00

Qwen 3.8 27b

Prompts: give me an essay on ontologies now create a website that will explain this in threeejs animation Running locally: Qwen3.8-27B UD-Q8_K_XL 220K context 2× RTX 3090 Up to 42.26 tok/s decode (~39.3 average) Hermes Recipe: Thanks

lazybutaiAug 17