FRAMEWIREIndonesiaUpdated Aug 16Live wire
0:00 / 0:00

I asked Qwen3.8 (q4) to build something interesting.

This is what it came out with before I ran out of memory for the context size (49.25 GB for 166k). Details in the replies.

Florian HerrengtAug 161
0:00 / 0:00

6 months ago, the smartest model on earth lived behind a paywall in the cloud.

Today something in the same league runs on a laptop. For free. On August 14th, Alibaba's Qwen team dropped Qwen 3.8-27B. 27 billion parameters. Tiny by frontier standards. But the benchmarks they

Julian Goldie SEOAug 163
0:00 / 0:00

Running 133 tok/s with QWEN 3.8 27b RTX 4090

32K context is what can fit inside of 24GB VRAM HyperInference stack: • Q4_K_M + full GPU offload • MTP speculative decoding • Flash Attention + Q8 KV cache • 32K context • Live tok/s Could try getting to 200K Context off this,

Rob McElvennyAug 16
0:00 / 0:00

Unlock AI power without the hefty price tag.

Qwen 3.8 offers a cost-effective edge over OpenAI & Claude, especially for C++ debugging. Python's context maintenance shines for coding, while high-frequency trading demands serious capital.

Bryan DowningAug 16