How i feel when i combine glm 5.2, qwen 3.8, deepseek v4 flash 0731 and ornith 397b across my DGX Sparks 5090s and 3090s .
Free GPUs 2x NVIDIA T4s & run Qwen 3.8 27B at 14 t/s with 120k context
120k context on FREE Kaggle T4s. - FP16 KV - 14 t/s - No credit card. - No expensive GPU. - Total VRAM used: only 26.6 GB across both cards. - Under 5 minutes setup. - Kaggle gives you 30 hours/week of
Preliminary experiment with Qwen 3.8/MiniMax H3
I'm developing a prompt that takes a comic page, cuts it out and turns it into a video.
Running Qwen 3.8 27B on 2×5090s with a few basic optimizations
~230 tokens/s (video is real-time) two years ago building a pubsub would have taken me an entire day now it takes 8 seconds
You have reached the end of the archive
All of qwen38