I also added a function to voice change only the vocals. Originally a Suno song.
Everything from songwriting to lip-syncing music video production can now be completed locally. Workflow to connect and run MiniMax H3, LTX 2.5, ACE-Step, Qwen 3.8 27B from iPhone (CloseBox) TechnoEdgeJP
50% more context unlocked for Qwen 3.8 27b Q4_K_XL dflash 2 on a single RTX 4090 (24 GB VRAM)
I found a hidden VRAM tax in llama.cpp. By combining my custom 2 bit DFlash 2 drafter with one overlooked server flag, I just unlocked another +80,000 tokens of context. Qwen3.8-27B
Deployed qwen 3.8 27b with dflash2 on 2 rtx 5090s
Steady generation: ~220 tok/s Normal range: ~190–257 tok/s Best observed: ~343 tok/s sgl_project with dflash2 is blazing fast
You have reached the end of the archive
All of qwen38