have used your qwen3 and qwen3.6 MOE, if you provide a quant Qwen3.8 27B GGUF smaller than 10G, that's much better for 12G VRAM user.

#2
by gemlincong - opened

Thanks for your contribution!!

I'm using the unsloth/Qwen3.8-27B-UD-IQ2_S (8.37 GB) on my RTX 3060 12GB, and it is very good at handling my requests in very long coding sessions. But, as we all know, the speed is very low (~9-16 t/s for a 100k context length). Even a one-digit improvement will be good here. Before the Qwen3.8 27B, my main model was the byteshape/Qwen3.6-35B-A3B-Q4_K_S, and I loved it so much.

I'm using the unsloth/Qwen3.8-27B-UD-IQ2_S (8.37 GB) on my RTX 3060 12GB, and it is very good at handling my requests in very long coding sessions. But, as we all know, the speed is very low (~9-16 t/s for a 100k context length). Even a one-digit improvement will be good here. Before the Qwen3.8 27B, my main model was the byteshape/Qwen3.6-35B-A3B-Q4_K_S, and I loved it so much.

suggest you use The tom's llama cpp turboquant version, you should get 24 token /s.

Sign up or log in to comment