https://qwen.ai/blog?id=qwen3.8
The model weights will be open-sourced on Hugging Face and ModelScope next week — stay tuned.

I’m hoping they also release a new 35b a3b, for us VRAM poors, a new 9b would also be great!
https://qwen.ai/blog?id=qwen3.8
The model weights will be open-sourced on Hugging Face and ModelScope next week — stay tuned.

I’m hoping they also release a new 35b a3b, for us VRAM poors, a new 9b would also be great!
With which quantization ? I have a 16GB GPU, running Qwen3.5 9b UD_Q8_K_XL basically max out the VRAM usage. Maybe a MOE model fits you. Gemma 4 e4b is one with a total 8b.
And what is the token speed you are getting ?