This is my upgraded rig:
- Ryzen9 5950x with 64gb DDR4
- Dual NVIDIA RTX A4000 (16+16GB VRAM)
to whom read my previous posts, i jumped the gun and upgraded my server, it was worthwhile and somewhat cheap given i already had the two GPUs and the DDR4 RAM.
Anyway, i am currently running Qwen3.6-35B-A3B-UD-Q5_K_XL all in VRAM with 65536 context and pretty happy with speed (80-90t/s) and overall responses (mostly chat).
I would like to experiment with something beefier, with CPU offload, that i can run with my llama.cpp. Of course t/s is not a goal here, but precision and accuracy of responses is.
I tried to find a good model with claude and gemini, but always got short. Once the model suggested fully crashed my server (guess fill up RAM and ended up in a swap loop), more then once i ended up chasing non existent models. Pretty annoying.
Considering i would only use between 32 and 48GB or system RAM, can you suggest (preferably with links to HF) some models?
I like qwen3.6, but open to anything.


What are you wanting to do with the model?
Agentic development with that setup, you can easily run a good quant of Qwen3.6-27B at full context. unsloth/Qwen3.6-27B-MTP Q5_K_XL and use a fixed template
Creative writing? Probably need a different model. I’ve heard good things about Gemma 4 31B.
Both of those are dense models so won’t be as fast as Qwen3.6-35B-A3B. Also play around with MTP, quants, KV cache quant, KV caching, etc.
The RAM size will limit you on larger models unless you stream from storage.
More chats i guess, maybe also agents in the future, i am keep to explore that but not yet there.
It’s funny; I was just reading about someone who went the other way
https://bitworking.org/news/2026/05/surprising-things-i-learned-putting-together-a-home-brain/
At a certain point, it becomes less about parameters and more about tools supporting those parameters. Something like Pithagoras, Understory, MCP tools etc.
https://github.com/thecodacus/pithagoras/
https://github.com/thecodacus/understory
https://www.youtube.com/watch?v=fpvF4n32lsE
https://www.youtube.com/watch?v=IwN-eK1s8og