Hybrid (CPU + GPU) inference is where it’s at these days. It opens up a whole world of huge MoE models, whereas on an 5090 you are stuck with Qwen 27B.
Why you do this to me? LOL Injecting your options. But seriously thank you for the advice. I’m really green in the AI arena. So I’m trying to feel my way around, trying not to spend money on equipment I’ll regret later.







I pay around $8/month USD. Now, there are factors involved. One, electricity in my locale is relatively cheap. In addition, I have solar. To contrast the Optiplex 7020 SFF, I used run the same containers/content on a Dell T320. It cost $40/month USD. which still isn’t an outrageous expenditure for someone with another type of hobby. Some of my other hobbies, like creating music, those prices climb steeply, rapidly. Last count I have 48 guitars, several banjos, 8 bass, an oud, mandolin, 4 full size Korg synth keyboards. three violins, and amps out the yang. So $40/month was peanuts, but I wanted something that wasn’t such a hog. Plus the heat dissipated by the T320 was significant, and that is passed on to your AC bill as well. The 7020 is much quieter, energy efficient, and produces very little heat. Win, win.