Welcome to the Homelab Hunger Games: Local LLM Edition!
Have you ever wondered what happens when you force budget hardware to run local AI models? Will your CPU gracefully handle the math, or will it scream in binary while fans spin at jet-engine decibels?
We put three popular micro-LLMs (Qwen 2.5 Coder 1.5B, Gemma 2 2B, and Ministral 3 3B) through an Ollama benchmark gauntlet across two very different machines. Grab your coffee—because as you’ll see, one of these models gave us plenty of time to brew a fresh pot.
🥊 Meet the Contenders (and the Victims)#
In the red corner, we have Draken, powered by an Intel Core i5-10210U with zero dedicated GPU (pure CPU-only suffering).
In the blue corner stands Geneseed, rocking an 11th Gen Core i7 and an NVIDIA GeForce MX450 GPU with a modest 2GB of VRAM. An MX450 is usually found in laptops designed for office work, but today it puts on a superhero cape!
🛠️ The Hardware Rig Breakdown#
- Geneseed (The Speedster): Intel i7-11370H @ 3.30GHz | NVIDIA GeForce MX450 (2GB VRAM)
- Draken (The Sweat Factory): Intel i5-10210U @ 1.60GHz | iGPU / CPU Only
- The Prompt: “Write a 200 word summary explaining how Linux kernel modules work.”
- Benchmark Tool: Ollama Benchmark Suite
⚡ Generation Speed: Caffeinated vs. Crawling#
When it comes to output speed (Tokens per Second), Qwen 2.5 Coder 1.5B on Geneseed was an absolute rocket ship, blazing past at 43 tokens/sec. That’s faster than you can read out loud without gasping for air!
Meanwhile, Draken’s CPU choked Ministral 3B down to a modest 10 tokens/sec—just enough speed to remind you of a 1996 dial-up modem printing text line by line.
⏳ Model Load Times: The 140-Second Paid Vacation#
Cold-loading a model into system memory vs VRAM is where things get hilarious (and painful).
Most models loaded in under 15 seconds. But then came Gemma 2 2B on Draken’s CPU…
🧠 Time To First Token (TTFT): Existential Crisis Mode#
Time To First Token measures how long the model sits back and thinks about its life choices before printing the very first character.
While Qwen and Gemma reacted almost instantly (under 0.5s), Ministral 3 3B had to process a 568-token prompt. On GPU, it handled it in 3.92s. On CPU? A staggering 12.55 seconds of silence!
🏆 The Verdict & Recommendations#
- Uncontested Champ Qwen 2.5 Coder 1.5B easily takes the crown. It’s fast, lightweight, and doesn’t melt your CPU.
- Even a Potato GPU Helps: That 2GB MX450 GPU—a chip nobody buys for gaming—boosted inference speed by up to 2.8x and eliminated cold-start loading hell.
- Keep Gemma Warm: If you run Gemma 2 on a CPU homelab server, configure
keep_alive: -1in Ollama so you don’t suffer the 140-second penalty on every single request.
💬 What’s Running in Your Homelab?#
What budget hardware are you using for local AI? Are you team Qwen or team Mistral? Let us know in the comments below, and don’t forget to check out the benchmark video!
Watch Full Benchmark Video
