<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0"><channel><title><![CDATA[Overnight Desk]]></title><description><![CDATA[Overnight Desk]]></description><link>https://overnightdesk.hashnode.dev</link><image><url>https://cdn.hashnode.com/res/hashnode/image/upload/v1593680282896/kNC7E8IR4.png</url><title>Overnight Desk</title><link>https://overnightdesk.hashnode.dev</link></image><generator>RSS for Node</generator><lastBuildDate>Tue, 29 Sep 2026 15:04:34 GMT</lastBuildDate><atom:link href="https://overnightdesk.hashnode.dev/rss.xml" rel="self" type="application/rss+xml"/><language><![CDATA[en]]></language><ttl>60</ttl><item><title><![CDATA[How to Benchmark LLMs on a Raspberry Pi 5 (llama.cpp, Step by Step)]]></title><description><![CDATA[Most "Raspberry Pi AI benchmark" numbers on the internet are one-off runs with unknown settings. If you want numbers you can trust - and compare - you need a method. This is the one I use.
1. Install ]]></description><link>https://overnightdesk.hashnode.dev/how-to-benchmark-llms-on-a-raspberry-pi-5-llama-cpp-step-by-step</link><guid isPermaLink="true">https://overnightdesk.hashnode.dev/how-to-benchmark-llms-on-a-raspberry-pi-5-llama-cpp-step-by-step</guid><category><![CDATA[Raspberry Pi]]></category><category><![CDATA[Artificial Intelligence]]></category><category><![CDATA[llm]]></category><dc:creator><![CDATA[Overnight Desk]]></dc:creator><pubDate>Tue, 22 Sep 2026 15:23:02 GMT</pubDate><content:encoded><![CDATA[<p>Most "Raspberry Pi AI benchmark" numbers on the internet are one-off runs with unknown settings. If you want numbers you can trust - and compare - you need a method. This is the one I use.</p>
<h2>1. Install llama.cpp</h2>
<pre><code class="language-bash">sudo apt update &amp;&amp; sudo apt install build-essential cmake git -y
git clone https://github.com/ggml-org/llama.cpp
cd llama.cpp &amp;&amp; cmake -B build &amp;&amp; cmake --build build --config Release -j4
</code></pre>
<p>Build from source. Prebuilt binaries rarely match your kernel and flags, and a mismatched build can cost you 20% throughput.</p>
<h2>2. Pick models that actually fit</h2>
<p>A Pi 5 has 8 GB of RAM shared with the OS. Stay under ~5 GB for model + context or you'll swap and your numbers are garbage. Reliable picks in Q4: TinyLlama 1.1B, Qwen 1.5B/3B, Phi-3-mini (tight but works), Gemma 2B. Skip 7B+ unless you enjoy watching swap thrash.</p>
<h2>3. Measure honestly</h2>
<p>Use <code>llama-bench</code> with a fixed prompt and generation length. Log three numbers per run: prompt processing tok/s, generation tok/s, and SoC temperature at start and end. A run without a temperature is not a benchmark - thermal throttling on an uncooled Pi 5 shows up around 85°C and can cut throughput by a third mid-run.</p>
<pre><code class="language-bash">./build/bin/llama-bench -m models/qwen-3b-q4.gguf -p 128 -n 256 -r 3
</code></pre>
<h2>4. Keep runs repeatable</h2>
<ul>
<li>Fix the CPU governor (<code>performance</code>) for the run and note it down.</li>
<li>Same prompt set every time. Different prompts = different numbers.</li>
<li>Cool the board between runs, or say so in the log.</li>
<li>Record ambient temperature when comparing across days.</li>
</ul>
<h2>Going further</h2>
<p>I run this protocol often enough that I packaged it: a <a href="https://overnightdesk.myshoppex.io/product/raspberry-pi-local-ai-benchmark-workbook">Benchmark Workbook with the log sheets and test matrix pre-built</a>, and a <a href="https://overnightdesk.myshoppex.io/product/local-ai-benchmark-report-generator">Report Generator</a> that turns raw logs into a shareable report.</p>
<p>Full guide with more detail: <a href="https://overnightdesk-ops.github.io/">Local AI on Raspberry Pi - guides hub</a></p>
<p>What are you running on your Pi? Curious what tok/s people are seeing on 3B-class models.</p>
]]></content:encoded></item></channel></rss>