<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Homelab on</title><link>/tags/homelab/</link><description>Recent content in Homelab on</description><generator>Hugo</generator><language>en</language><lastBuildDate>Fri, 20 Feb 2026 00:00:00 +0000</lastBuildDate><atom:link href="/tags/homelab/index.xml" rel="self" type="application/rss+xml"/><item><title>Upgrading to llama.cpp</title><link>/2026/02/upgrading-to-llama.cpp/</link><pubDate>Fri, 20 Feb 2026 00:00:00 +0000</pubDate><guid>/2026/02/upgrading-to-llama.cpp/</guid><description>&lt;h2 id="the-first-time-i-tried-llamacpp">The first time I tried llama.cpp&lt;/h2>
&lt;p>llama.cpp was one of the first inference servers available for running an LLM model. At the time, I wasn&amp;rsquo;t super impressed with AI but they were better than the old fashioned auto complete. I didn&amp;rsquo;t think a subscription was worth paying for but if I could run my own, i&amp;rsquo;d take it.&lt;/p>
&lt;p>When I first got to the llama.cpp repo on github, running it still required compiling it from source code. Annoying but it wasn&amp;rsquo;t the end of the world. What actually stopped me was that there wasn&amp;rsquo;t something like hugging face available to pull models from. The models that were around were the hundreds of billion parameter versions or I could write code to create my own from training data.&lt;/p></description></item></channel></rss>