<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Small Language Models on</title><link>/categories/small-language-models/</link><description>Recent content in Small Language Models on</description><generator>Hugo</generator><language>en</language><lastBuildDate>Fri, 20 Feb 2026 00:00:00 +0000</lastBuildDate><atom:link href="/categories/small-language-models/index.xml" rel="self" type="application/rss+xml"/><item><title>Upgrading to llama.cpp</title><link>/2026/02/upgrading-to-llama.cpp/</link><pubDate>Fri, 20 Feb 2026 00:00:00 +0000</pubDate><guid>/2026/02/upgrading-to-llama.cpp/</guid><description>&lt;h2 id="the-first-time-i-tried-llamacpp">The first time I tried llama.cpp&lt;/h2>
&lt;p>llama.cpp was one of the first inference servers available for running an LLM model. At the time, I wasn&amp;rsquo;t super impressed with AI but they were better than the old fashioned auto complete. I didn&amp;rsquo;t think a subscription was worth paying for but if I could run my own, i&amp;rsquo;d take it.&lt;/p>
&lt;p>When I first got to the llama.cpp repo on github, running it still required compiling it from source code. Annoying but it wasn&amp;rsquo;t the end of the world. What actually stopped me was that there wasn&amp;rsquo;t something like hugging face available to pull models from. The models that were around were the hundreds of billion parameter versions or I could write code to create my own from training data.&lt;/p></description></item><item><title>Setting up Ollama server with Open WebUI</title><link>/2025/12/setting-up-ollama-with-open-webui/</link><pubDate>Sat, 20 Dec 2025 00:00:00 +0000</pubDate><guid>/2025/12/setting-up-ollama-with-open-webui/</guid><description>&lt;p>Ollama is a powerful AI model serving platform that makes it easy to run large language models locally. In this post, we&amp;rsquo;ll walk through setting up Ollama with Open WebUI in a Docker containerized environment, including GPU acceleration support. Visit &lt;a href="https://ollama.com/" target="_blank" rel="noopener noreffer ">Ollama&amp;rsquo;s website&lt;/a> for more information.&lt;/p>
&lt;h2 id="what-is-ollama">What is Ollama?&lt;/h2>
&lt;p>Ollama is an AI platform that provides an easy way to run large language models locally. It offers:&lt;/p>
&lt;ul>
&lt;li>Easy model management&lt;/li>
&lt;li>GPU acceleration support&lt;/li>
&lt;li>A clean API for model interactions&lt;/li>
&lt;li>Containerized deployment&lt;/li>
&lt;/ul>
&lt;h2 id="what-is-open-webui">What is Open WebUI?&lt;/h2>
&lt;p>Open WebUI provides a web-based interface for interacting with Ollama models, making it easy to use AI models without needing to write code or deal with complex command-line interfaces. Visit the &lt;a href="https://github.com/open-webui/open-webui" target="_blank" rel="noopener noreffer ">Open WebUI GitHub repository&lt;/a> for more information.&lt;/p></description></item></channel></rss>