<?xml version="1.0" encoding="UTF-8"?><rss version="2.0" xmlns:content="http://purl.org/rss/1.0/modules/content/">
  <channel>
    <title>inti-tidball</title>
    <link>https://inti-tidball.writeas.com/</link>
    <description> code &amp; chaos &amp; a bit of everything else</description>
    <pubDate>Mon, 05 Oct 2026 23:40:57 +0000</pubDate>
    <item>
      <title>Por qué uso (y recomiendo) dominios .xyz numéricos para self-hosting</title>
      <link>https://inti-tidball.writeas.com/por-que-uso-y-recomiendo-dominios-xyz-numericos-para-self-hosting?pk_campaign=rss-feed</link>
      <description>&lt;![CDATA[Self-hosting desde tu red local (sin IP pública): .xyz numérico + Cloudflare Tunnel&#xA;&#xA;El problema&#xA;&#xA;Self-hostear algo como Grafana, tu app, una API, necesita tres cosas: un nombre estable que resuelva a tu IP, un certificado TLS válido (Let&#39;s Encrypt no emite para IPs), y  no depender de un subdominio de DDNS gratuito que puede cambiar o desaparecer. Un dominio propio resuelve las tres. El asunto es el costo y la disponibilidad del nombre.&#xA;&#xA;Ahí entra la clase numérica de .xyz.&#xA;&#xA;Y si hosteás desde tu casa, se suma un cuarto: el ISP normalmente no te da IP pública sino que te deja detrás de CGNAT (una IP compartida entre varios clientes). Con CGNAT el port forwarding ni siquiera funciona ya que no hay a dónde forwardear.&#xA;&#xA;Un dominio propio + un túnel saliente resuelven las cuatro.&#xA;&#xA;Qué es la clase 1.111B&#xA;&#xA;  En 2017 el registro de .xyz abrió un bloque: todos los dominios de 6 a 9 dígitos con terminación .xyz (de 000000.xyz a 999999999.xyz), 1.111 mil millones de combinaciones, a US$0,99&#xA;  por año para siempre — registro, renovación y transferencia, con las tasas ICANN incluidas. La renovación cuesta lo mismo que el alta.&#xA;&#xA;Por qué lo recomiendo&#xA;&#xA;Costo irrisorio: ~US$1/año, para siempre y sin sorpresas en la renovación.&#xA;Disponibilidad: con 1.111 mil millones de combinaciones, encontrás uno libre sin pelear.&#xA;Sin riesgo de marca: un número no colisiona con ninguna marca ni te lo van a disputar.&#xA;Funcional, no de marca: para self-hosting no necesitás un nombre lindo ni memorable, solo necesitás que resuelva.&#xA;Wildcard + subdominios: un solo dominio te da infinitos subdominios con un registro wildcard.&#xA;Let&#39;s Encrypt gratis: con el desafío DNS-01 sacás un certificado wildcard (.12345678.xyz) y cubrís todos los subdominios con uno solo.&#xA;Independencia: es tu dominio. No dependés de un DDNS de terceros que puede cerrarse.&#xA;&#xA;Cómo lo uso&#xA;&#xA;El dominio numérico (ej. 12345678.xyz) como raíz, con DNS en Cloudflare (gratis) y un wildcard . Y lo clave para hostear desde casa es el Cloudflare Tunnel (el demonio cloudflared) que me dá un túnel saliente hacia Cloudflare, sin IP pública y sin abrir puertos en el router. Cada servicio tiene su subdominio (app., api., grafana.).&#xA;&#xA;Contras&#xA;&#xA;Es feo y poco memorable. Para self-hosting da igual; para una marca, no.&#xA;Los TLD baratos (.xyz incluido) arrastran cierta reputación de spam/phishing; algunos filtros corporativos los miran con recelo. Para uso propio no molesta, pero puede aparecer.&#xA;No es para producción &#34;seria&#34; de cara a terceros (un banco no va a confiar), pero para infraestructura propia y demos va perfecto.&#xA;]]&gt;</description>
      <content:encoded><![CDATA[<p><img src="https://i.postimg.cc/MTpj9Cb9/diagrama.png" alt="Self-hosting desde tu red local (sin IP pública): .xyz numérico + Cloudflare Tunnel"/></p>

<h2 id="el-problema">El problema</h2>

<p>Self-hostear algo como Grafana, tu app, una API, necesita tres cosas: un <strong>nombre estable</strong> que resuelva a tu IP, un <strong>certificado TLS válido</strong> (Let&#39;s Encrypt no emite para IPs), y  <strong>no depender de un subdominio de DDNS gratuito</strong> que puede cambiar o desaparecer. Un dominio propio resuelve las tres. El asunto es el <strong>costo</strong> y la <strong>disponibilidad</strong> del nombre.</p>

<p>Ahí entra la clase numérica de .xyz.</p>

<p>Y si hosteás desde tu casa, se suma un cuarto: el ISP normalmente no te da IP pública sino que te deja detrás de CGNAT (una IP compartida entre varios clientes). Con CGNAT el port forwarding ni siquiera funciona ya que no hay a dónde forwardear.</p>

<p>Un <strong>dominio propio</strong> + un <strong>túnel saliente</strong> resuelven las cuatro.</p>

<h2 id="qué-es-la-clase-1-111b">Qué es la clase 1.111B</h2>

<p>  En 2017 el registro de .xyz abrió un bloque: todos los dominios de 6 a 9 dígitos con terminación .xyz (de 000000.xyz a 999999999.xyz), 1.111 mil millones de combinaciones, a US$0,99
  por año para siempre — registro, renovación y transferencia, con las tasas ICANN incluidas. La renovación cuesta lo mismo que el alta.</p>

<h2 id="por-qué-lo-recomiendo">Por qué lo recomiendo</h2>
<ol><li><strong>Costo irrisorio</strong>: ~US$1/año, para siempre y sin sorpresas en la renovación.</li>
<li><strong>Disponibilidad</strong>: con 1.111 mil millones de combinaciones, encontrás uno libre sin pelear.</li>
<li><strong>Sin riesgo de marca</strong>: un número no colisiona con ninguna marca ni te lo van a disputar.</li>
<li><strong>Funcional, no de marca</strong>: para self-hosting no necesitás un nombre lindo ni memorable, solo necesitás que resuelva.</li>
<li><strong>Wildcard + subdominios</strong>: un solo dominio te da infinitos subdominios con un registro wildcard.</li>
<li><strong>Let&#39;s Encrypt gratis</strong>: con el desafío DNS-01 sacás un certificado wildcard (*.12345678.xyz) y cubrís todos los subdominios con uno solo.</li>
<li><strong>Independencia</strong>: es tu dominio. No dependés de un DDNS de terceros que puede cerrarse.</li></ol>

<h2 id="cómo-lo-uso">Cómo lo uso</h2>

<p>El dominio numérico (ej. 12345678.xyz) como raíz, con DNS en Cloudflare (gratis) y un wildcard *. Y lo clave para hostear desde casa es el Cloudflare Tunnel (el demonio <em>cloudflared</em>) que me dá un túnel saliente hacia Cloudflare, sin IP pública y sin abrir puertos en el router. Cada servicio tiene su subdominio (app., api., grafana.).</p>

<h2 id="contras">Contras</h2>
<ul><li>Es <strong>feo y poco memorable</strong>. Para self-hosting da igual; para una marca, no.</li>
<li>Los TLD baratos (.xyz incluido) arrastran cierta <strong>reputación de spam/phishing</strong>; algunos filtros corporativos los miran con recelo. Para uso propio no molesta, pero puede aparecer.</li>
<li>No es para producción “seria” de cara a terceros (un banco no va a confiar), pero para infraestructura propia y demos va perfecto.</li></ul>
]]></content:encoded>
      <guid>https://inti-tidball.writeas.com/por-que-uso-y-recomiendo-dominios-xyz-numericos-para-self-hosting</guid>
      <pubDate>Mon, 05 Oct 2026 20:58:30 +0000</pubDate>
    </item>
    <item>
      <title>Build Your Own AI Bunker: Privacy-First Agents on Outdated Hardware</title>
      <link>https://inti-tidball.writeas.com/build-your-own-ai-bunker-privacy-first-agents-on-outdated-hardware?pk_campaign=rss-feed</link>
      <description>&lt;![CDATA[In an era where AI productivity often comes at the cost of data privacy, building a local &#34;Bunker&#34; is the sweet and often sought after solution. I was inspired by Liz Howard and their focus on open-source model accessibility, and Simon Willison’s experiments running Claude Code in &#34;dangerous&#34; containerized environments. So I set out to make my own local stack.&#xA;&#xA;While Simon might be running an Nvidia Spark, I am humbly proof-of-concepting this on an old Thinkpad with 16GB of RAM. The goal was to test the limits, see what breaks, and keep my data offline. Hopefully I too will one day be able to do similar experiments!&#xA;&#xA;The Hardware Stack&#xA;&#xA;Host: Proxmox VE 9 running on a ThinkPad T420&#xA;CPU: Intel Core i5-2520M (4 Cores, 2 Threads each)&#xA;RAM: 16GB DDR3&#xA;Storage: 480GB SSD&#xA;&#xA;1. The Infrastructure (Proxmox LXC)&#xA;&#xA;The stack lives in a Proxmox LXC. If you wanted to allow the container to run its own instances, you must grant the LXC specific permissions on the Proxmox host. My initial idea was to run agents with Podman, so I had to do this:&#xA;&#xA;Edit /etc/pve/lxc/ID.conf&#xA;features: nesting=1,keyctl=1&#xA;&#xA;Allows nesting and keyctl=1 allows the container to manage its own cryptographic keys and process-specific tokens which is required when running Podman in a nested environment. &#xA;&#xA;2. The Engine (llama.cpp)&#xA;&#xA;To squeeze performance out of an older i5, we compile locally to utilize the AVX (AVX1) instruction set.&#xA;&#xA;Install dependencies&#xA;apt update &amp;&amp; apt install -y   git curl tmux direnv   build-essential cmake ccache   libssl-dev libsqlite3-dev liblzma-dev   zlib1g-dev libbz2-dev libreadline-dev&#xA;&#xA;Clone repo&#xA;git clone https://github.com/ggerganov/llama.cpp&#xA;&#xA;Compile&#xA;cd llama.cpp&#xA;cmake -B build -DGGMLAVX=ON -DGGMLNATIVE=ON&#xA;cmake --build build --config Release --target llama-server -j 4&#xA;&#xA;I included ccache in the dependencies. It’s optional, but it will cut recompilation time.&#xA;&#xA;Run a few tests:&#xA;&#xA;The llama-server binary should exist&#xA;ls -lh build/bin/llama-server&#xA;&#xA;3. Orchestration &amp; Persistence (TMUX)&#xA;&#xA;On limited hardware, high CPU load can cause SSH lag. Use TMUX inside the LXC to decouple the LLM from your connection.&#xA;&#xA;Step-by-Step Setup:&#xA;&#xA;Start Session&#xA;tmux new -s aibunker&#xA;&#xA;Prepare Models&#xA;mkdir -p models&#xA;wget -O models/qwen-0.5b.gguf https://huggingface.co/Qwen/Qwen2.5-Coder-0.5B-Instruct-GGUF/resolve/main/qwen2.5-coder-0.5b-instruct-q4km.gguf&#xA;&#xA;Fire up the Server&#xA;mkdir -p bunkermemory&#xA;./build/bin/llama-server   \&#xA;-m models/qwen-0.5b.gguf  \&#xA;--port 8080  \&#xA;--ctx-size 2048  \&#xA;--threads 4   \&#xA;--mlock   \&#xA;--no-mmap  \&#xA;--host 0.0.0.0   \&#xA;--slot-save-path ./bunkermemory&#xA;&#xA;4. Direct Inference (curl)&#xA;&#xA;curl http://localhost:8080/v1/chat/completions   -H &#34;Content-Type: application/json&#34;   -d &#39;{&#xA;    &#34;model&#34;: &#34;local&#34;,&#xA;    &#34;maxtokens&#34;: 128,&#xA;    &#34;temperature&#34;: 0.2,&#xA;    &#34;messages&#34;: [&#xA;      {&#xA;        &#34;role&#34;: &#34;user&#34;,&#xA;        &#34;content&#34;: &#34;Write a simple Python function that reverses a string.&#34;&#xA;      }&#xA;    ]&#xA;  }&#39;&#xA;&#xA;5. Connecting via browser&#xA;&#xA;Since the llama-server is bound to 0.0.0.0:8080, it’s accessible to any tool on your network. The lightest method is via a browser calling the endpoint directly.&#xA;&#xA;6. Connecting locally (Wave &amp; Claude &amp; Aider)&#xA;&#xA; Most modern terminals and AI extensions expect an OpenAI-compatible endpoint. However, these tools inject massive system prompts, and will cause context overflow errors and eat up more tokens than you might have. &#xA;&#xA;I got it working, but it was slow. If you have a better machine and want to go ahead, you may want to configure your tools to use your local model. To do so for Wave, you can follow the documentation. Most AI terminals have similar setup. You need to edit the configuration at  ~/.config/waveterm/waveai.json or run a custom command to configure Wave:&#xA;&#xA;wsh editconfig waveai.json&#xA;&#xA;Inside the JSON file, add a new entry. Based on your existing llama-server setup, use this exact structure:&#xA;&#xA;{&#xA;  &#34;ia-bunker&#34;: {&#xA;    &#34;display:name&#34;: &#34;IA Bunker - Qwen Coder&#34;,&#xA;    &#34;display:order&#34;: 1,&#xA;    &#34;display:icon&#34;: &#34;shield&#34;,&#xA;    &#34;display:description&#34;: &#34;Local Qwen 2.5 Coder on Thinkpad T420&#34;,&#xA;    &#34;ai:apitype&#34;: &#34;openai-chat&#34;,&#xA;    &#34;ai:model&#34;: &#34;qwen-coder&#34;,&#xA;    &#34;ai:endpoint&#34;: &#34;http://192.168.1.XX:8080/v1/chat/completions&#34;,&#xA;    &#34;ai:apitoken&#34;: &#34;not-needed&#34;&#xA;  }&#xA;}&#xA;&#xA;Replace 192.168.1.XX with your local IP.&#xA;&#xA;wsh setconfig waveai:defaultmode=&#34;ia-bunker&#34;&#xA;&#xA; I also experimented with running higher-level terminal agents (for example, Claude and Aider) inside Podman and pointing them at the bunker via environment variables. &#xA;&#xA;On a 2 k token context, this usually results in exceedcontextsizeerror failures or multi-minute stalls before the first token is produced.&#xA;&#xA;These are the commands I tried inside the container that successfully loaded Aider &amp; then Claude, but unfortunately, they just stalled with my setup (you may have more luck): &#xA;&#xA;podman run -it --rm \&#xA;  -v $(pwd):/app \&#xA;  -e OPENAIAPIBASE=http://172.17.0.1:8080/v1 \&#xA;  -e OPENAIAPIKEY=not-needed \&#xA;  paulgauthier/aider-full \&#xA;  --model openai/gpt-3.5-turbo \ #  used to select prompt profile&#xA;  --no-auto-commits&#xA;&#xA;podman run -it --rm \&#xA;  --network host \&#xA;  -v $(pwd):/app \&#xA;  -e ANTHROPICBASEURL=&#34;http://127.0.0.1:8080&#34; \&#xA;  -e ANTHROPICAUTHTOKEN=&#34;local-bunker&#34; \&#xA;  -e CLAUDECODEDISABLENONESSENTIALTRAFFIC=1 \&#xA;  node:20-alpine \&#xA;  sh -c &#34;apk add --no-cache git &amp;&amp; npm install -g @anthropic-ai/claude-code@latest &amp;&amp; claude --model qwen-coder&#34;&#xA;7. The Global Tunnel: Tailscale in the Bunker &#xA;&#xA; To access your &#34;Bunker&#34; from a coffee shop you could install Tailscale, CloudFlare or Zerotier directly inside the LXC. This turns your old Thinkpad into a secure, private node on your personal VPN. It’s probably worth securing the entry with some auth if you wanted to do that. &#xA;&#xA;But it’s a topic for another post!&#xA;&#xA;PS: Tweak your Engine &#xA;&#xA; Initially, I was getting a painfully slow 0.25 tokens/s. To make it usable on this vintage i5, I had to aggressively tune the llama-server flags:&#xA;&#xA;--ctx-size 2048: Dropping this down helped because smaller context equals faster &#34;thinking&#34;.&#xA;-ctk q80 -ctv q8_0: I tried this KV-cache quantization (-ctk/-ctv), but opted out, as on a small-context (2k) Sandy Bridge CPU, the added CPU overhead often outweighs the memory savings. If I had a newer CPU, this might have helped.&#xA;--mlock: Essential to prevent the model from being paged to the Swap partition, which would turn the 20-minute wait into a total system hang.&#xA;&#xA;]]&gt;</description>
      <content:encoded><![CDATA[<p>In an era where AI productivity often comes at the cost of data privacy, building a local “Bunker” is the sweet and often sought after solution. I was inspired by <a href="https://www.lizthe.dev/" rel="nofollow">Liz Howard</a> and their focus on open-source model accessibility, and <a href="https://simonwillison.net/" rel="nofollow">Simon Willison</a>’s experiments running Claude Code in “dangerous” containerized environments. So I set out to make my own local stack.</p>

<p>While Simon might be running an Nvidia Spark, I am humbly proof-of-concepting this on an old Thinkpad with 16GB of RAM. The goal was to test the limits, see what breaks, and keep my data offline. Hopefully I too will one day be able to do similar experiments!</p>

<h2 id="the-hardware-stack">The Hardware Stack</h2>
<ul><li><strong>Host:</strong> Proxmox VE 9 running on a ThinkPad T420</li>
<li><strong>CPU:</strong> Intel Core i5-2520M (4 Cores, 2 Threads each)</li>
<li><strong>RAM:</strong> 16GB DDR3</li>
<li><strong>Storage:</strong> 480GB SSD</li></ul>

<h2 id="1-the-infrastructure-proxmox-lxc">1. The Infrastructure (Proxmox LXC)</h2>

<p>The stack lives in a Proxmox LXC. If you wanted to allow the container to run its own instances, you must grant the LXC specific permissions on the Proxmox host. My initial idea was to run agents with Podman, so I had to do this:</p>

<pre><code class="language-bash"># Edit /etc/pve/lxc/ID.conf
features: nesting=1,keyctl=1
</code></pre>

<p>Allows nesting and keyctl=1 allows the container to manage its own cryptographic keys and process-specific tokens which is required when running Podman in a nested environment.</p>

<h2 id="2-the-engine-llama-cpp">2. The Engine (llama.cpp)</h2>

<p>To squeeze performance out of an older i5, we compile locally to utilize the AVX (AVX1) instruction set.</p>

<pre><code class="language-bash"># Install dependencies
apt update &amp;&amp; apt install -y   git curl tmux direnv   build-essential cmake ccache   libssl-dev libsqlite3-dev liblzma-dev   zlib1g-dev libbz2-dev libreadline-dev

# Clone repo
git clone https://github.com/ggerganov/llama.cpp

# Compile
cd llama.cpp
cmake -B build -DGGML_AVX=ON -DGGML_NATIVE=ON
cmake --build build --config Release --target llama-server -j 4
</code></pre>

<p>I included ccache in the dependencies. It’s optional, but it will cut recompilation time.</p>

<p>Run a few tests:</p>

<pre><code class="language-bash"># The llama-server binary should exist
ls -lh build/bin/llama-server
</code></pre>

<h2 id="3-orchestration-persistence-tmux">3. Orchestration &amp; Persistence (TMUX)</h2>

<p>On limited hardware, high CPU load can cause SSH lag. Use TMUX inside the LXC to decouple the LLM from your connection.</p>

<p>Step-by-Step Setup:</p>

<pre><code class="language-bash"># Start Session
tmux new -s ai_bunker

# Prepare Models
mkdir -p models
wget -O models/qwen-0.5b.gguf https://huggingface.co/Qwen/Qwen2.5-Coder-0.5B-Instruct-GGUF/resolve/main/qwen2.5-coder-0.5b-instruct-q4_k_m.gguf

# Fire up the Server
mkdir -p bunker_memory
./build/bin/llama-server   \
-m models/qwen-0.5b.gguf  \
--port 8080  \
--ctx-size 2048  \
--threads 4   \
--mlock   \
--no-mmap  \
--host 0.0.0.0   \
--slot-save-path ./bunker_memory
</code></pre>

<h2 id="4-direct-inference-curl">4. Direct Inference (curl)</h2>

<pre><code class="language-bash">curl http://localhost:8080/v1/chat/completions   -H &#34;Content-Type: application/json&#34;   -d &#39;{
    &#34;model&#34;: &#34;local&#34;,
    &#34;max_tokens&#34;: 128,
    &#34;temperature&#34;: 0.2,
    &#34;messages&#34;: [
      {
        &#34;role&#34;: &#34;user&#34;,
        &#34;content&#34;: &#34;Write a simple Python function that reverses a string.&#34;
      }
    ]
  }&#39;
</code></pre>

<h2 id="5-connecting-via-browser">5. Connecting via browser</h2>

<p>Since the llama-server is bound to 0.0.0.0:8080, it’s accessible to any tool on your network. The lightest method is via a browser calling the endpoint directly.</p>

<h2 id="6-connecting-locally-wave-claude-aider">6. Connecting locally (Wave &amp; Claude &amp; Aider)</h2>

<p> Most modern terminals and AI extensions expect an OpenAI-compatible endpoint. However, these tools inject massive system prompts, and will cause context overflow errors and eat up more tokens than you might have.</p>

<p>I got it working, but it was slow. If you have a better machine and want to go ahead, you may want to configure your tools to use your local model. To do so for Wave, you can follow the documentation. Most AI terminals have similar setup. You need to edit the configuration at  ~/.config/waveterm/waveai.json or run a custom command to configure Wave:</p>

<pre><code class="language-bash">wsh editconfig waveai.json
</code></pre>

<p>Inside the JSON file, add a new entry. Based on your existing llama-server setup, use this exact structure:</p>

<pre><code class="language-json">{
  &#34;ia-bunker&#34;: {
    &#34;display:name&#34;: &#34;IA Bunker - Qwen Coder&#34;,
    &#34;display:order&#34;: 1,
    &#34;display:icon&#34;: &#34;shield&#34;,
    &#34;display:description&#34;: &#34;Local Qwen 2.5 Coder on Thinkpad T420&#34;,
    &#34;ai:apitype&#34;: &#34;openai-chat&#34;,
    &#34;ai:model&#34;: &#34;qwen-coder&#34;,
    &#34;ai:endpoint&#34;: &#34;http://192.168.1.XX:8080/v1/chat/completions&#34;,
    &#34;ai:apitoken&#34;: &#34;not-needed&#34;
  }
}
</code></pre>

<p>Replace <code>192.168.1.XX</code> with your local IP.</p>

<pre><code class="language-bash">wsh setconfig waveai:defaultmode=&#34;ia-bunker&#34;
</code></pre>

<p> I also experimented with running higher-level terminal agents (for example, Claude and Aider) inside Podman and pointing them at the bunker via environment variables.</p>

<p>On a 2 k token context, this usually results in exceed<em>context</em>size_error failures or multi-minute stalls before the first token is produced.</p>

<p>These are the commands I tried inside the container that successfully loaded Aider &amp; then Claude, but unfortunately, they just stalled with my setup (you may have more luck):</p>

<pre><code class="language-bash">podman run -it --rm \
  -v $(pwd):/app \
  -e OPENAI_API_BASE=http://172.17.0.1:8080/v1 \
  -e OPENAI_API_KEY=not-needed \
  paulgauthier/aider-full \
  --model openai/gpt-3.5-turbo \ #  used to select prompt profile
  --no-auto-commits

</code></pre>

<pre><code class="language-bash">podman run -it --rm \
  --network host \
  -v $(pwd):/app \
  -e ANTHROPIC_BASE_URL=&#34;http://127.0.0.1:8080&#34; \
  -e ANTHROPIC_AUTH_TOKEN=&#34;local-bunker&#34; \
  -e CLAUDE_CODE_DISABLE_NONESSENTIAL_TRAFFIC=1 \
  node:20-alpine \
  sh -c &#34;apk add --no-cache git &amp;&amp; npm install -g @anthropic-ai/claude-code@latest &amp;&amp; claude --model qwen-coder&#34;
</code></pre>

<h2 id="7-the-global-tunnel-tailscale-in-the-bunker">7. The Global Tunnel: Tailscale in the Bunker</h2>

<p> To access your “Bunker” from a coffee shop you could install Tailscale, CloudFlare or Zerotier directly inside the LXC. This turns your old Thinkpad into a secure, private node on your personal VPN. It’s probably worth securing the entry with some auth if you wanted to do that.</p>

<p>But it’s a topic for another post!</p>

<h2 id="ps-tweak-your-engine">PS: Tweak your Engine</h2>

<p> Initially, I was getting a painfully slow 0.25 tokens/s. To make it usable on this vintage i5, I had to aggressively tune the llama-server flags:</p>
<ul><li>—ctx-size 2048: Dropping this down helped because smaller context equals faster “thinking”.</li>
<li>-ctk q8<em>0 -ctv q8</em>0: I tried this KV-cache quantization (-ctk/-ctv), but opted out, as on a small-context (2k) Sandy Bridge CPU, the added CPU overhead often outweighs the memory savings. If I had a newer CPU, this might have helped.</li>
<li>—mlock: Essential to prevent the model from being paged to the Swap partition, which would turn the 20-minute wait into a total system hang.</li></ul>
]]></content:encoded>
      <guid>https://inti-tidball.writeas.com/build-your-own-ai-bunker-privacy-first-agents-on-outdated-hardware</guid>
      <pubDate>Mon, 26 Jan 2026 00:11:23 +0000</pubDate>
    </item>
  </channel>
</rss>