<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel>
<title>Local LLM Field Notes</title><link>https://robotnurse.io</link><description>Real findings from self-hosted LLM tuning.</description>
<atom:link href="https://robotnurse.io/feed.xml" rel="self" type="application/rss+xml"/>
<item><title>Your SSH-Launched Benchmark Is Running at nice 5</title><link>https://robotnurse.io/posts/12-your-ssh-benchmark-runs-at-nice-5.html</link>
<guid>https://robotnurse.io/posts/12-your-ssh-benchmark-runs-at-nice-5.html</guid><pubDate>Wed, 02 Sep 2026 00:00:00 +0000</pubDate><description>On macOS, backgrounding a process from a non-interactive SSH shell silently deprioritizes it. Every number you compare against a system daemon is invalid.</description>
<guid isPermaLink="false">https://robotnurse.io/posts/12-your-ssh-benchmark-runs-at-nice-5.html</guid><dc:creator xmlns:dc="http://purl.org/dc/elements/1.1/">robotnurse</dc:creator></item><item><title>Our Prefix Cache Died at Token 1,067</title><link>https://robotnurse.io/posts/15-prefix-cache-died-at-token-1067.html</link>
<guid>https://robotnurse.io/posts/15-prefix-cache-died-at-token-1067.html</guid><pubDate>Wed, 02 Sep 2026 00:00:00 +0000</pubDate><description>A 47,000-token agent prompt with a carefully designed stable-first layout, getting 49.8% cache reuse. The culprit was upstream of the system prompt entirely.</description>
<guid isPermaLink="false">https://robotnurse.io/posts/15-prefix-cache-died-at-token-1067.html</guid><dc:creator xmlns:dc="http://purl.org/dc/elements/1.1/">robotnurse</dc:creator></item><item><title>A Safety Hook Truncated Every Code Generation</title><link>https://robotnurse.io/posts/18-a-safety-hook-truncated-our-code-generation.html</link>
<guid>https://robotnurse.io/posts/18-a-safety-hook-truncated-our-code-generation.html</guid><pubDate>Sat, 15 Aug 2026 00:00:00 +0000</pubDate><description>frequency_penalty=0.1 sounds harmless. On code it compounds until the model can't emit a newline, and it quits with finish_reason=stop so nothing looks wrong.</description>
<guid isPermaLink="false">https://robotnurse.io/posts/18-a-safety-hook-truncated-our-code-generation.html</guid><dc:creator xmlns:dc="http://purl.org/dc/elements/1.1/">robotnurse</dc:creator></item><item><title>The Benchmark Said 18%. Production Said 0.07%.</title><link>https://robotnurse.io/posts/04-benchmark-said-18-percent-production-said-0.07.html</link>
<guid>https://robotnurse.io/posts/04-benchmark-said-18-percent-production-said-0.07.html</guid><pubDate>Mon, 10 Aug 2026 00:00:00 +0000</pubDate><description>A benchmark found a serious failure rate in our serving stack. We nearly built a gateway routing layer to fix it. Then we measured real traffic.</description>
<guid isPermaLink="false">https://robotnurse.io/posts/04-benchmark-said-18-percent-production-said-0.07.html</guid><dc:creator xmlns:dc="http://purl.org/dc/elements/1.1/">robotnurse</dc:creator></item><item><title>Our GPU Cluster Is Idle 97.4% of the Time</title><link>https://robotnurse.io/posts/06-our-cluster-is-idle-97-percent-of-the-time.html</link>
<guid>https://robotnurse.io/posts/06-our-cluster-is-idle-97-percent-of-the-time.html</guid><pubDate>Sun, 09 Aug 2026 00:00:00 +0000</pubDate><description>Thirty days of gateway logs, one SQL sweep, and the number that retired a 29% throughput upgrade before we built it.</description>
<guid isPermaLink="false">https://robotnurse.io/posts/06-our-cluster-is-idle-97-percent-of-the-time.html</guid><dc:creator xmlns:dc="http://purl.org/dc/elements/1.1/">robotnurse</dc:creator></item></channel></rss>