[{"content":"I was so excited that Suitcase AI had done real work for me. 604,649 tokens, the GPU staying around 63°C, unit tests passing, code committed, guard rails all showing green. Feeling pretty good about buying a GX10. Unfortunately, a warm GPU does not infer a GPU that produces code.\nThe Starter Task: spark-to-k8s First, on my clean Suitcase AI repository, I needed to get default benchmarks for the chosen model, nvidia/llama-3.3-nemotron-super-49b-v1.5. My smoke test confirmed the model was responding, but for rigorous baseline metrics, I ran llama-benchy and posted to Spark Arena.\nSpark Arena grounds its results on a recipe.yaml that defines the model and the settings used for the inference engine. Because my GX10 is joined directly to the Suitcase AI Kubernetes cluster rather than running standalone, I couldn\u0026rsquo;t just use sparkrun out of the box. To unblock the benchmark, I had my agents whip up a rough prototype, spark-to-k8s, to take my recipe.yaml and transform it into the kustomize manifest part file I needed. Nemotron 49B on my Suitcase AI benchmarked with strong numbers across the 28-task matrix, sustaining up to 1,417 prefill tok/s and 31.7 decode tok/s (view the Spark Arena benchmark results).\nWith baseline performance locked in, I was ready to supply Camp Colt (my extension to Gas City) with real work that mattered to me. Productionizing that rough spark-to-k8s prototype was the perfect starter task.\nThe assignment was straightforward: extract that early prototype out of Suitcase AI into a dedicated tools repository, write proper unit tests, and generalize it. Instead of being hardcoded to one or two specific models, it needed to reliably compile any generalized YAML recipe into clean cluster manifests.\nFirst Day on the Job: Breaking the Tools Of course, I had all the normal problems that you would expect when working with a new staff at a new location with new tools. Right out of the gate, the tools failed. vllm was not parsing tool calls appropriately, stalling the agents for an hour.\nThe problem was that the default Nemotron 49B tool parser did not support streaming mode (HTTP 501: Tool calling is not supported in streaming mode!). So, I had the agents write code that handled this for me: a custom streaming tool-parsing plugin. My agents grabbed that code, mounted it into the inference engine via a ConfigMap, and ran our smoke tests. They passed cleanly, and we were off and running. (The full engineering breakdown is documented in the pilot overview report).\nMy Guardrails Were in Place I wasn\u0026rsquo;t flying blind. I had three explicit operational guardrails defined for the pilot:\nZero-Fallback: 100% of inference must run on local silicon with zero calls routed to cloud models. Empirical Outcome: Code must compile, unit tests must pass, and manifests must be generated. Autonomous Local Delivery: The local agent workers must drive the migration, refactoring, and commit without human intervention. As far as my dashboard was concerned, everything was ✅. The vLLM gateway showed a 0.00% cloud fallback rate across more than 604,000 tokens. The Go unit tests passed cleanly. A commit was in the git log. The migration was done.\nThe results are very satisfying: full autonomous software generation using sovereign local inference. I sat down and wrote the article with all its ups, downs, and drama. All that was left to do was choose an image.\nI wanted to visualize the longest query1, the 66-minute turn where the code was finally committed. I thought I would use actual thought traces from that query. However, everything fell apart when I dug into that hour-long prompt response.\nRequest latency showing the monolithic reasoning turn executing on local silicon. The agent was trapped in a reasoning loop over a quote mismatch in a bash heredoc, spinning in circles while producing zero output. Yet somehow, the code had still been committed.\ncolt-utils/gc.implementation-worker-3 REASONING LOOP TRACE \u0026lt;think\u0026gt;\nThe previous command had a JSON parsing error due to an unterminated string. Let me examine why... In JSON, newlines must be escaped as \\n, and double quotes must be escaped as \\\"... \u0026lt;/think\u0026gt; [FAIL] Invalid input for tool bash: JSON parsing failed: Unterminated string \u0026lt;think\u0026gt;\nThe user provided a long script that's supposed to be executed using the bash tool. But every time I try to call the bash command with that script, there's a JSON parsing error about an unterminated string. Looking at the script, is there a quote issue?... \u0026lt;/think\u0026gt; [FAIL] Invalid input for tool bash: JSON parsing failed: Unterminated string After an hour, the supervising agent running Antigravity that assigned this task finally gave up, wrote the code itself on the cloud, and committed it for me.\nFixing the Trap and Tuning the Rig Of course, that heredoc problem had already been resolved upstream in Gas City via PR #204 in gascity-packs. I was just pinned to an older version during the pilot. Updating the harness resolved the issue by moving away from brittle inline heredocs in favor of clean CLI calls.\nDigging deeper into the telemetry also surfaced why my KV cache hit rates weren\u0026rsquo;t where they should have been on my Grace Blackwell silicon. Gas City was injecting a dynamic timestamp at the very beginning of the session start prompt. In an inference engine using Radix prefix caching (like vLLM or NIM), any changing token at the start of a prompt busts the cache for the entire context, forcing the GPU to recompute prefill from scratch on every new session. I opened issue #5732 in Gas City to get prompt timestamps moved out of the prefix header so prefix caching can do its job.\nBut the biggest revelation wasn\u0026rsquo;t about prompt formatting or cache invalidation. It was about how I measured agentic systems.\nFailed Smoke Tests and Guardrails My smoke test, which tested tool calling, didn\u0026rsquo;t test it in a way that mattered to me: it tested in batch mode instead of streaming mode. The fix was easy enough, using custom tool parsing code my agents had written earlier.\nThe guardrail for local execution also tested the wrong thing. Yes, “worker” agents only executed on the local LLM. And yes, good results (code) were produced. But the workers never did work I cared about. Instead the supervisor did the work that matters on the cloud. I assumed token draw meant code output. I had confused movement with progress.\nOf course, this is why I paid for a GX10 and built Suitcase AI: to learn, get better, and share.\nWhat\u0026rsquo;s Next This week, I turned the local model loose on issue #5732 in Gas City to fix prompt timestamp cache invalidation.\nNext, I will experiment with selective reasoning: testing different agents in the hierarchy with and without thinking tokens, tuning effort levels, and evaluating alternative model weights.\nI will also do a deep dive into my agent staff architecture, detailing why I use military staff structure (S-3, S-6) to orchestrate autonomous development.\nJoin the Discussion Discuss on X / Twitter Join the conversation on LinkedIn Why does the chart show a 2-hour spike when the prompt actually took 1 hour? The actual inference turn ran for 66.01 minutes (3,960 seconds), generating 15,511 tokens with zero queue wait. The 2.05-hour (7,392s) spike on the dashboard is a known mathematical artifact of Prometheus histogram_quantile(0.95, ...). Because the vLLM histogram had a wide bucket gap between 32 minutes (le=\u0026quot;1920.0\u0026quot;) and 128 minutes $1920 + (7680 - 1920) \\times 0.95 = 7392\\text{s}$. (See the full investigation report on GitHub for the mathematical derivation and forensic breakdown).\u0026#160;\u0026#x21a9;\u0026#xfe0e;\n","permalink":"https://showmetheinference.com/posts/604649-tokens-and-not-a-single-line-of-code/","summary":"Suitcase AI ran 604,649 tokens and guardrails read green. But the local agent was stuck in a thought loop while a cloud supervisor wrote all the code.","title":"604,649 Tokens and Not a Single Line of Code"},{"content":"When you leave Google, the first thing you lose is the illusion of infinite, free compute. Suddenly, every token has an invoice attached to it. Watching Steve Yegge burn $5,000 a month on Claude tokens for Wyvern was inspiring. I—alas—am not Steve Yegge, burning $5,000 a month on a video game with active players that I\u0026rsquo;ve been working on for 30 years. However, I did drop almost $5,000 on an ASUS Ascent GX10 to enable me to continue my AI learning without blowing my monthly budget on tokens.\nI wanted a self-contained, portable, sovereign compute node rather than a giant rack in a server closet. Thus was born Suitcase AI—because the name reminds me of a suitcase nuke, so why not?\nThe general idea is a 10\u0026quot; minirack packed with the inference, storage, and general compute needed to handle real problems with a predictable cost: Unit 02 (the ASUS GX10 on top) dedicated to sovereign GPU inference, and Unit 01 (below) handling the control plane, hypervisor, and utility services.\nThe Suitcase AI rig: A 10-inch minirack with the ASUS Ascent GX10 inference unit (Unit 02) on top and the local control plane node (Unit 01) below. I spent several weeks just playing with different models and learning all kinds of things. For example, oh my gosh, the models are huge and take a long time to download—then as much as 10 minutes just to get the model loaded into memory! This means I really need to plan more instead of just trying something new because it\u0026rsquo;s there.\nAfter three weeks of conflicting environments, orphaned model weight caches, and half-baked shell scripts, I finally decided I had made enough of a mess and started over with a clean, reproducible setup for Suitcase AI. My suitcase-ai repository is where all my infrastructure as code lives. Of course, I couldn\u0026rsquo;t resist using Kubernetes, Vault, and all the other tools required for a reasonable homelab. But I was putting more time into the infrastructure than actually learning about AI. The infrastructure is up, and I can now turn my attention back to AI.\nThe First Experiment: Hybrid Gas City For my first experiment, I\u0026rsquo;m going to look at how I can use the Gas City harness with frontier models for human interaction, but rely on local models for the actual coding and other smaller tasks. I experimented with this previously and got stuck mostly on the watchdogs and other mechanisms needed to push things along; the smaller models seemed to get stuck on the simplest tasks.\nIn \u0026ldquo;Fences, Not Sandboxes,\u0026rdquo; Steve talked about how his agents wound up creating a large legal system to try and codify answers to previous mistakes. That kind of overhead will completely overwhelm local AI.\nSteve was using all frontier models that have 200k+ context windows and unlimited cloud compute; they can act like medieval parliamentarians drafting bylaws. But a local model running on a maximum of 80GB of memory needs lean, deterministic instructions. If you bury a local model under a mountain of procedural \u0026ldquo;fences\u0026rdquo; and legal rules, it stalls under its own weight.\nIn fact, Steve later found that even Fable seemed to legislate itself into a corner.\nNext week I will be back at it, and I\u0026rsquo;ll let you know which model I selected and how it worked on a sample task assigned to Gas City.\nJoin the Discussion Discuss on X / Twitter Join the conversation on LinkedIn ","permalink":"https://showmetheinference.com/posts/why-i-spent-5000-on-hardware-instead-of-tokens/","summary":"Why I bought a $5,000 local inference box instead of burning a monthly cloud token budget, and how local hardware changes the way you build AI agents.","title":"Why I Spent $5,000 on Hardware Instead of Tokens"},{"content":"There was a time when, if your grandmother gave you a handmade quilt, it was a valuable piece of property. You took care of it and maybe even treasured it. A quilter\u0026rsquo;s skill was valued by the community.\nThen Walmart started selling blankets for $20. They\u0026rsquo;re not as nice as what Grandma made, but they\u0026rsquo;re $20. You can buy a bunch of blankets, and you throw them away when they get threadbare or stained.\nStitching Code — Generated with Google Nano Banana 2. Image prompt A detailed oil painting in the style of Norman Rockwell. A mid 30s, focused man with black rim glasses, wearing a tie and short sleeved white shirt. We view him from the side at his standard desk with an intricate quilt on a small wooden frame on top of the desk. He is meticulously stitching code into the quilt line by line using a needle and thread. The code is embroidered in a vintage 1950s monospace font. The quilt features patterns of geometric shapes and computer punch cards. The scene busy 1950 office with warm, soft sunlight from a side window. Expressive facial features, nostalgic Americana aesthetic, rich textures. I could lament the fact that I used to write beautiful code—that my software was a piece of art meant to be maintained for years or decades. I could lament, but that\u0026rsquo;s not my point, and that\u0026rsquo;s not the problem I think tech is facing. I\u0026rsquo;m over the fact that my ability to write beautiful code is about as valuable as the ability to make a beautiful quilt. That is what machines do. That is what AI does.\nThe problem I think we\u0026rsquo;re really facing is\u0026hellip; well, I\u0026rsquo;ll tell a story.\nMy youngest son was getting married, and we rented a large VRBO house for all of my family—seven kids, spouses, and grandchildren. That was a lot of people, and the house, while large and well-appointed, just didn\u0026rsquo;t have enough blankets.\nSo, I went to Walmart and bought three $20 blankets.\nAfter the wedding, which was beautiful, it was time to go home. We packed up those blankets and brought them home. The problem was, our linen closet was full. I already had all the blankets I needed. I didn\u0026rsquo;t have a spot to store those extra blankets. Really, I probably should have just thrown them away, but they were good blankets and I had just bought them.\nSo\u0026hellip; my wife looked at the linen closet and asked if we should throw away the quilt her grandmother made her.\nAnd that is the problem we\u0026rsquo;re facing with AI coding.\nIt\u0026rsquo;s practically free to generate, so the impulse is, \u0026ldquo;Why not? Have it build that feature.\u0026rdquo; But cheap creation ignores the hidden tax: every line of code must be stored, audited, and maintained. We haven\u0026rsquo;t learned how to manage hundreds of bespoke programs, or how to triage what AI code is actually worth keeping.\nI think our industry is going to have to get used to viewing code as disposable—like a $20 blanket.\nIn my own work with local AI, that means exploring a few practical questions:\nHow do we design a staff of agents that actively decides what code to keep and what to throw away? How do we protect what actually has value—like tests and APIs—while letting the implementation be completely disposable? I\u0026rsquo;ll let you know what I learn along the way.\nJoin the Discussion Discuss on X / Twitter Join the conversation on LinkedIn ","permalink":"https://showmetheinference.com/posts/coding-like-quilting/","summary":"I\u0026rsquo;m over the fact that my ability to write beautiful code is about as valuable as the ability to make a beautiful quilt.","title":"Coding, Like Quilting, Used to Be a Valuable Skill"},{"content":"I have always said if I am not coding, I am not happy. If my job did not involve enough coding, I would end up writing code late into the night.\nYet I have not written more than a few lines of code in more than a year—and I am having more fun solving problems than I have had in a decade.\nYears ago, I ran a blog called Show Me the Code (a nod to Linus Torvalds\u0026rsquo;s famous quote). Today, I add value by choosing the right models, running them efficiently, and designing a staff of agents that produce results.\nSo now I say: Show Me the Inference.\nThe shift from line-by-line coding to directing autonomous agents. (Illustration generated with Gemini) Image prompt An editorial split illustration in the tactile style of vintage magazine art and modern technical linework. Left side (1990s): A warm, nostalgic developer desk with a bulky beige CRT monitor displaying green monospaced Java code, a thick vintage textbook titled \"Effective Java\", a classic white ceramic mug with the red-and-blue \"Java\" logo, a stack of 3.5-inch floppy disks, and a beige trackball mouse. Right side (Modern): A contemporary dark-mode developer workspace with a printed research paper titled \"Attention Is All You Need\" and a sleek monitor displaying: User: 'Add 2-factor authentication to checkout.' Code Agent: 'Which standard should we use? OAuth 2.0, WebAuthn, MFA, or SMS-OTP?' What I am exploring You can expect me to talk about:\nLocal AI: Running local inference, managing GPU infrastructure, and deploying private, sovereign models. Agent Staff: How I use multi-agent systems to both write code and interact with the world in direct, practical ways. Sea Stories: Real-world lessons from decades of systems design, infrastructure, and operations. Welcome. Let\u0026rsquo;s dig in.\n","permalink":"https://showmetheinference.com/posts/show-me-the-inference/","summary":"I have not written more than a few lines of code in more than a year—and I am having more fun solving problems by directing an agent staff than I have had in a decade.","title":"Show Me the Inference"},{"content":" Nick Chalko Ex-Google, Ex-Apple, Retired Lt Col USMCR.\nI have spent decades writing and designing software, building distributed systems, and leading engineering teams.\nToday, my focus has shifted to local AI compute, sovereign models, and directing autonomous multi-agent staffs.\nTo read why I stopped writing code by hand and what this site is about, check out the opening essay on why I shifted from Show Me the Code to Show Me the Inference.\nThis site is where I share what I\u0026rsquo;m learning along the way—from private inference architectures to systems design lessons and other sea stories.\nConnect X / Twitter: @chalko LinkedIn: linkedin.com/in/chalko GitHub: github.com/chalko RSS: showmetheinference.com/index.xml ","permalink":"https://showmetheinference.com/about/","summary":"About Nick Chalko and Show Me the Inference.","title":"About"}]