Show Me the Inference

Local AI, GPU compute architecture, and autonomous agentic staffs by Nick Chalko.

End-to-End Request Latency chart showing the monolithic reasoning turn on local silicon alongside the agent thought loop.

604,649 Tokens and Not a Single Line of Code

Suitcase AI ran 604,649 tokens and guardrails read green. But the local agent was stuck in a thought loop while a cloud supervisor wrote all the code.

September 29, 2026 · 6 min · 1249 words · Nick Chalko
The Suitcase AI 10-inch minirack with Unit 01 (control) and Unit 02 (ASUS Ascent GX10 inference), sitting next to the retail ASUS Ascent GX10 box.

Why I Spent $5,000 on Hardware Instead of Tokens

Why I bought a $5,000 local inference box instead of burning a monthly cloud token budget, and how local hardware changes the way you build AI agents.

September 17, 2026 · 3 min · 573 words · Nick Chalko
A 1950s engineer with glasses stitching code line-by-line into an intricate quilt at his desk

Coding, Like Quilting, Used to Be a Valuable Skill

I’m over the fact that my ability to write beautiful code is about as valuable as the ability to make a beautiful quilt.

September 2, 2026 · 3 min · 580 words · Nick Chalko
An editorial split illustration showing green Java code on a 1990s CRT monitor beside the classic Effective Java book and Java mug on the left, contrasting with a modern dark-mode terminal displaying the dialog User: 'Add 2-factor authentication to checkout.' and Code Agent: 'Which standard should we use? OAuth 2.0, WebAuthn, MFA, or SMS-OTP?' above the 'Attention Is All You Need' paper on the right.

Show Me the Inference

I have not written more than a few lines of code in more than a year—and I am having more fun solving problems by directing an agent staff than I have had in a decade.

August 28, 2026 · 2 min · 290 words · Nick Chalko