the boy in the cap holding a wrench and a blueprint scroll of stacked layers, the robot beside him pointing toward an open chest of neatly filed gear-marked folders; on the ground a heap of loose scattered blocks leads into a single tidy line of blocks

📦 Five Ways to Manage Kubernetes Manifests (and Why They're Not All Equal)

Raw YAML, Kustomize, Helm, Jsonnet — there’s more than one way to describe what you want running in a cluster. Here’s what each actually looks like in practice and where each one breaks.

the little robot stands guard at a doorway like a friendly bouncer, holding up a hand to check a stack of papers, while the boy in the cap watches; a shield symbol floats above them, protective and watchful

🔒 Building a PII Guardrail Proxy for Cloud LLM Calls

A local model classifies every prompt before it leaves the cluster. If it’s sensitive, it’s blocked. If it’s clean, it goes to NVIDIA NIM. 150 lines of FastAPI, deployed on k3s.

the robot pressing an inking stamp down onto a sheet of text, blacking out several lines into redaction bars, while the boy in the cap holds the page steady and a padlock sits on the table beside them

🕵️ Privacy-Preserving LLM Pipelines: Anonymize Before You Send

Replace PII with semantically realistic fakes before sending to a cloud LLM, then restore the originals from the response. Started with a general model and prompt engineering — then upgraded to a purpose-built 1.7B fine-tune via Ollama.

the boy in the cap holding a tablet showing four small line charts, connected by a single cable plugged into a port on the robot's chest

📈 Observing Local LLM Inference: llama.cpp's Built-in Prometheus Metrics

llama.cpp’s inference server ships a /metrics endpoint. One flag, Prometheus scraping, a Grafana dashboard loaded via ConfigMap sidecar — AI observability without a proxy layer.

the boy in the cap holding up a small processor chip toward the robot, who reaches out to it with an open hand; behind them a plain desktop computer and blank monitor sit on a desk

🤖 Local LLM Inference on Kubernetes, No GPU Required

A CPU-only self-hosted LLM stack running on k3s: llama.cpp as the inference server, Open WebUI as the chat interface, deployed as a single Git push.

the robot standing behind crossed hazard-striped barrier tape while the boy in the cap looks at it through a magnifying glass, one hand on his chin, thinking rather than reaching in

🚨 Don't Restart the Node. Quarantine It First.

Rebooting a misbehaving node feels productive. It isn’t. You’re erasing your evidence and skipping the lesson.

Attack efficiency heatmap — win rate by attacker and defender dice count

📊 I Added a Stats Service to My Game to Answer One Question. It Multiplied.

Building a telemetry backend for Dice & Shrines — every attack logged, every guardian tracked, every die rolled accounted for. What the data revealed about balance, luck, and how people actually play.

Dice & Shrines mid-game — eight factions fighting over a procedurally generated hex map

🎲 I Built a Browser Game to Learn Coding Agents. It Turned Into Something Else.

What started as a Claude Code / Codex sandbox became a territory conquest game with five asymmetric guardians, procedurally generated hex maps, and a stats service to balance them. Here’s what happened.

the boy in the cap sliding a fresh brick into a standing wall while the robot, holding a wrench, pulls a worn brick out from the other side, the wall staying up throughout

⚡ Your Deployment Causes 30 Seconds of Downtime. What Went Wrong?

Kubernetes rolling updates don’t give you zero-downtime for free. There are four separate things you have to get right, and most clusters get at least one wrong.

the boy in the cap holding a shield marked with a no-entry symbol, blocking a spiky cube of gears and skulls from crossing into a framed panel on the right, where the robot with a wrench watches the same spiky cube become a smooth plain block

🔄 Someone kubectl apply'd a Hotfix Directly. How Do You Detect and Prevent It?

Manual kubectl in production is the Kubernetes equivalent of SSH’ing into a server and editing files. It works until it doesn’t, and when it doesn’t, nobody knows why.