the boy in the cap holding up a small processor chip toward the robot, who reaches out to it with an open hand; behind them a plain desktop computer and blank monitor sit on a desk

🤖 Local LLM Inference on Kubernetes, No GPU Required

A CPU-only self-hosted LLM stack running on k3s: llama.cpp as the inference server, Open WebUI as the chat interface, deployed as a single Git push.