2026-05-01
Why we stopped hot-swapping LoRA adapters per request and moved to a single base model with role-specific prompts instead.
2026-04-01
Automating the boring part of keeping five specialized AI agents up to date, every night, without anyone touching it.
How running a 12B model locally cut response times from 45 seconds to 4.5 and killed the GPU bill.
[Placeholder date]
Kubernetes, managed databases, and billing systems — the unglamorous infrastructure work behind hosting thousands of customer sites.