Run AI in production: latency budgets, cost per call, caching, fallbacks, and what you do when it breaks at 3am - taught through incidents of our own.