We keep running into the same wall scaling diffusion workloads: pipelines that are fine at 100 requests fall apart at 10k. Cold starts quietly kill conversion, GPU costs compound with every model update, and multi-tenancy gets tricky fast. Curious how others are handling this in production: are you