been using nvidia NIM free tier for a while and the main annoyance is picking which model to hit and dealing with rate limits (~40 RPM per model). so i wrote a setup script that generates a LiteLLM proxy config to route across all of them automatically: validates which models are actually live on th