Ran 10 realistic product tasks (classification, RAG QA, multi turn conversation, an agentic plan then execute task, etc) against the live APIs of OpenAI, Anthropic, Gemini and Kimi, using each provider's cost optimized tier. Total cost spread was 10.6x despite published rates differing by only 2x, m