Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

Evaluated a RAG chatbot and the most expensive model was the worst performer. Notes on what actually moved the needle.

Via r/LocalLlama
Friday, May 15, 2026 · 12:24PM
Summary

We had a customer support RAG bot. Standard setup: ChromaDB, system prompt, an LLM doing generation. Nobody had actually measured the response quality. In the name of evaluation, I only had a keyword matching script producing numbers that looked like scores and meant nothing. I went in to fix this p

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories