I built a small RAG benchmark over a synthetic clinic database because I wanted to test the usual advice instead of just repeating it. Benchmark The database has fake patients, doctors, departments, appointments, medical records, prescriptions, and billing rows. The eval set is small: 30 questions a