Built my 1st inference machine and have been tweaking models trying to get the most out of my modest hardware. I think I’m at a good place but I’m testing with my own prompts. I’ve looked into some of the popular benchmarks but I’m honestly lost. I use my models for Hermes agent mainly and a little