2 days ago there was a very cool post by u/nickl: https://reddit.com/r/LocalLLaMA/comments/1s7r9wu/ Highly recommend checking it out! I've run this benchmark on a bunch of local models that can fit into my RTX 5080, some of them partially offloaded to RAM (I have 96GB, but most will fit if you have