I think the biggest unlock for local models over the next year is not another benchmark jump. It’s making the whole stack feel boring and dependable. Right now the average workflow still has too many sharp edges: model format mismatch, VRAM roulette, broken tool calling, inconsistent evals, and setu