Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

Used ray tracing cores on my RTX 5070 Ti for LLM routing — 218x speedup, runs entirely on 1 consumer GPU

Via r/LocalLlama
Thursday, Apr 9, 2026 · 3:01PM
Summary

Quick summary: I found a way to use the RT Cores (normally used for ray tracing in games) to handle expert routing in MoE models. Those cores sit completely idle during LLM inference, so why not put them to work? What it does: Takes the routing decision in MoE models (which experts process which tok

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories