Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

Orthrus: Memory-Efficient Parallel Token Generation via Dual-View Diffusion [R]

Via r/MachineLearning
Friday, May 15, 2026 · 5:21PM
Summary

Paper: https://arxiv.org/abs/2605.12825 Code: https://github.com/chiennv2000/orthrus Disclosure: co-author. Idea: Inject a trainable diffusion attention module into each layer of a frozen AR Transformer. Both heads share one KV cache. Diffusion head projects K=32 tokens in parallel; AR head verifies

Continue reading the full article
Read at r/MachineLearning
www.reddit.com
Back to all stories