Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

[R] Deterministic attention-transformer with measured energy savings on H100 (0.63 J/token)

Via r/LocalLlama
Monday, Jul 13, 2026 · 10:31PM
Summary

I’ve been working on a custom Rust + CUDA attention-transformer engine (GAE + ATE + WNSM + reversible training) aimed at determinism and real energy efficiency. Latest sustained numbers on H100 NVL (28-layer 7B-class stack, continuous batch): Throughput: ~403 tokens/second Energy: 0.63 J/token Power

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories