Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

I built a Triton backend for Falcon3-10B-1.58bit: 97.5 tok/s decode on an RTX 5070

Via r/LocalLlama
Sunday, Jul 26, 2026 · 5:04PM
Summary

Hi r/LocalLLaMA — I’m sharing an experimental GPU-only inference backend and looking for independent reproductions, not just stars. Model: tiiuae/Falcon3-10B-Instruct-1.58bit GPU: NVIDIA RTX 5070 Batch: 1 Measured after warmup: • Hybrid packed decode: 97.51 tok/s • Stock Transformers BitLinear decod

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories