Best AI News โ€” Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

Ornith-1.0-35B GGUF update: native MTP speculative-decode graft + full serving/TTFT/long-context numbers (llama.cpp, tp=1)

Via r/LocalLlama
Sunday, Jun 28, 2026 ยท 6:35PM
Summary

Follow-up to my previous Ornith-1.0-35B Q3_K_M post. I grafted a native MTP draft head onto the IQ4_XS body (head at Q6) for self-speculative decode, single GPU, llama.cpp: 1.3-1.35x single-stream decode (172.6 -> 233.8 tok/s). Next-token distribution is byte-identical to target-only (KLD 0.0, 32/32

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories