Best AI News โ€” Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

PSA: You may not need to quantize spec draft when using MTP

Via r/LocalLlama
Friday, Jun 5, 2026 ยท 4:41AM
Summary

Using `--spec-draft-type-k q4_0 --spec-draft-type-v q4_0` might actually decrease your context size! With quantized spec draft, my context size is 83200. Without it (i.e. using the default fp16 spec draft), context size increased to 91648. I reported this in a llama.cpp discussion and am17an (the GO

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories