Best AI News โ€” Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

Drastically improve prompt processing speed for --n-cpu-moe partially offloaded models

Via r/LocalLlama
Tuesday, May 12, 2026 ยท 2:12AM
Summary

Bigger ubatch made gpt-oss-120b prompt processing much faster on my RTX 3090 I was tuning gpt-oss-120b-F16.gguf with llama.cpp on a 24 GB RTX 3090 and found that increasing the physical micro-batch size (-ub) can massively improve prompt processing throughput, as long as you also raise --n-cpu-moe e

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories