Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

llama: limit max outputs of `llama_context` by am17an · Pull Request #23861 · ggml-org/llama.cpp

Via r/LocalLlama
Monday, Jun 1, 2026 · 3:29PM
Summary

Overview continue #23764, this PR only reserves logits space for n_seqs when possible. With -ub 2048 and MTP, this saves another 1.2GB of VRAM for me. I've tested llama-perplexity also and it seems to work fine. But maybe there is a better API, putting up as a draft for now According to me an API in

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories