Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

llamacpp patch - DeepSeek V4 Flash running with full 1M token context locally on RTX 5090

Via r/LocalLlama
Thursday, Jul 2, 2026 · 11:54PM
Summary

Wanted to try running DeepSeek V4 Flash locally but found it asking for absurd amounts of VRAM at higher context lengths (~256GB at 1M). Turned out the DSA lightning indexer lacks proper llamacpp support. Did a bit of digging and there's an upstream PR to address the issue (shoutout u/fairydreaming,

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories