Best AI News — Updated Every 3 Hours
Best
AI
News
Story Page
← All Stories
Home
→
Community
→
Story
Community
llama: use f16 mask for FA to save VRAM by am17an · Pull Request #23764 · ggml-org/llama.cpp
Via
r/LocalLlama
Friday, May 29, 2026 · 7:49AM
Summary
now you can download more VRAM ;) (by downloading new llama.cpp version)
Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Related in Community
Hopfield Memory in VLA [R]
r/MachineLearning
Qwen 3.6 27B overdoing it
r/LocalLlama
Building a monokernel for LLM inference on AMD MI300X - up to 3,300 output tokens/s per request [P]
r/MachineLearning
How do I make MTP work in llama-server?
r/LocalLlama
Use HTML as the primary chat language for your agents so they can draw diagrams
r/LocalLlama
More from Best AI News
Amazon kills internal AI leaderboard after employees gamed it with pointless tasks
The Decoder · Industry & Money
Behavior-Induced Mirror-Prox Temporal-Difference Learning for Faster Off-Policy Prediction
ArXiv cs.AI · Papers
Behavior-Aware Auxiliary Corrections for Off-Policy Temporal-Difference Prediction
ArXiv cs.AI · Papers
The Cognitive Categorical Transformer: Category-Theoretic Inductive Biases for Language Modeling
ArXiv cs.AI · Papers
Back to all stories