Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

Running SmolLM2‑360M on a Samsung Galaxy Watch 4 (380MB RAM) – 74% RAM reduction in llama.cpp

Via r/LocalLlama
Thursday, Apr 2, 2026 · 8:22AM
Summary

I’ve got SmolLM2‑360M running on a Samsung Galaxy Watch 4 Classic (about 380MB free RAM) by tweaking llama.cpp and the underlying ggml memory model. By default, the model was being loaded twice in RAM: once via the APK’s mmap page cache and again via ggml’s tensor allocations, peaking at 524MB for a

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories