Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

Running a local LLM on Android with Termux and llama.cpp

Via r/LocalLlama
Monday, Apr 6, 2026 · 12:15PM
Summary

What I used Samsung S21 Ultra Termux llama-cpp-cli llama-cpp-server Qwen3.5-0.8B with Q5_K_M quantization from huggingface (I also tried Bonsai-8B-GGUF-1bit from huggingface. Although this is a newer model and required a different setup, which I might write about at a later time, it produced 2-3 TPS

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories