Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

Using Gemma 4 E4B with the LiteRT engine - ~2.4x speedup over Q4 GGUF in text generation, image processing roughly the same

Via r/LocalLlama
Tuesday, Jun 2, 2026 · 5:46PM
Summary

I know there is a PR in llama.cpp to support MTP for the 26b and 31b versions of Gemma 4, but as far as I can tell there is nothing yet for the E2B and E4B models. Using Hermes Agent, I had it set up Gemma 4 E4B in Google's Lite RT format, and then write a Python wrapper around it to create an OpenA

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories