Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

Running Gemma4 31b-it on vLLM 0.21.0 A100s (bad quality or what am I doing wrong)

Via r/LocalLlama
Wednesday, May 27, 2026 · 10:21PM
Summary

Okay fun time I got access to two Nvlinked A100s for some research project I benchmarked my work against the Gemma 4 31b-it available through Google, but my dataset is rather massive, so I need to run it on the "local" resources. Basically I use vLLM to run the model liteLLM to proxy to it and some

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories