Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

how to run gemma-4-12b-it-qat-w4a16-ct in vllm or any version quantized of the model

Via r/LocalLlama
Monday, Jun 8, 2026 · 12:16AM
Summary

when running by using transformers it runs by using vllm some weird error come up plese can any body share the command of running it on vllm ?

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories