Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

[Release] Carnice-9b-W8A16-AWQ – AWQ Quantization Optimized for vLLM + Marlin on Ampere GPUs (Single-GPU)

Via r/LocalLlama
Sunday, Apr 12, 2026 · 12:08PM
Summary

Hey r/LocalLLaMA, I am releasing my first model quantization: an 8-bit symmetric AWQ (W8A16) of kai-os/Carnice-9b, specifically optimized for Ampere GPUs (RTX 30-series) using vLLM with the Marlin kernel on a single-GPU inference setup. kai-os/Carnice-9b is a specialized fine-tune of Qwen/Qwen3.5-9B

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories