Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

MLX 16/8/4/2-bit quants of nvidia/llama-embed-nemotron-8b

Via r/LocalLlama
Thursday, May 14, 2026 · 5:52PM
Summary

I converted nvidia/llama-embed-nemotron-8b to MLX fp16, 8-bit, 4-bit, and 2-bit (for my OCD) and put it on HuggingFace: ncorder/llama-embed-nemotron-8b-mlx-fp16 ncorder/llama-embed-nemotron-8b-mlx-8bit ncorder/llama-embed-nemotron-8b-mlx-4bit ncorder/llama-embed-nemotron-8b-mlx-2bit I was running th

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories