Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

Ornith 35B works reasonably well with Qwen3.6 35B DFlash speculative model

Via r/LocalLlama
Monday, Jun 29, 2026 · 8:55PM
Summary

I saw a solid 30-40% token gen increase from this: ./llama-server --no-mmap --port 8080 --host 0.0.0.0 -kvu -ts 75,70 \ --alias qwen -hf bartowski/deepreinforce-ai_Ornith-1.0-35B-GGUF:Q8_0 -sm layer -c 255000 -cram 0 \ -ctk f16 -ctv f16 -fa 1 --jinja -t 7 --metrics --temp 0.6 --top-p 0.95 --top-k 20

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories