Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

High E2E latency on fine-tuned Gemma 4 26B despite low TTFT [R]

Via r/MachineLearning
Thursday, May 21, 2026 · 6:12AM
Summary

Recently fine-tuned a Gemma 4 26B model, and I’m seeing surprisingly high end-to-end latency despite the effective inference footprint being much smaller (~4B-ish behavior during serving). Current setup: Model: Gemma 4 26B (fine-tuned) Engine: vLLM Quantization: FP8 Hardware: H100 Observed latency:

Continue reading the full article
Read at r/MachineLearning
www.reddit.com
Back to all stories