Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

torch-nvenc-compress: GPU NVENC silicon as a PCIe bandwidth multiplier — PCA + pure-ctypes Video Codec SDK wrapper. Parallel-path overlap measured at 67% of theoretical max on a real GEMM + encode workload. [P]

Via r/MachineLearning
Sunday, May 3, 2026 · 10:43PM
Summary

I've been working on the consumer-multi-GPU PCIe bottleneck — Nvidia removed NVLink from the 4090/5090, and splitting a 70B model across two consumer cards drops you to ~30 GB/s over PCIe peer-to-peer. Spent the last few months building a Python library that uses the GPU's otherwise-idle NVENC/NVDEC

Continue reading the full article
Read at r/MachineLearning
www.reddit.com
Back to all stories