I've been trying to keep as much of my AI workflow local as possible. One thing I couldn't find was a good way to work with videos without repeatedly sending them through a multimodal model. My use case is mostly screen recordings, bug reports, product demos, and Loom videos. The approach I ended up