Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

Training a number-aware embedding model + Text JEPA doesn't work too well + Text auto-encoders have a strange frequency bias [R][P]

Via r/MachineLearning
Wednesday, May 13, 2026 · 11:55AM
Summary

Hi guys! I've spent 1y trying to predict company growth from the full text of their 10-k filings. It completely failed. But I've had a lot of fun playing with encoder transformers and making them good at numbers (bypassing the tokenizer/prediction head for numbers). I've MLM-trained a modified Moder

Continue reading the full article
Read at r/MachineLearning
www.reddit.com
Back to all stories