Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

Mid research got me thinking what about reversed alignment, would trained "bad" model exhibit"good" behavior later and/or secretly [D]

Via r/MachineLearning
Tuesday, Jul 7, 2026 · 7:08PM
Summary

late night thoughts as I was working on my paper that is about specific behavior that arises from RHLF, it got me thinking what if train a model in an environment where bad behavior is rewarded: deception, selfishness, harmful behavior etc. and then find it occasionally and/or secretly exhibit good

Continue reading the full article
Read at r/MachineLearning
www.reddit.com
Back to all stories