Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

SycoFact 4B - Open model for detecting sycophancy & confirmation of delusions, 100% on psychosis-bench, generates feedback for model training, trained without human labels

Via r/LocalLlama
Monday, Mar 30, 2026 · 6:07PM
Summary

I published a model you can use now to help detect sycophantic AI responses. It rejects 100% of the sycophantic delusion affirming responses from psychosis-bench. It also does well on the AISI Harmful Advice, PKU-SafeRLHF, and safety subsets of RewardBench. It's only 4B parameters, so it's of partic

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories