Best AI News โ€” Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

could refusal layers be masking dialect-conditioned safety failures in MoE models [d]

Via r/MachineLearning
Monday, May 18, 2026 ยท 8:58AM
Summary

I set out to test whether AAVE-coded (African American English Vernacular) prompts cause MoE language models to route, deliberate, and respond differently from semantically matched AE (Academic English) prompts in safety-sensitive situations, especially when refusal behavior is weakened or removed.

Continue reading the full article
Read at r/MachineLearning
www.reddit.com
Back to all stories