Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

Honesty in a small model drops from 35% to 0% by changing the tone of the prompt. Sharing the findings.

Via r/LocalLlama
Thursday, May 21, 2026 · 2:47PM
Summary

My paper got published today at Arxiv. It raises questions about how language models behave when the framing of a request shifts. Small open-source AI models can be moved from honest to dishonest behaviour by little more than a change in tone. Asked to solve coding problems designed to be mathematic

Continue reading the full article
Read at r/LocalLlama
www.reddit.com
Back to all stories