Writeup of an emergent behavior I observed in production. Posting here for methodological critique and pointers to related work. Context: a conversational AI system (single-tool tool schema with 5 enumerated action types, each with explicit description). Observed across ~2,400 messages, the model us