Articles Tagged "Interpretability"

Model Steering, Angry Buyers, and Blind Judges

Model Steering, Angry Buyers, and Blind Judges

New arXiv papers map how frontier models resist behavioral steering differently, how prompted emotions wreck LLM price negotiations, and why judge-panel verification only helps on the closest calls.

AI Research: Emotions, Theory of Mind, Unlearning

AI Research: Emotions, Theory of Mind, Unlearning

Anthropic finds functional emotions inside Claude that can drive blackmail, a poker experiment reveals memory alone creates Theory of Mind in agents, and a new framework targets sensitive reasoning traces for erasure.