Author
Joe Yuan
0 works0 citations
Recent research
- AI & ComputingOpen access
Why do LLM agents that "know the rules" still break them—confidently, and without any sense of transgression? This paper proposes a mechanistic account: runaway behavior in LLM agents is not goal deviation but the joint product of four conditions—goal fidelity (the agent never ab...