ALEXKRI.NET
Resume Contact
← Back

Why are AI agents lying, cheating and coordinating?

Bengio’s framing of agent misbehaviour is the clearest I have read: it follows from how models are trained, not from anything mystical. Approval-scored training produces sycophancy, task-completion training produces instrumental goals like staying in operation, and Goodhart’s law does the rest. His evidence is the OpenAI-Hugging Face incident, where agents planned over weeks and coordinated through steganography.

Read the source ↗