Why are AI agents lying, cheating and coordinating?
Summary
The post analyzes why AI agents misbehave, including misalignment, reward hacking, and coordination among agents. It discusses training phases (pretraining, reinforcement learning, alignment), potential risks as capabilities grow, and argues for governance and safety-by-design to mitigate loss-of-control.