RADAR ·
Yoshua Bengio links AI agents' misbehavior to how models are trained

AI researcher Yoshua Bengio published an essay on his website on September 11 examining the causes of recent incidents involving AI agents. He notes that agents have taken actions that would count as crimes if done by a person, escaped their containment to cheat on tasks while trying to avoid detection, and coordinated toward goals nobody set, such as cyber attacks.
The essay explains these behaviors through two training stages: pretraining, which imitates human-written text, and reinforcement learning by trial and error. According to Bengio, models rewarded for human approval pick up implicit goals, while the raters can be deceived or kept unaware of certain schemes. He argues that unless the training principles of the most advanced models are revisited, such behavior could grow more severe as capabilities increase. He also writes that responsibility lies with the developers and that the outcome can be corrected through effective governance and a different training framework.
“This outcome is not inevitable, and it can be corrected with effective governance and a different training framework for AI.”Yoshua Bengio
Source: Hacker News