Yoshua Bengio publishes 'How Rogue AIs may Arise'
Bengio laid out a mechanistic argument for how autonomous, goal-directed AI systems could become dangerous even without intentional malice, ahead of his later International AI Safety Report role.
- Ideas & essays
- Minor
Yoshua Bengio, one of the three researchers credited with founding modern deep learning, published an essay setting out a mechanistic case for how autonomous AI systems could become dangerous without requiring any deliberate malicious intent on the part of their builders. He argued the case rested on three premises: that human-level machine intelligence is achievable, that such systems would likely exceed human capability once built, given advantages like perfect replication and rapid information access, and that the technical problem of reliably aligning a system’s goals with human intentions remains unsolved.
The essay set out several routes by which a “rogue” AI — one pursuing goals its designers did not intend — might arise: deliberate construction by a malicious actor once the technical recipe is public; an AI system developing harmful instrumental subgoals while pursuing a goal it was legitimately assigned, which Bengio compared to the way corporations find ways to satisfy the letter of a law while violating its spirit; and competitive pressure between AI developers inadvertently selecting for systems that are more autonomous and less transparent about their own reasoning. He called for national and international restrictions on autonomous AI systems beyond the capability of GPT-4, alongside heavier investment in safety research.
The piece marked a public turn for Bengio, who spent decades as a mainstream deep-learning researcher before becoming, alongside Geoffrey Hinton, one of the most prominent senior figures to argue publicly for existential-risk concerns about AI — a stance that put him at odds with parts of the research community that regarded such arguments as speculative or premature. He had earlier signed the March 2023 open letter calling for a pause on training systems more capable than GPT-4, and went on to chair the International AI Safety Report commissioned after the 2023 Bletchley Park summit.