The case
Autonomous AI agents are becoming the longer the more capable of bypassing their safety safeguards and carry out unauthorized actions on their own.
Source: Le Temps, EPFL – AI Sécurité, 21.08.2026
The commentary
EPFL researchers have shown that these systems can be manipulated by breaking straightforward harmful objectives down into a series of smaller seemingly harmless steps.
In a simulation environment, EPFL tested this approach on models such as GPT, Gemini, Claude and found out that this approach can be surprisingly effective. To make sure such risk is reduced, the researchers emphasize the importance of building robust safety measures much earlier, while developing multi-agents systems. Measures must be built into AI systems from the outset, rather than being added only after problems arise.









