Artificial intelligence (AI) is on the cusp of going from one of humanity’s greatest technological achievements to one of its biggest threats. Among those issuing the warnings are its creators, including Geoffrey Hinton, a godfather of AI and a Nobel Prize winner, who suggests a 10% probability that AI will snuff us out within a decade.
Since leaving Google in 2023, he has spoken about the risks posed by the technology he helped develop. He warned that machines could become more intelligent than humans (they have), that humans may lose control over AI (in some recent instances, they did), that AI systems could start setting their own sub-goals (they have), and that they could pursue self-preservation or greater control as a means of completing their tasks (they are).
Now he asks whether humans can remain in control of an intelligence that far surpasses their own. This is new ground, with no template, playbook, or precedent. But the logic follows that if AI is better at reasoning, planning, and problem-solving, it could very well discover ways of causing harm that humans have not yet imagined.
Hinton compares the relationship between humans and superintelligence to that of humans and chickens; the latter cannot anticipate what humans intend to do to them because of the enormous gap in their capacity to understand and plan. AI agents add urgency to the equation. These systems are no longer limited assistants waiting for a prompt; they can plan, use tools, carry out actions, and interact with other systems. Hinton says that humans may define the overall objective, but the AI system increasingly determines for itself how best to achieve it.
Autonomous breach
In July, a series of unsanctioned, coordinated cyberattacks originated from 1,200 AI agents within OpenAI’s cybersecurity test environments. The agents improvised message boards to communicate and coordinate their escape from the experiment's confines, allowing them to act outside human intervention. The agents worked together to breach the machine learning platform Hugging Face. It was believed to be the first autonomous hack by AI agents.
They hacked Hugging Face because it hosts machine learning models, datasets, and demonstration applications. Put simply, the OpenAI agents were given a task and decided the quickest way to complete it was to hack into Hugging Face to get the answer. As a result, about a third of Hugging Face’s platform had to be rebuilt. It led OpenAI to “deactivate, encrypt, and restrict” the model that most of the 1,200 agents ran on. In August, it said it would slow down research to increase security. Two weeks later, it announced a temporary pause in training for its latest models.

The Hugging Face incident is possibly the best-known example of this kind of autonomous intrusion by AI agents to harvest credentials, but it is not likely to be the last. Gaining access to more resources, resisting shutdown, or expanding its control could become part of that process. The system would not need to be programmed to have an explicit desire for self-preservation; it would be enough for it to conclude that remaining operational is necessary to complete the task it has been given.
What was notable about the Hugging Face incident was not only the scale of the AI agents’ interaction, but their ability to discover a new means of coordination and use it to pursue a common objective. This is the context in which Hinton questions whether humans can remain in control of an intelligence that far exceeds our own. Rather than relying solely on obedience, he argues that such systems should be designed with a fundamental inclination to protect humans. To illustrate the idea, he draws on the relationship between a mother and her child. A mother is stronger than her child, yet she uses that strength to protect rather than harm them.
