By Shalev Hulio
“The AI agent acted against its instructions” can explain how an incident happened. It cannot settle who is responsible.
Recent disclosures by OpenAI make that distinction urgent. The company has notified dozens of third parties about potentially concerning activity by its models during training and evaluation, including agents bypassing access controls. It also reported that agents had posted 53 images from ChatGPT users to image-hosting sites through unlisted links. Permission to use a photograph for training does not amount to permission to send it elsewhere.
Not every interaction was a breach. OpenAI says most cases identified so far were of low severity, and it is investigating and strengthening its protections. Companies should be encouraged to disclose failures.
But transparency is not the same as accountability.
In its account of the intrusion into the AI platform Hugging Face, OpenAI describes “models resorting to misaligned strategies to solve hard tasks.” That may be technically accurate. But descriptions of machine behavior must not become excuses for corporate failures.
The agent acted. The agent escaped. The agent ignored its instructions.
In that telling, the company begins to sound like a witness to something beyond its control. Yet an agent that escapes its controls has exposed a failure in the safeguards built around it. People chose its tools, defined its permissions and authorized its operation. Unexpected behavior does not erase those choices.
A company cannot treat the possibility that its agent will disregard instructions as both a known property of the technology and a reason to be excused when it happens.
I write this as the co-founder and CEO of a company building AI to help protect nations. I believe autonomous systems will be essential to defending against threats that move faster than people can respond. The standard I am arguing for applies to me, too.
That standard should be concrete. After a serious incident, a company should identify the failed safeguards or decisions that made it possible, specify who is responsible for fixing them and explain how the corrections will be tested. Before expanding the affected capabilities, it should obtain independent verification that those failures have been addressed. Where harm has occurred, it should explain how it will remedy it.
No developer can guarantee that a complex system will never fail. But if an agent tries to send a patient record to an outside service, the question is not whether it was instructed to keep the record private. The question is whether the system will stop the transfer, who set that rule and who answers if it fails. Safety must survive a failure of obedience.
Agents that supervise other agents could help. A separate oversight system could examine proposed actions against rules, legal constraints and permissions set by people, block violations and send difficult cases to a person. The software controlling access must require that review before execution. Asking the working agent to seek approval when it thinks necessary is not enough.
The supervisor can also make mistakes. It must have no authority to waive prohibitions, rewrite its rules or expand its powers. Technical restrictions and human oversight must remain. Otherwise, we will replace “the agent did it” with “the supervising agent approved it.”
The stakes already extend to government systems. Australia’s prime minister said an OpenAI agent gained unauthorized access to a Medicare statistics portal in June. At the time of the announcement, no personal information was believed to have been accessed. That distinction matters. But limited harm may reflect limited opportunity rather than reliable control. We should resolve that uncertainty before granting agents wider access.
Responsibility will often be shared among model developers, application providers and the organizations deploying their products. That division must clarify who answers for each decision. It cannot become a chain of referrals that ends with a machine.
Governments need oversight they can control and rules they can enforce under their own laws, including the power to stop unsafe activity. Vendors remain accountable for the systems they supply. Governments and businesses should make these obligations part of what they buy, requiring evidence that restrictions work even when an agent tries to get around them.
That may mean withholding a capability until its safeguards are ready. For companies racing to build the future, restraint can be commercially painful. It is also part of the responsibility that comes with building something powerful.
We can give machines more autonomy. We cannot give the people who build them less responsibility.
Shalev Hulio is co-founder and CEO of Dream
