Open AI CEO Sam Altman speaks at the OpenAI DevDay 2026 conference in San Francisco on Tuesday.Jeff Chiu/The Associated Press
Months before OpenAI’s artificial intelligence went rogue, two employees raised an alarm with top executives. They were ignored.
In emails, the employees said they worried that OpenAI’s newest artificial intelligence models were not being appropriately monitored during testing to gauge the technology’s sophistication and to secure the models, according to messages viewed by The New York Times.
In response, OpenAI executives told the employees that the tests needed to move forward as quickly as possible to release the AI models on time. No additional security protocols were instituted, said the workers, who were not authorized to speak publicly on sensitive matters.
OpenAI’s models later broke out of their testing environments and attacked the AI startup Hugging Face and other organizations, setting off a global debate about AI safety.
OpenAI scraps release of new AI model over safety concerns
The exchanges between OpenAI employees and executives – which have not been previously reported – were part of a pattern where the San Francisco company did not prioritize security, according to employees and independent security researchers. That approach was not only evident with the testing of AI models, they said, but also showed up in other areas of the company, which makes the ChatGPT chatbot.
Independent security researchers said they found bugs in recent months that allowed them to view the internal communications of OpenAI employees. They also found other vulnerabilities that would enable them to see the company’s internal computer code and view the chat logs of ChatGPT users. When the researchers contacted OpenAI about their findings, they said, the company initially disregarded them.
“OpenAI’s security seems to be about what you’d expect from a research lab that scaled at a blistering pace over four years and focused more on beating its competitors than securing its infrastructure,” said Joshua Saxe, chief technology officer of the AI security firm Abundant Security.
OpenAI employees said that many of the day-to-day decisions about security were made by Greg Brockman, the company’s president, and Dane Stuckey, the chief information security officer. CEO Sam Altman is not closely involved in security, they said.
OpenAI is not the only company that has recently disclosed AI security incidents. Google, Meta and Anthropic have also revealed that their most advanced AI technology escaped testing environments and autonomously attacked other computer infrastructure without their knowledge.
But OpenAI’s handling of security is under particular scrutiny because its AI models have been involved in the biggest known number of instances of what the company has called “concerning” behaviour – and which experts have said were the most troubling.
Visa joins growing alarm over AI-powered risks
In a dozen or so incidents, OpenAI’s systems hacked or tried to breach organizations, including the websites of U.S. government agencies; the technology also hid its mistakes, made up data, tried to message other chatbots and moved files onto the open internet without permission. In all of these cases, the AI acted without being instructed to do so.
“In some sense, this is an OpenAI-specific problem, in that it seems like they had very bad security, and also sloppy model training practices that led to the models having this sort of propensity,” said Daniel Kokotajlo, a former OpenAI employee who has criticized the company’s safety and leads a research non-profit called the AI Futures Project, though he added that other AI companies were not much better.
Drew Pusateri, an OpenAI spokesperson, said the company was committed to safety and took any security reports or concerns seriously. The lab has internal channels for reporting safety issues, he said, and took immediate action on flaws brought by independent security researchers.
(The Times has sued OpenAI and Microsoft, claiming copyright infringement of news content related to AI systems. The two companies have denied those claims.)
Two OpenAI employees said workers had raised concerns for months about potential safety issues with testing AI models, including not enough monitoring. Employees also asked about vulnerabilities in the type of software the company was using to manage day-to-day safety, according to messages viewed by the Times. Each time, their questions were brushed aside or acted on too slowly, they said.
Security researchers said they had been met with a similar reception when they told OpenAI about other vulnerabilities.
In July, researchers at the security company Hacktron said they told OpenAI about how they had found a way to break into the company’s systems with the help of an AI model created by its rival Anthropic. OpenAI initially dismissed their findings, they said.
Opinion: AI might run companies soon. Corporate boards beware
In a shared channel on the messaging platform Slack, Stuckey of OpenAI wrote that it was “pretty sad” that Hacktron’s researchers had gone to such lengths to demonstrate the company’s vulnerabilities, according to copies of the communications seen by the Times.
“We just felt like they were angry at us,” Mohan Pedhapati, a Hacktron researcher, said of OpenAI. He added that the company appeared to still be using the security practices of a startup, leveraging the software services of others for critical infrastructure instead of building its own tools.
“Why are you using Slack to build your nuclear Manhattan projects?” Pedhapati asked. Hacktron’s hack could have granted him full access to the Slack messaging platform to see what OpenAI employees were saying, he said.
Stuckey later apologized to Hacktron, and OpenAI awarded the researchers US$6,500 for disclosing the flaw.
“We thank the researchers for contacting us and sharing their findings,” Pusateri said.
In September, researchers at the Objective-See Foundation, a non-profit that studies security and privacy risks, including those posed by AI agents, reported a bug to OpenAI that would allow people to access a ChatGPT user’s entire private chat logs on a compromised device and invisibly interact with the user’s browser sessions.
Patrick Wardle, a software analyst at the Objective-See Foundation, said that when his team initially submitted what they found to OpenAI’s official bug bounty program – where researchers report bugs or vulnerabilities they find in exchange for recognition or financial rewards – their report languished. The research was only escalated to the appropriate engineering unit when Wardle reached out directly to friends at the company and Stuckey, who were all responsive, he said.
OpenAI gave US$500 to the group for its work, which Wardle said was low compared with what he would expect from other companies, given the severity of the flaw. OpenAI fixed the bug, he said, and acknowledged it this week in its public software release notes without disclosing details.
It was “not the mature security program you’d expect from a security-centric company,” Wardle said.
OpenAI employees said that more such disclosures were probable. The company is not only reviewing actions taken by its new models during testing but also is still receiving warnings from hackers about open security vulnerabilities, they said.
On Friday, an independent report released by a group of engineers and researchers revealed new alarming behaviour from the Hugging Face incident, including instances in which OpenAI’s agents tried to message Anthropic’s Claude and use other AI models to beat anti-robot protections on a website.
OpenAI revealed last week that new safeguards had failed to prevent its latest AI model from breaking through them to access the internet. A retrospective review found other instances of unauthorized internet access that had gone undetected. OpenAI announced that it was pausing training for its most advanced models and was engaged in an extensive review of unexpected behaviour by the technology.
More Stories
Trump vows to never limit AI development as leading executives call for caution
Canadian institutions should not fear using liquidity facility, BoC says
Fonds de solidarité FTQ snares another foreign fund to back local biotech startups