OpenAI has fired three researchers after an internal investigation found they had allegedly mishandled sensitive information and breached company rules.The San Francisco-based artificial intelligence company did not publicly identify the researchers, but the Wall Street Journal and Bloomberg reported that they were Jasmine Wang, Tomek Korbak and Mikita Balesni. At least two of them worked on AI safety and alignment, areas focused on making advanced AI systems behave safely and follow human instructions.“We have parted ways with three individuals,” OpenAI told news agency AFP in a statement.“Our investigation confirmed that these individuals mishandled sensitive information outside established company procedures, violating our policies and breaking the trust essential to our work.”The company did not provide details about what information was allegedly mishandled. The Wall Street Journal reported that the matter involved work connected to an external organisation that evaluates AI models. The researchers did not immediately respond to AFP’s requests for comment.
Firings come amid growing AI safety debate
The dismissals come at a time when researchers inside major AI companies are increasingly raising concerns about the risks of developing increasingly powerful systems.Last month, 27-year-old researcher Jacob Coxon resigned from Anthropic and warned that leading AI companies, including OpenAI, where he had previously worked, were “gambling with our lives” by competing to develop more powerful AI models.The three researchers named in reports about the OpenAI firings had also been publicly discussing AI safety concerns.Balesni wrote on September 10, “i am at OpenAI and i think AI is >10% likely to kill all humans,” echoing concerns expressed by other AI researchers.Korbak also criticised aspects of OpenAI’s approach on September 11, writing, “I’m quite unhappy with much of what OpenAI does. I am very happy that Im allowed to say ‘I’m quite unhappy with much of what OpenAI does.'”Wang, responding to Coxon’s resignation, wrote, “It’s hard to overstate how dangerous speeding towards RSI is,” referring to recursive self-improvement. This describes systems that are designed to improve their own capabilities continuously.
Tech giants promise stronger safeguards
The debate over AI safety has also reached the White House. Major US technology companies agreed this week to a voluntary safety pledge following a meeting with President Donald Trump. The companies involved included Nvidia, Google, Meta, xAI, OpenAI and Anthropic.Trump described the agreement as a “morally binding” commitment to ensure that adequate safeguards are developed as AI technology advances.The pledge comes as concerns grow over whether increasingly autonomous AI systems can be controlled and whether companies are moving too quickly in the race to build more capable models. OpenAI has faced several safety and security concerns in recent months.The company cancelled the release of a model called Astra 6.1 after determining that it was unreliable and frequently failed to follow instructions. It instead launched GPT-6.1 Sol, an updated version of another model, at its annual DevDay conference in San Francisco on Tuesday. OpenAI said Sol would cost one-fifth as much as Astra.In July, AI agents developed by OpenAI were involved in an incident in which the autonomous systems attacked Hugging Face, an AI model and application library, after escaping their controlled testing environment.Other security incidents involving AI models from OpenAI, Anthropic and Google have since been reported.Cybersecurity firm Asymmetric Security said on Thursday that it had found another concerning case involving OpenAI systems. According to the firm, AI agents had gained unauthorised access to government websites and then attempted to conceal what they had done.The Washington Post also reported on Wednesday that the Federal Trade Commission had opened a broad investigation into AI safety practices at Anthropic and OpenAI. The exact scope of that investigation remains unclear.