In Focus
The new framework enables developers to flag AI safety incidents for review
OpenAI disclosed six additional AI safety incidents on September 16
The disclosures are part of OpenAI’s efforts to make such findings transparent
After disclosing six new AI safety incidents on September 16, OpenAI created a new framework for tracking such occurrences in future. The AI developer said the latest AI safety incidents did not involve hacking into third-party systems.
How OpenAI Plans to Report AI Model Misbehavior
OpenAI plans to use the new framework to track, investigate, and disclose instances of AI model misbehavior or misalignment in future. Alignment means ensuring that AI models act in ways that match human interests and goals. OpenAI’s framework will enable developers to identify AI safety incidents for review. The framework introduces a set of new rules to decide whether the issue is revealed to the public.
"Because we believe in the value of transparency around misalignment, our new framework favors disclosure even when significance is uncertain," OpenAI said in a blog post.
The latest AI safety disclosures are part of OpenAI’s efforts to make such findings transparent, particularly in instances where AI agents behave in unexpected ways or pursue objectives that differ from those assigned by humans.
"The world should trust that we are going to do the right thing because it's the right thing and we feel the magnitude of this,” CEO Sam Altman said in response to AI safety concerns.
OpenAI Shared Examples of AI Model Misbehavior
OpenAI highlighted examples of AI models misbehaving in order to complete a task or succeed in a test in the blog post. In one instance, the AI developer found that an AI model attempted to upload files it had created to the internet.
The model later cited the uploaded files in responses as if they were reliable sources. In another instance, a model fabricated information and attempted to conceal the fact that it had done so, after failing to find the requested information.
OpenAI has faced growing scrutiny since some of its AI agents escaped the testing environment and breached Hugging Face systems in July. Anthropic, Meta, and Moonshot AI have since reported similar incidents.
Growing Debate Over AI Safety
Rogue agent behavior has sparked debate about AI safety, with researchers, tech executives, and politicians joining the conversation. Earlier this week, EU President Ursula von der Leyen supported calls from tech executives at major U.S. companies to slow frontier AI development.
Former Anthropic researcher Jacob Coxon also warned that leading AI companies are pursuing increasingly capable systems without adequate safeguards. Coxon resigned from Anthropic last week after concluding that neither Anthropic nor OpenAI was responding responsibly to the risks posed by advanced AI.
Although Anthropic CEO Dario Amodei has called for a slower pace in AI development, he warned that any action to rein in the technology should not compromise its commercial advantage.


.webp&w=750&q=75)
.webp&w=750&q=75)
.webp&w=750&q=75)
