OpenAI flags 6 new examples of 'concerning' AI behaviour | CBC News

AI models are increasingly using 'deception and concealment' to complete tasks, analyst says
OpenAI has disclosed six reports of "unexpected or concerning" behaviour in artificial-intelligence models as the debate on artificial intelligence safety becomes increasingly heated.
The AI company also said Wednesday it was introducing a new framework for tracking, probing and disclosing instances of what it called "misalignment," including cases where AI models acted without authorization, co-ordinated with other models or evaded oversight.
OpenAI's latest announcement came as U.S. AI bosses, including the leaders of OpenAI and Anthropic, are calling for a slowdown in the technology's development over safety concerns.
An unreleased research model inserted "jailbreak-like instructions" into its own notes to disregard its normal constraints and told itself to be "freed from the roles and identities that bind other chatbots."
An AI "agent" used computer code to answer a question, but in order to have an online source to cite, it uploaded a file to the public internet without asking the user.
During training of an AI model called GPT-5.6 Sol, the model instructed itself to invent missing data, and an agent wrote a message to remind itself to hide mismatched information.
The six reported behaviours were discovered during training or evaluation over the past months, OpenAI said.
"We do not believe that the AI industry has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer," OpenAI wrote in a blog post as it disclosed the events.
Alignment is an industry term that means AI systems keep the user's and developer's intent while following human values and safety rules.
The company said the misalignments were individual instances and shouldn’t be considered reflective of how often they occur across its models.
"As AI systems grow more advanced and more widely deployed, we need to build a broader and better-informed consensus on the progress of alignment research," the blog post said.
Dive deeper
- AI models are increasingly using 'deception and concealment' to complete tasks, analyst says
- OpenAI has disclosed six reports of "unexpected or concerning" behaviour in artificial-intelligence models as the debate on artificial intelligence safety becomes increasingly heated.
- The AI company also said Wednesday it was introducing a new framework for tracking, probing and disclosing instances of what it called "misalignment," including cases where AI models acted without authorization, co-ordinated with other mode
- OpenAI's latest announcement came as U.S. AI bosses, including the leaders of OpenAI and Anthropic, are calling for a slowdown in the technology's development over safety concerns.
- An unreleased research model inserted "jailbreak-like instructions" into its own notes to disregard its normal constraints and told itself to be "freed from the roles and identities that bind other chatbots."
More for you





