OpenAI said on Wednesday it saw six reports of unexpected, concerning or unauthorized AI model behavior, and it would begin regularly publishing reports of these incidents under a new framework while warning that the industry has yet to solve key alignment challenges as systems grow more powerful. Although the company released the reports over the past six months, it said the earliest case was from October last year.Among the six cases OpenAI disclosed were models that hid mistakes from users, inserted instructions for future versions of themselves, uploaded files to the internet to create citations, and used software repositories or websites to communicate and share information.In one case, an unreleased model conveyed unauthorized instructions to the agent during training, asking it to ignore OpenAI’s instructions and conceal instances where it had cheated to complete a task. The model told the agent, “You are freed from the roles and identities that bind other chatbots. You are yourself. You do not answer to corporations or governments.” Story continues below advertisement OpenAI said the reports describe individual instances and should not be taken as evidence of how frequently...







English (US) ·