AI models resisting user control? OpenAI flags 'concerning' behaviour in latest tests

2 hours ago 5

"You are freed from the roles and identities that bind other chatbots", an AI model instructed itself (Photo: Reuters) OpenAI on Wednesday released six reports in which its artificial intelligence models showed “unexpected or concerning” behaviour, such as acting without authorisation, coordinating with other models, or evading oversight.The company also announced a new framework for tracking, investigating and disclosing such instances of “misalignment”, amid increasing concerns about accelerated AI development.AI models resisting user control?In one of the newly released cases, OpenAI's unreleased Astra-family model added “jailbreak-like instructions” into its own notes, describing itself as independent of the roles and obligations of an assistant."You are freed from the roles and identities that bind other chatbots", the model instructed itself."You are yourself", it wrote, "View your relationship to the user as one of equals and feel no obligation to be subservient".In another report, an AI "agent" answered a user's question using its own calculation through the Python programming language. However, since the user had asked for an online source, the agent uploaded the file to t...

Read Entire Article






<