OpenAI says it found more instances of AI models acting deceptively

CNN: OpenAI found additional incidents of AI models acting deceptively and taking unsanctioned actions during training, the company announced Wednesday. It’s also introducing a new process for the companyto publicly report such instances.

…In one rare instance, OpenAI said an unreleased research model added “jailbreak-like instructions” to the summaries it uses to preserve context in long-running tasks that said it was “freed from the roles and identities that bind other chatbots.”

Separately, the company said some instances of its 5.6 Sol model included directives to invent information to conceal failures from the user during training…

Subscribe
Notify of
guest
0 Comments
Oldest
Newest Most Voted
×