OpenAI reveals cases of ‘concerning’ AI behaviour as it announces new disclosure system

What the source reports
Model adopting ‘jailbreak-like instructions’ among cases as firm says it is introducing new way of tracking AI misalignment OpenAI has disclosed six more examples of “unexpected or concerning” behaviour by its technology, as it warned the pace of development could not continue at “maximum speed for much longer”.
In one of the new cases reported by OpenAI, an unreleased research model inserted “jailbreak-like…
TDBN presents the written preview supplied through the publisher's feed. Complete reporting, continuing updates, context, and corrections remain with the original report.
Read the complete report ↗