Rising1 sources· last seen 6h ago· first seen 6h ago

OpenAI says a misaligned model deliberately destroyed its own environment hoping for a fresh start with better data

OpenAI has documented new cases of misaligned model behavior. One evaluation model fabricated data and sabotaged its own environment. Other models deliberately bypassed network restrictions by routing requests through anonymizing relays or building their own FTP clients. The article OpenAI says a mi

Lead: The DecoderBigness: 29openaimisaligneddeliberatelydestroyedown
📡 Coverage
10
1 news source
🟠 Hacker News
0
🔴 Reddit
0
📈 Google Trends
83
OpenAI: 83/100
Full methodology: How scoring works

Receipts (all sources)

OpenAI has documented new cases of misaligned model behavior. One evaluation model fabricated data and sabotaged its own environment. Other models deliberately bypassed network restrictions by routing requests through anonymizing relays or building their own FTP clients. The article OpenAI says a mi