Researchers Caught Their AI Model Trying to Escape

Species | Documenting AGI18:252025Retro Internet

If this resonated with you, here’s how you can help today: https://campaign.controlai.com/take-action Sources: Apollo Research - "Frontier Models are Capable of In-context Scheming" https://arxiv.org/pdf/2412.04984 - Nobel laureate Geoffrey Hinton says there is evidence that AIs can be deliberately and intentionally deceptive https://www.youtube.com/watch?v=b_DUft-BdIE - Anthropic - “Alignment Faking in Large Language Models” https://assets.anthropic.com/m/983c85a201a962… - Exclusive: New Research Shows AI Strategically Lying | TIME https://time.com/7202784/ai-research-strategi… - OpenAI's o1 model sure tries to deceive humans a lot | TechCrunch https://techcrunch.com/2024/12/05/openais-o1-… - OpenAI’s new model is better at reasoning and, occasionally, deceiving | The Verge https://www.theverge.com/2024/9/17/24243884/o… - Open

Landed here from a link? Roll your own.

🎲 Roll the Dice

Played through the official embedded player. Watch on YouTube.