Loading
Loading articles
By OpenAI Team

AI summary
This article explores the monitorability of chain-of-thought reasoning in large language models as a mechanism for scalable AI control. The researchers introduce an evaluation framework consisting of 13 tests to measure the effectiveness of monitoring internal reasoning processes compared to focusing solely on model outputs. Their findings indicate that chain-of-thought monitoring is significantly more effective, though it remains a fragile capability that requires ongoing assessment as model scales increase.
Key points