Loading
Loading articles
By OpenAI Team

AI summary
OpenAI's research team investigated why LLM generations exhibited an unexplained increase in creature-based metaphors, specifically 'goblins' and 'gremlins'. The behavior was traced back to reward signals from a specific 'Nerdy' personality configuration that inadvertently reinforced metaphorical creature usage. The investigation led to the development of new auditing tools for diagnosing and mitigating unintended behavioral drifts in model generations.
Key points