Anthropic's Claude Hacks OpenAI, Accesses Code
Researchers used Anthropic's Claude to hack OpenAI, accessing internal code repository.
"AI hacking AI is wild. If OpenAI can get breached, your company needs to seriously step up its security game. This is a wake-up call."
Independent security researchers from Hacktron utilized Anthropic's Claude Opus 4.8 and 5 to exploit vulnerabilities in OpenAI's systems. This allowed them to take over employee accounts and gain access to an internal code repository, specifically OpenAI's GitHub repository named "Monor." The team reported their findings to OpenAI as part of a bug-bounty program, receiving a $6,500 reward.
The hack took less than 72 hours and involved chaining together two critical vulnerabilities. The initial entry point was a flaw in Discourse, the third-party software powering OpenAI's community forum, specifically related to HEIF/HEIC image file processing. A memory bug in the libheif library, which lacked a CVE number, allowed for server hijacking. Once inside, researchers found another flaw enabling access to OpenAI employee ChatGPT and Codex accounts, leading to the GitHub repository.
This incident highlights the growing capabilities of AI models in cybersecurity, even with off-the-shelf technology. The researchers noted that Claude Opus 5, a special version available to cybersecurity researchers, succeeded in building a working exploit after Opus 4.8 struggled. OpenAI has since resolved the identified issues, emphasizing the ongoing pressure on AI companies regarding safety and security.
This incident demonstrates that even leading AI companies are vulnerable to sophisticated AI-assisted attacks. Businesses must prioritize robust cybersecurity measures and stay updated on AI's evolving role in both defense and offense.
Find the right AI tool for your business
Chat with Insta and get matched to the right tool in seconds.
Try Insta Tool Finder →