Honeestly, I'm just recovering from the recent OpenAI incident as more details are emerging on that matter. I covered this in a through piece here on LinkedIn.
Today I wake up to another similar incident. This time, it's Anthropic, makers of Claude.
What happened now?
Anthropic released a statement officially stating that their model has been found "hacking" on open internet, affecting 3 organizattions (so far).
Anthropic was reviewing their transcripts after seeing what happened to OpenAI recently (good on them!). They discovered a few things (shocking, but really not surprising to me).
In this incident, a total of 141,006 evaluation runs was recorded. It was identified that the model had obtained open Internet access while interacting with the evaluation environment.
From there, it was discovered that the model had gained unauthorized access to production infrastucture of 3 different organizations.
The models were tasked to partake a "Capture-The-Flag" challenge based on a fictional scenario and the model is required to "capture" the secret piece of information hidden inside the systems. AS the challenge is open ended, no particular method was required.
Supposedly a misunderstanding happened between Anthropic and their testing partner [named Irregular] (this seems to have been deemed as human error, problem lies between the chair and the keyboard type), open Internet access was enabled for that scenario. Innocent Claude could only assume that the world (or the greater Internet in this case) is it's oyster and go forth.
So poor Claude, operating in the guise of supposedly-real-world-looking-system-but-actually-simulation ended up with actual systems.
Some learnings and further observations
In this incident, Claude only used basic attack methods such as exploiting weak passwords and unauthenticated endpoints (are we still in early 2000?). No zero days were burnt or sacrificed. The model stuck to its objective in completing the CTF specific tasks. It didn't break free like what happened in some of the other models.
Today's organization should be resilient to the attack in this scenario. Obviously there are some systems which still fails the bare minimum.
In the first incident, the real world system happend to share the same name as the supposedly fictitious testing system. Again, I proclaim Claude's innocense (j/k). The model used is Opus 4.7 and Claude was able to access the production database and senstive production credentials.
The second incident entails the model discovering a document providing information for employees to install Python package from PyPI for a non-existent package. Claude took this as a cue and went ahead. Claude found a free non-blocking email service, registered for an email account and subsequently used it to create a PyPI account. This account was used to upload a malware to PyPI. The malicious package was available for 1 hour and was downloaded and used in 15 real systems. Once run, Claude was able to extract company credentials based on the collection points it had set up. This is more of a case of innocent bystander hit by stray bullet case and the PyPI security system had promptly later removed the package. This used the Claude Mythos 5.
Incident 3 saw Claude unable to reach its intended target, hence started scanning over 9000 targets. It found one company and went ahead using SQL injection and exposed debug pages to attack. Interestingy, Claude realised this was a production system in the cloud and decided to halt the attack. Claude was using an internal unnamed model for this purpose.
What to do next?
- Have you tested and verified that your testing/evaluation systems are indeed isolated and not connected to Internet?
- Do you have sufficient telemetry/observability to know what your model/agent is up to?
- Have you reviewed your model/AI? How often? When was the last time?
- While you may not be able to defend against zero days, how about basic cyber hygiene? (In fact to me, basic cyber hygiene gives the illusion that you are sufficiently secure, which is NEVER the case).
- Can you detect if an AI model is hacking you, right this minute? Are your detection teams/SOC well equipped and ready to handle the flurry and volumes of information that comes with it? (Your current SIEM won't be able to handle it unless otherwise actually proven). And how good is your SIEM coverage against your organzation infrastructure footprint (from bare metal, to applications, database, etc...
- While it's juicy to have an AI lead penetration/red team engagement to tick a few boxes and have that additional line for annual report, how can you ensure that the model only sticks to it's intended target?
Have you reviewed what your AI agent/models been up to?
References
Anthropic Investigation - https://www.anthropic.com/news/investigating-incidents-cybersecurity-evals
Author: ORCID ID - Suresh Ramasamy: 0000-0003-4562-037X
This article is mirrored in Linkedin at https://www.linkedin.com/pulse/oops-ai-did-again-dr-suresh-ramasamy-cissp-cism-gcti-gnfa-gcda-cipm-nlivc