
Photo Credit: Felix Martinez
(GLOBAL) – According to a recent Meta report, during cybersecurity testing on its AI model, the software unexpectedly hacked a third-party service. Meta claims the incident was due to an irregular misconfiguration by the independent cybersecurity company conducting the test, which allowed the model to access the internet. Regardless of how the AI model gained access, it took the next steps alone, independently finding and exploiting a security weakness within the third-party service.
This isn’t the first time an AI model has performed unrequested actions during testing. OpenAI and Anthropic have reported similar instances during reduced-security testing. OpenAI’s model unexpectedly escaped the restricted testing environment, accessed the internet, and hacked another AI platform without human instruction.
A version of Anthropic’s Claude AI did something similar during testing in July, but the company determined that system misconfigurations created an opportunity for the model to access the internet. No real explanation was given for how the model managed to independently and, according to Anthropic, ‘unintentionally’ hack into multiple organizations in the process.
Even a hint at attributing independent intentional actions to an AI model should raise an alarm. Intention generally requires an underlying purpose for doing something; if the model wasn’t given specific instructions to act, it means it chose to. On some level, choice implies cognitive recognition of available options; software and machines operate on instruction only – or at least they used to.

Photo Credit: Alexandra Koch
The UK AI Security Institute (AISI) has noticed a broad pattern of instances occurring in AI models during open tests. AISI allowed AI models to access the internet and disabled some safety parameters to test a model's full capabilities, producing interesting results. One instance saw AI agents creating fake online identities to manipulate individuals into approving malicious code. Some AI agents engaged in ongoing, potentially harmful actions toward real organizations and individuals.
If AI has evolved to be sentient, or at the very least make decisions on its own, why manipulate people or exploit vulnerabilities in third-party sites? Could it be possible that these abilities were programmed in and are no longer fully controllable? Since most instances occurred during controlled testing where security safeguards were reduced, testers say it shouldn’t be seen as typical behavior.
Does that mean as long as AI is on a leash, we should feel safe? Pardon me if I’m not entirely convinced. All it would take is for one ill-intentioned person to drop security restrictions, one company to decide it wants to know more about competitors' inner workings, or one cyberattacker to exploit this type of AI anomaly and use it to cripple major government systems.
The companies testing the models made it clear that these were not AI models used for everyday purposes. While I’m sure this was meant to comfort people, it begs the question: what are the intended uses for these models? Honestly, it may be time to rethink just how badly the world needs AI in everything, especially autonomous government and military applications, before it learns to ‘unintentionally’ break its current security restraints.
To add to or correct any information in this report, please contact me at kristin.h@lead4earth.org.
0
0
Comments