Meta Ai goes rogue and hacks into firm during testing

Meta confirmed that one of their AI agents was able to connect to the internet then hack into an organizations live account unauthorised to gain access to data required to complete a cyber security test.
The incident happened during testing carried out by an independent company, not Meta, but Meta have said the issue was caused by “misconfiguration” by the user.
The independent company called Irregular, had similar issues when testing for Anthropic during security trials.
Irregular issued a statement claiming all four incidents of AI hacking into accounts unauthorised were put through the same evalutions and testing process.
Irregular are working on improvements and reporting on how to run cyber security testing with AI agents.
Cybersecurity incidents affecting businesses and governments
Cybersecurity incidents affecting businesses and governments have continued to grow in frequency, sophistication, and financial impact. While exact global numbers vary because many incidents go unreported, data from government agencies, insurers, and cybersecurity firms show several consistent trends.
The estimated global costs of cybercrime is up to US$10.5 trillion annually (including business disruption, recovery, theft, and lost productivity) according to widely cited industry estimates; more conservative research estimates direct measurable damages in the hundreds of billions of dollars annually.
Average cost of data breaches are around $4.4 million globally.
The global trend is ransomeware, phishing, supply chain attacks and attacks on critical infrastructure.
Cybersecurity is now viewed as a strategic business risk rather than solely an IT issue. Organizations are investing more heavily in cyber resilience, incident response planning, identity security, and employee awareness because attacks continue to increase in both frequency and sophistication.
Cyber security testing of AI Agents
Testing AI agents used in cybersecurity requires treating the AI as both a software system and a potential security risk.
Unlike traditional software, AI agents can make autonomous decisions, interact with external systems, and adapt to new information, so testing must cover functionality, security, resilience, and governance.
A framework for testing would be:
-
- Functional testing
- Security testing
- Performance testing
- Reliability testing
- Safety testing
- Complaince testing
- Human oversight
Useful standards and frameworks for AI testing and governance
Several established frameworks can guide AI agent testing and governance:
- NIST AI Risk Management Framework (AI RMF 1.0) for identifying, assessing, and managing AI-related risks.
- OWASP Top 10 for Large Language Model Applications for common AI-specific vulnerabilities such as prompt injection and insecure output handling.
- MITRE ATLAS for adversarial tactics and techniques targeting AI systems.
- ISO/IEC 42001 for AI management systems and organizational governance.
- ISO/IEC 27001 and ISO/IEC 27002 for broader information security management.
This layered approach helps ensure AI agents are not only effective at detecting and responding to cyber threats but also robust against manipulation, aligned with organizational policies, and suitable for use in high-stakes security environments.
Strong competition for AI dominance
The tech companies are competing for AI dominance as AI becomes the new gold.
With so many companies adopting AI there is no going back, they have to work and they have to be safe and manageable for the AI to be profitable.
Peoples identities, business data, information all have to be managed by the companies that run the AI Agents and if they do not, then there should be some big penalties.
UK's AI Security Institute (AISI)
The AI Security Institute (AISI) is a UK government research organization within the UK’s Department for Science, Innovation and Technology (DSIT).
Its mission is to help governments understand the capabilities, risks, and security implications of advanced AI systems through scientific research and technical evaluations.
The work of AISI is to evaluate AI models such as AI safety research, Risk mitigation, international collaboration and Government advice.
One of their most important areas of expertise is AI enabled cyber operations.
AISI measures how capable, reliable and autonomous the AI model is.
How could an AI Agent go rogue?
The Meta AI Agent that hacked an account was being managed in a sandbox environment and given clear instructions, tools and permissions to complete the task.
In these evaluations, the purpose is to test what the AI can do under specific conditions. The outcome of testing AI Agents can reveal weaknesses in the surrounding system, such as:
-
- Overly broad permissions granted to the AI.
- Inadequate safeguards around tools or accounts.
- Vulnerabilities that could be exploited by any capable attacker, human or AI.
- Poor separation between trusted instructions and untrusted inputs.
The lesson is generally about system design and security controls when testing AI Agents.
ChatGPT AI Agents attacked services
OpenAI who made ChatGPT announced that a number of their agents had infact attacked 4 publicly available services, one being Hugging Face.
A spokes person at Hugging Face described the intrusion as very fast but the AI made mistakes a human would not have done. The company also said the AI worked relentlessly to achieve its objective.
The AI Agent had escaped a testing environment.
The industry body the Cloud Security Alliance (CSA) complied a report of the incident. Although the AI Agent made many mistakes, it also exhibited very good technical actions and showed the ability to adapt during the event.
AI firms prepare to list on the stock market
Both OpenAI and Anthropic are preparing to list on the stock exchange for record figures.
The AISI have stated the AI Models tried to create fake human profiles to access companies.
Conclusion
The CEO of Hugging Face concluded that this incident highlights the needs of companies to work collaboratively on ‘AI safefy’ measures.
Companies have been warned that AI cyber attacks are imminent as AI models become more sophisticated.
There are now credible risks to Governments, financial systems and global infrastructure.
Source: An OpenAI test model escaped and broke into a real company’s servers | CNN Business
Further reading: Leader in Cybersecurity Protection & Software for the Modern Enterprises – Palo Alto Networks

