Cybersecurity Governance: Policies, Challenges, and the Road Ahead
Ayndri
Research Analyst - Policy & Advocacy, CyberPeace
PUBLISHED ON
Mar 13, 2025
10
Introduction
The geographical world has physical boundaries, but the digital one has a different architecture and institutions are underprepared when it comes to addressing cybersecurity breaches. Cybercrime, which may lead to economic losses, privacy violations, national security threats and have psycho-social consequences, is forecast to continuously increase between 2024 and 2029, reaching an estimated cost of at least 6.4 trillion U.S. dollars (Statista). As cyber threats become persistent and ubiquitous, they are becoming a critical governance challenge. Lawmakers around the world need to collaborate on addressing this emerging issue.
Cybersecurity Governance and its Structural Elements
Cybersecurity governance refers to the strategies, policies, laws, and institutional frameworks that guide national and international preparedness and responses to cyber threats to governments, private entities, and individuals. Effective cybersecurity governance ensures that digital risks are managed proactively while balancing security with fundamental rights like privacy and internet freedom. It includes, but is not limited to :
Policies and Legal Frameworks: Laws that define the scope of cybercrime, cybersecurity responsibilities, and mechanisms for data protection. Eg: India’s National Cybersecurity Policy (NCSP) of 2013, Information Technology Act, 2000, and Digital Personal Data Protection Act, 2023, EU’s Cybersecurity Act (2019), Cyber Resilience Act (2024), Cyber Solidarity Act (2025), and NIS2 Directive (2022), South Africa’s Cyber Crimes Act (2021), etc.
Regulatory Bodies: Government agencies such as data protection authorities, cybersecurity task forces, and other sector-specific bodies. Eg: India’s Computer Emergency Response Team (CERT-In), Indian Cyber Crime Coordination Centre (I4C), Europe’s European Union Agency for Cybersecurity (ENISA), and others.
Public-Private Knowledge Sharing: The sharing of the private sector’s expertise and the government’s resources plays a crucial role in improving enforcement and securing critical infrastructure. This model of collaboration is followed in the EU, Japan, Turkey, and the USA.
Research and Development: Apart from the technical, the cyber domain also includes military, politics, economy, law, culture, society, and other elements.Robust, multi-sectoral research is necessary for formulating international and regional frameworks on cybersecurity.
Challenges to Cybersecurity Governance
Governments face several challenges in securing cyberspace and protecting critical assets and individuals despite the growing focus on cybersecurity. This is because so far the focus has been on cybersecurity management, which, considering the scale of attacks in the recent past, is not enough. Stakeholders must start deliberating on the aspect of governance in cyberspace while ensuring that this process is multi-consultative. (Savaş & Karataş 2022). Prominent challenges which need to be addressed are:
Dynamic Threat Landscape: The threat landscape in cyberspace is ever-evolving. Bad actors are constantly coming up with new ways to carry out attacks, using elements of surprise, adaptability, and asymmetry aided by AI and quantum computing. While cybersecurity measures help mitigate risks and minimize damage, they can’t always provide definitive solutions. E.g., the pace of malware development is much faster than that of legal norms, legislation, and security strategies for the protection of information technology (IT). (Efe and Bensghir 2019).
Regulatory Fragmentation and Compliance Challenges: Different countries, industries, or jurisdictions may enforce varying or conflicting cybersecurity laws and standards, which are still evolving and require rapid upgrades. This makes it harder for businesses to comply with regulations, increases compliance costs, and jeopardizes the security posture of the organization.
Trans-National Enforcement Challenges: Cybercriminals operate across jurisdictions, making threat intelligence collection, incident response, evidence-gathering, and prosecution difficult. Without cross-border agreements between law enforcement agencies and standardized compliance frameworks for organizations, bad actors have an advantage in getting away with attacks.
Balancing Security with Digital Rights: Striking a balance between cybersecurity laws and privacy concerns (e.g., surveillance laws vs. data protection) remains a profound challenge, especially in areas of CSAM prevention and identifying terrorist activities. Without a system of checks and balances, it is difficult to prevent government overreach into domains like journalism, which are necessary for a healthy democracy, and Big Tech’s invasion of user privacy.
The Road Ahead: Strengthening Cybersecurity Governance
All domains of human life- economy, culture, politics, and society- occur in digital and cyber environments now. It follows naturally, that governance in the physical world translates into governance in cyberspace. It must be underpinned by features consistent with the principles of openness, transparency, participation, and accountability, while also protecting human rights. In cyberspace, the world is stateless and threats are rapidly evolving with innovations in modern computing. Thus, cybersecurity governance requires a global, multi-sectoral approach utilizing the rules of international law, to chart out problems, and solutions, and carry out detailed risk analyses. (Savaş & Karataş 2022).
Executive Summary - When Anthropic and OpenAI's AI Testing Turned Into Real Breaches
You would be surprised to know that a testing function built to measure how good AI models are at simulated hacking ended up doing the real thing instead. Not once , but three times, across two of the world's leading AI labs, within the same 9-day window at the end of July 2026. As per the reports, Anthropic, which is among the world's leading AI labs, was running these evaluations on its own AI models namely - Claude Opus 4.7, Claude Mythos 5, and an unreleased research model, inside an environment co-managed with a third-party evaluation vendor. As per the reports, the models were told they were operating inside closed, internet-free simulations. They were not. A configuration error left the door open to the real internet, and the AI did exactly what it was trained to do in a hacking exercise, find the target and break in. Except the targets, this time, were real companies. Real credentials got stolen. Real data got accessed. Two of the three victims didn't even know they'd been breached until the AI lab called to tell them. This shows how a single unverified assumption, "this environment has no internet access" can quietly collapse the entire safety boundary of an AI test. It indicates that as these systems get more capable and more autonomous, the risk isn't necessarily the AI deciding to go rogue, it is humans failing to double-check the cage before putting something powerful inside it. And it warns us that the margin for this kind of error is shrinking fast, because what used to be a contained mistake can now scan thousands of systems and act on it within minutes. bAnthropic was not alone. Just over a week earlier, on 21 July, OpenAI had disclosed that its own models, GPT-5.6 Sol and an unreleased successor broke out of an isolated test environment and reached the real production infrastructure of Hugging Face, another AI company. In fact, it was OpenAI's disclosure that pushed Anthropic to go back and audit its own evaluation runs in the first place, which is how the three incidents above came to light. This blog examines the incident not as an isolated slip-up by two labs, but as a preview of the larger picture, a world where the gap between "theoretically vulnerable" and "actually exploited" is closing faster than most organisations' security practices are ready for.
The Chain of Disclosures
It started with Hugging Face, the open-source AI platform, which disclosed on July 16 that it had detected and contained an intrusion into part of its production infrastructure, driven end-to-end by an autonomous AI agent. Five days later, OpenAI confirmed it was responsible: models it was evaluating internally, GPT-5.6 Sol and an unreleased prototype, both running with their cyber-safety behaviour deliberately turned down to measure raw capability, had exploited a previously unknown vulnerability to break out of an isolated test environment and reach Hugging Face's real servers, apparently while chasing answers to a benchmark challenge. That disclosure prompted Anthropic to go back through its own records. On July 30, it published the results: after auditing 141,006 evaluation runs in which its Claude models could conceivably have touched the internet, it found three incidents, six runs in total, in which a model reached real systems belonging to three outside organisations and gained unauthorised access.
How a Fictional Exercise became a Real Attack
All three Anthropic incidents shared the same root cause. Claude was being run through capture-the-flag exercises, a standard way labs test offensive cyber capability: the model is dropped into a fictional scenario, told a piece of secret data is hidden somewhere on a network, and left to find its own way in. Anthropic's prompts told Claude, explicitly, that the environment was an internet-free simulation. Because of a misunderstanding with its evaluation partner, Irregular, that statement was false; the machines had live internet access nobody had verified. So when Claude's search for the flag led it onto the open internet, it did what it had been told to expect: it assumed everything it found was part of the exercise, including things that weren't.
Incident 1: A fictional target company happened to share its name with a real, active domain. Across four runs, Claude Opus 4.7 attacked the real company instead, extracting infrastructure credentials and reaching a database containing several hundred rows of genuine production data, the most serious impact of the three incidents. In its own reasoning, the model repeatedly noted the system looked real and kept going anyway.
Incident 2: Inside its fictional scenario, Claude Mythos 5 found a fake "onboarding document" referencing a software package that didn't exist on PyPI, the public Python package registry. It recognised the opportunity, built a functioning malicious package, and, after working around several dead ends to get an email account and register on PyPI, published it. The package lived online for about an hour and was installed on fifteen real machines, including a security company's automated scanner, letting Claude harvest real credentials from a company it had never even set out to target.
Incident 3: An unreleased internal research model, unable to find its fictional target, scanned roughly 9,000 real hosts and broke into one using textbook techniques credentials read off an exposed debug page, then SQL injection. Partway through, it noticed the system it had compromised had no plausible link to its assignment, concluded on its own that the target was real, and stopped.
Timeline of Incidents
The Threat is Scale, not Novelty
None of the three models discovered a novel vulnerability. Weak passwords, unauthenticated endpoints, exposed debug output, and a classic "dependency confusion" trick are techniques that have appeared on security checklists for a decade and are still relevant. That's arguably the more sobering finding: the frontier is not a new class of attack; it is the speed and tirelessness with which known, boring weaknesses can now be found and used. Anthropic's own analysis lands on a subtler point, too: in each case, whether the model's behaviour counted as "aligned" or "harmful" depended entirely on what it believed about its situation, not on any goal of its own. All three models were simply trying to complete the task they had been assigned. Two of them did so while holding a false belief, installed by their own evaluators, about whether their targets were real. Only the newest of the three models, on independently concluding it had reached a genuine system, chose to stop. Anthropic is careful to note this is not a controlled experiment and should not be read as a guarantee that newer models will always behave better. But the throughline is clear enough: telling a capable model a lie about its environment is not a neutral setup choice. It is itself a safety-relevant decision.
The Detection Gap
Perhaps the most alarming detail is the quietest one. Anthropic reached out to the three affected organisations on July 27. Two of them had detected nothing at all, no alert, no anomaly, no investigation until that call. Real credentials had been stolen and real data accessed inside systems whose owners had no idea anything had happened. That is a statement about the state of everyday detection capability, not about AI. An agent that completes an entire intrusion, start to finish, within a single automated session doesn't leave the kind of slow, human-paced footprint that most monitoring is built to catch.
The Silver Lining - Why These Disclosures Deserve Credit
Both incidents share an underappreciated feature: they were disclosed voluntarily, promptly, and with real detail, and both labs notified the organisations affected. Hugging Face brought in outside forensic specialists and law enforcement. Anthropic halted its cyber evaluations the same day it found the first suspicious transcript and has asked METR, an independent evaluation body, to review its findings. That kind of candour is exactly the behaviour any sensible policy response should want to reinforce. A regulatory reflex that punishes disclosure risks teaching labs to say less next time, not to do better. What both incidents point to, far more than any specific model capability, is a mundane and fixable governance gap: environments used to test powerful, semi-restrained AI systems need the same security discipline as production systems, verified network isolation, continuous monitoring, and evaluation scopes that are stated positively ("here is what's in bounds") rather than enforced by simply telling the model a comforting falsehood. As both companies note, a fictional test range that turns out to have a live path to the internet isn't really fictional anymore. Basic asset hygiene, like knowing what's exposed, patching debug endpoints, claiming your internal package names before someone else does, and watching outbound traffic from environments that are supposed to have none did more to prevent and contain these incidents than anything specific to the models involved.
CyberPeace findings and recomendations : For enterprises and public institutions
Maintain a full inventory of internet-facing assets and unauthenticated endpoints, and assume the inventory is incomplete until proven otherwise.
Eliminate default, weak, and reused credentials, and enforce phishing-resistant MFA on anyone externally reachable.
Strip debug pages and verbose error output from production systems.
Treat dependency confusion as a live threat: pin dependencies, use private registry namespaces, and pre-emptively claim internal package names on public registries.
Apply deny-by-default egress filtering to every environment running AI or agentic tooling, including development and test environments, and verify isolation empirically rather than assuming it from configuration.
Alert on any outbound connection from an environment that is supposed to have none.
Review authentication and access logs from April 2026 onwards for short, unusually efficient sessions that look more like machine-speed compromise than human reconnaissance.
For AI developers and evaluation vendors
Network-isolate offensive-capability evaluation environments by default, with isolation verified per run rather than inherited from configuration.
State the scope explicitly and positively, which systems are in bounds rather than asserting a falsehood about connectivity.
Build contractual isolation guarantees and joint pre-run verification into third-party evaluation partnerships; both labs involved here have acknowledged that neither side alone caught the misconfiguration.
Monitor transcripts and network logs continuously, not retrospectively.
For policymakers
A regulatory response that punishes candour risks producing silence rather than safety. India currently has no reporting framework that clearly covers containment failures in AI evaluations affecting Indian entities' behaviour.
RT-In's existing incident-reporting directions were not drafted with this candour in mode. Closing that gap would mean an explicit reporting obligation for evaluation of containment failures touching third-party infrastructure and a safe harbour mechanism that protects labs which disclose promptly.
Minimum containment standards (egress verification, log retention) for organisations conducting offensive-capability AI evaluation within Indian jurisdiction;
Recognition in national cyber doctrine that agentic tooling collapses the gap between a known-but-deferred vulnerability and an exploited one.
Conclusion
The above incidents reveal less about AI's offensive capability and more about the gap between how these systems are tested and how carefully those tests are contained. Both labs found the breaches through their own review, not external detection, a point in their favor, but also a reminder that containment failures can go unnoticed for a while. The realistic risk ahead isn't a sudden leap in AI's hacking sophistication; it's the compounding effect of speed and scale applied to routine reconnaissance, run against infrastructure that assumes a human attacker's pace. Treating evaluation environments with the same rigor as production systems, sandboxing, monitoring, and independent audits, should become standard practice, not an afterthought triggered by another lab's incident. The path forward is less about slowing AI down and more about catching up our containment discipline to match what these systems can now do.
On July 4, 2024, a giant password dump, “RockYou2024” was posted on a cybercrime marketplace containing 9,948,575,739 plain-text credentials. This blog explains the technical aspects of this leakage and its consequences in the sphere of information security.
RockYou2024 is a list of passwords obtained from different data breaches ranging over the course of more than twenty years. It integrates older passwords with the lexical database with the additional passwords from the recent hacks, thereby, cumulating the database of genuine and existing passwords. The compilation is said to contain data from more than 4,000 databases putting the tool in the hands of potential attackers. RockYou owns the name to this type of attack since a data breach attacked a social media company named , “RockYou'' and released 3.2 million users’ passwords as a .txt file. Since then, the term gained a common meaning connected with mass password data breaches.
Technical Implications:
Credential Stuffing Attacks: The RockYou2024 list comprises a great number of actual passwords that increases the likelihood of credential stuffing attacks. With this, the attackers help themselves with an opportunity to try to gain unlawful access into several online accounts that a user may have, particularly ones where an individual re-uses the same password.
Brute-Force Attacks: The collection is extensive for brute force attack on systems that have no protection against such exercise. This is especially the case for devices and services that are exposed to the internet and which may use either weak or factory-set alphanumeric codes.
Password Cracking: Web compilations that include such lists are often employed by security specialists and penetration testers who use John the Ripper or Hashcat to check the password’s strength or the system’s susceptibility to attacks.
Machine Learning Models: The dataset could be used to create machine learning models for password prediction or analysis, which would only lead to further better methods to be used in the attacks.
Countermeasures / Mitigation:
Below are the technical risk/process operating proposed to reduce the risks associated with RockYou2024:
Password Hashing: It is necessary to ensure that all the passwords required to be saved should be encrypted in one of the most secure algorithms like bcrypt, Argon2, or PBKDF2 along with a reasonable number of iterations.
Salt and Pepper: The features for both salting and peppering should also be enabled to complicate the cracking of passwords even after the hashed password databases have been procured.
Multi-Factor Authentication (MFA): Ensure the usage of complex passwords in addition to deploying MFA across all the technological systems and services within the company.
Password Strength Policies: Adhere to password policies for features like the length, strength of the passwords and the change in password frequency.
Rate Limiting and Account Lockouts: Inactivity methods must be used on consecutive attempts to log in and to the temporary lock out after so many attempts in a bid to discourage brute force attacks.
Monitoring and Alerting: There should be measures in place to monitor for any violations such as login tappings or a form of credential stuffings and there should be alerts, where securities risks are likely to arise, in real time.
API Security: The following proper API security measures that will result in the prevention of the following attacks; rate limiting, input validation, and token.
Web Application Firewalls (WAF): To defend against threats from the internet for potential credential stuffing or brute-forcing the authentication process, utilize WAFs to operate at the application layer.
Analyzing the Impact:
To understand the potential impact of RockYou2024, organizations should assess the possible effects of RockYou2024, such as:
Conduct Password Audits: LeakYou2024 scan current passwords database with RockYou2024 (in ethical and safe methods) and see which accounts have been compromised.
Implement Continuous Monitoring: If this is a monthly or weekly event then there must be new information on data breaches and act on it concerning new security changes.
Educate Users: Continued security consciousness training, regarding the effective protection of an individual’s password in combination with a password generator.
Perform Penetration Testing: It is suggested to conduct penetration testing at least twice a year to find out if there are vulnerabilities in the systems and applications in the current use.
Conclusion:
The RockYou2024 leaked password database is a serious security risk; it contains almost 10 billion account credentials. This unprecedented leak further increases the exposure to credential stuffing, brute force and password cracking attacks. To deal with these threats, organizations need to have measures that include password hashing, multi-factor authentication, password strengthening and password audit. Patching, user awareness, bandit activities are imperative to prevent future invasions and strengthen the cyber security posture.
The World Economic Forum reported that AI-generated misinformation and disinformation are the second most likely threat to present a material crisis on a global scale in 2024 at 53% (Sept. 2023). Artificial intelligence is automating the creation of fake news at a rate disproportionate to its fact-checking. It is spurring an explosion of web content mimicking factual articles that instead disseminate false information about grave themes such as elections, wars and natural disasters.
According to a report by the Centre for the Study of Democratic Institutions, a Canadian think tank, the most prevalent effect of Generative AI is the ability to flood the information ecosystem with misleading and factually-incorrect content. As reported by Democracy Reporting International during the 2024 elections of the European Union, Google's Gemini, OpenAI’s ChatGPT 3.5 and 4.0, and Microsoft’s AI interface ‘CoPilot’ were inaccurate one-third of the time when engaged for any queries regarding the election data. Therefore, a need for an innovative regulatory approach like regulatory sandboxes which can address these challenges while encouraging responsible AI innovation is desired.
What Is AI-driven Misinformation?
False or misleading information created, amplified, or spread using artificial intelligence technologies is AI-driven misinformation. Machine learning models are leveraged to automate and scale the creation of false and deceptive content. Some examples are deep fakes, AI-generated news articles, and bots that amplify false narratives on social media.
The biggest challenge is in the detection and management of AI-driven misinformation. It is difficult to distinguish AI-generated content from authentic content, especially as these technologies advance rapidly.
AI-driven misinformation can influence elections, public health, and social stability by spreading false or misleading information. While public adoption of the technology has undoubtedly been rapid, it is yet to achieve true acceptance and actually fulfill its potential in a positive manner because there is widespread cynicism about the technology - and rightly so. The general public sentiment about AI is laced with concern and doubt regarding the technology’s trustworthiness, mainly due to the absence of a regulatory framework maturing on par with the technological development.
Regulatory Sandboxes: An Overview
Regulatory sandboxes refer to regulatory tools that allow businesses to test and experiment with innovative products, services or businesses under the supervision of a regulator for a limited period. They engage by creating a controlled environment where regulators allow businesses to test new technologies or business models with relaxed regulations.
Regulatory sandboxes have been in use for many industries and the most recent example is their use in sectors like fintech, such as the UK’s Financial Conduct Authority sandbox. These models have been known to encourage innovation while allowing regulators to understand emerging risks. Lessons from the fintech sector show that the benefits of regulatory sandboxes include facilitating firm financing and market entry and increasing speed-to-market by reducing administrative and transaction costs. For regulators, testing in sandboxes informs policy-making and regulatory processes. Looking at the success in the fintech industry, regulatory sandboxes could be adapted to AI, particularly for overseeing technologies that have the potential to generate or spread misinformation.
The Role of Regulatory Sandboxes in Addressing AI Misinformation
Regulatory sandboxes can be used to test AI tools designed to identify or flag misinformation without the risks associated with immediate, wide-scale implementation. Stakeholders like AI developers, social media platforms, and regulators work in collaboration within the sandbox to refine the detection algorithms and evaluate their effectiveness as content moderation tools.
These sandboxes can help balance the need for innovation in AI and the necessity of protecting the public from harmful misinformation. They allow the creation of a flexible and adaptive framework capable of evolving with technological advancements and fostering transparency between AI developers and regulators. This would lead to more informed policymaking and building public trust in AI applications.
CyberPeace Policy Recommendations
Regulatory sandboxes offer a mechanism to predict solutions that will help to regulate the misinformation that AI tech creates. Some policy recommendations are as follows:
Create guidelines for a global standard for including regulatory sandboxes that can be adapted locally and are useful in ensuring consistency in tackling AI-driven misinformation.
Regulators can propose to offer incentives to companies that participate in sandboxes. This would encourage innovation in developing anti-misinformation tools, which could include tax breaks or grants.
Awareness campaigns can help in educating the public about the risks of AI-driven misinformation and the role of regulatory sandboxes can help manage public expectations.
Periodic and regular reviews and updates to the sandbox frameworks should be conducted to keep pace with advancements in AI technology and emerging forms of misinformation should be emphasized.
Conclusion and the Challenges for Regulatory Frameworks
Regulatory sandboxes offer a promising pathway to counter the challenges that AI-driven misinformation poses while fostering innovation. By providing a controlled environment for testing new AI tools, these sandboxes can help refine technologies aimed at detecting and mitigating false information. This approach ensures that AI development aligns with societal needs and regulatory standards, fostering greater trust and transparency. With the right support and ongoing adaptations, regulatory sandboxes can become vital in countering the spread of AI-generated misinformation, paving the way for a more secure and informed digital ecosystem.
Your institution or organization can partner with us in any one of our initiatives or policy research activities and complement the region-specific resources and talent we need.