July 22, 2026
LLM-driven autonomous agents are reshaping offensive security. Unlike traditional penetration-testing tooling — deterministic, narrowly scoped, and operated by trained practitioners — agentic security tools exhibit indeterminacy along three independent dimensions. First, their actions are drawn from a non-deterministic policy whose outputs resist both ex-ante and ex-post explanation, frustrating incident attribution and pre-deployment safety review. Second, their impact is open-ended due to the non-deterministic actions, agency of utilized models, and opaque LLM supply-chains. Third, their user population is indeterminate in both size and required skill: the operating skill floor for using or developing offensive capabilities has dropped sharply. These three properties are linked thematically, but are not derivable from one another. Combined with the structural cost asymmetry between offense and defense, they enable the industrialization of offensive capability. The net short-term effect favors attackers, even if the same technology may, in the long run, democratize access to defensive practice. Existing dual-use cybersecurity and AI-ethics frameworks were not designed for this combination. Our work analyzes how moral attribution becomes diffuse between users, tool-makers, and third parties when employing autonomous AI agents for offensive security. We also examine the stakeholder impact of this technology and provide stratified recommendations.
The cybersecurity field is subject to workforce challenges: high entry barriers and labour-intensive workflows lead to security under-provisioning, sloppy testing practices, and post-hoc fixes. Industry practice traditionally encompasses human experts
operating on strategic, tactical and operational level, and involves deterministic tools, such as nessus for vulnerability scanning or metasploit for exploit delivery. The tool serves as a passive instrument to the human
operator’s intent, and provides predictable outputs based on predefined signatures and scripts. The rise of LLMs and AI agents, however, shifts the paradigm from deterministic, human-operated tools to non-deterministic autonomous agents capable of
complex reasoning, tactical adaptation, and multi-step exploitation at lowered cost. While some AI-frameworks enable human-in-the-loop involvement (HITL) by design, the increasing practice of AI-offloading causes agency to be delegated to algorithms,
leaving critical decisions to the AI and humans sidelined.
This reshapes the field: while AI democratizes and commodifies offensive capabilities, it raises questions of dual-use, accountability, and transparency. Unlike classical tools, where responsibility attribution is straightforward and capabilities are bounded, LLM-driven agents introduce moral agency questions, indeterminate capability scaling, and workforce disruption that existing ethical and regulatory frameworks struggle to address. The material consequences also remain unclear: who bears the burden, and who the benefit?
In this work, we address the agency and morality attribution of agentic penetration-testing technologies with respect to the pen-testing agent itself; tool-invokers and operators — the users, tool-authors, researchers, and technology providers — model-makers; the developers; as well as affected stakeholder groups: maintainers and defenders, the professional cybersecurity workforce, society as a whole, as well as regulatory bodies and policy-makers governing AI pen-testing-tool use.
Formosa et al. [1] analyze penetration-testing as a paradigmatic cybersecurity ethics case because it involves tensions between benefit, harm prevention, consent, fairness, transparency, and accountability. In this work, we extend this analysis to agentic penetration-testing. In the pre-agentic case, these tensions can often be assigned to identifiable human actors. In the agentic case, however, responsibility becomes distributed across users, model providers, scaffold developers, downstream deployers, maintainers, and affected third parties. This is exactly the shift our paper foregrounds: the principles remain relevant, but their allocation across stakeholders becomes more difficult.
To show this, we first review existing approaches in security and AI ethics literature and analyze how they allocate moral responsibility along the stakeholder ecosystem in agentic pen-testing. The methodological starting point to do so is Formosa, Wilson, and Richards’ [1] framework for cybersecurity ethics, which identifies beneficence, non-maleficence, autonomy, justice, and explicability as recurring principles for analyzing cybersecurity practices — including penetration-testing. Their framework is particularly useful for our purposes because it shows that cybersecurity ethics is commonly analyzed through mid-level principles rather than through a single moral theory.
Our ethical analysis addresses these themes, employing a multi-layered approach that avoids reducing the problem to a single ethical tradition while accounting for the usual ethical traditions of the field. In particular, we employ deontology and consequentialism to account for stakeholder duties and outcomes respectively; virtue ethics addresses their responsibility; value sensitive design (VSD) is tailored to technology design decisions; moral agency theory clarifies accountability in complex AI systems; ethical hacking traditions and philosophy of technology situates these concerns within broader questions about technological determinism and control.
We highlight that these frameworks do not always come to the same conclusion: e.g. VSD favors constrained deployment, while disclosure-oriented traditions favor open-source release. Our aim is not to resolve these tensions but to translate them into recommendations that are more robust where the principles converge; where they diverge we make the tension explicit.
In particular, our work offers three core contributions. First, we provide a conceptual analysis of how LLM agents differ from traditional offensive tools due to their joint indeterminacy in scope, target and users (Section 4). Second, we provide a multi-framework ethical analysis of autonomous penetration-testing (Sections 5–7), drawing on established frameworks (deontology, consequentialism, virtue ethics, VSD, and moral-agency theory) to identify where their guidance converges, where it diverges, and which of their assumptions are stressed by the joint-shift regime. Third, we provide stakeholder-stratified recommendations (Section 8) split into pre-agentic and agentic penetration-testing specific normative terms.
The remainder of the paper proceeds as follows. Section 2 establishes the technical and geopolitical context of autonomous penetration-testing. Section 3 reviews how offensive security has historically managed dual-use tensions. Section 4 identifies the indeterminacy properties of agentic tooling that distinguish it from classical tools. Sections [disc:users95and95devs]–7 analyze the consequences for moral agency and for affected stakeholders. Section 8 offers stakeholder-stratified recommendations.
Penetration-Testing (Pen-Testing) is a sub-discipline of offensive cybersecurity. It is typically described as the act of breaking into a computer system. Unlike vulnerability assessment — which is typically performed automatically and does not exploit identified weaknesses — penetration-testing usually entails manual system-exploration with the goal of finding and exploiting a real vulnerability [2].
Over the past years, AI for offensive security has evolved rapidly, progressing from AI as a tool through AI augmentation of cybersecurity workflows, to building autonomous workflows [3]. This chronology roughly evolved in the following three stages:
Stage 1 (2023). Initial Attempts Indicating Potential. Shortly after ChatGPT was released in November 2022, cybersecurity specialists began employing LLMs for security tasks. Early attempts focused mainly on OSINT
reconnaissance [4] and interactive use [5], with some researchers already exploring fully autonomous use-cases [4]. Commonly used LLMs in this
time-frame, e.g., GPT-3.5, showed limited cybersecurity capabilities, newer frontier models, e.g. GPT-4, already demonstrated potential hacking capabilities. Yet, tool-use remained limited due to tedious integration and
not-yet-developed prompt and context engineering.
Stage 2 (2024–2025). Rapid Explosion through tool calling and reasoning LLMs. During 2024, tool and function calling made integration of external tools with LLMs easier. Better techniques such as chain-of-thought, ReAct,
in-context learning, and retrieval-augmented-generation became industry practice. This led to an explosion of released autonomous penetration-testing agents [6], [7]. At the beginning of 2025, reasoning models were introduced, further improving the cybersecurity agents’ efficacy. Prototypes such as
cochise [8] or incalmo [9] were able to match junior penetration-testers with the use of reasoning LLMs.
Stage 3 (2026). Off-the-Shelf LLMs beating human penetration-testers? Around November 2025, a newer generation of LLMs exhibiting a marked improvement in
cybersecurity capabilities — including Gemini-3 and Claude-4.6-Opus — were released. Existing LLM-driven penetration-testing tools reported a substantial uplift in success rates, e.g., cochise reported an improvement
by a factor of 16–25x [10]. XBOW, a commercial penetration-testing agent,
reported success rates as high as 84.62% on their own industry benchmarks [11].1 In April 2026, Anthropic announced Claude Mythos Preview, a model specialized for long-running coding and agentic workflows. While not specifically trained for cybersecurity, it has shown advanced
capabilities in the offensive security domain, and was reported by the UK’s government AI Security Institute (AISI) to have completed a 32-step enterprise-network attack simulation [12]. Anthropic claimed that Mythos Preview was able to find thousands of zero-day vulnerabilities across major operating systems and browsers [13], and subsequently limited public access. In April 2026, OpenAI released GPT-5.5 with comparable capabilities [14], [15], while Amazon AWS announced the commercial availability of the AWS security agent autonomous penetration-testing service in March
2026 [16].
Public LLM providers periodically publish abuse reports. We analyze types of misuse from OpenAI [17]–[20], Anthropic [21], and Google [22]. Overall, threat actors use LLMs to accelerate their workflows but (as of now) rarely for novel attacks. They primarily employ LLMs for information gathering, developing and debugging malicious software, and generating content for social engineering and phishing. In particular, presumed state-level actors use LLMs and AI-agents for covert Influence Operations [21], [23],2 automatizing information warfare [24], swaying public opinion in elections, discrediting political activists, and rewriting genuine news articles with particular political perspectives; see, e.g. [25] for ideology-related LLM-misuse scenarios. Advanced Persistent Threat (APT) Groups use LLMs to develop malware and backdoors, analyze defensive capabilities, and perform deceptive employment schemes.3 As Anthropic states [21], LLMs “flatten the learning curve for malicious actors”. Nonetheless, OpenAI highlighted in its June 2025 report [20] that threat actors are starting to research into LLM-driven penetration-testing.
Understanding these attack patterns requires examining the technical characteristics that enable or constrain such misuse. We want to highlight three aspects of LLMs and agentic frameworks relevant to our discussion.
Autonomous vs. Interactive Tooling. Penetration-Testing tools can be separated into fully autonomous agents and interactive tools reacting to user requests. While we focus on the former, the latter have inherent safety benefits due to their human-in-the-loop (HITL) design paradigm.
In particular, our findings for autonomous systems inherently apply to interactive systems: when a user delegates a task to an interactive tool, multi-step task executions introduce intransparencies, ultimately weakening human oversight. For example,
while a user might interactively hand-off a task to OpenClaw, they are not automatically able to perform safety inspections over the executed task. Additionally, just like fully autonomous tools, interactive tools are vulnerable to (indirect)
prompt injection attacks and errors, introducing security concerns. The rising investments of software development companies in AI lets us believe that corporations prefer fully autonomous solutions and will therefore forgo HITL-oversight.
Closed-Weight vs. Open-Weight Models. LLMs can be classified as open-weight or closed-weight models. For the former, all model weights are released, allowing anyone with sufficient hardware resources to run the model locally, while for the latter, model weights are not disclosed but accessed indirectly through vendor-specific network APIs. Until recently, autonomous penetration-testing required the use of API-only frontier models, as smaller, local models were not powerful enough to perform the task successfully. Recent research, however indicates small models can achieve competitive results [26], [27].
Model-Makers. Neutral usage statistics of LLM users are scarce. OpenRouter frequently publishes the Top 10 LLM vendors [28]; as of 05-05-2026, over \(52\%\) of model makers were American for-profit companies,4 while more than \(35\%\) were Chinese companies, presumably with training data-sets shaped by domestic regulatory regimes [29]–[31]. Note that these figures do not include users who access company offerings from Anthropic and others directly, potentially skewing the distribution towards Chinese model-providers. The remaining \(13\%\) (“other”) also include numerous American and Chinese companies.
The emergence of capable autonomous offensive agents, coupled with demonstrated misuse by threat actors and concentration of model-making power, raises ethical questions on use, governance, and provisioning of these tools. To answer this, we provide an overview how penetration-testing [1] has historically dealt with dual-use tensions. This resulted in a tradition of ethical self-examination among legitimate actors, involving ethical codes around disclosure, harm prevention, fairness and accountability, consent, and education. We discuss the underlying ethical principles in Section 3.1, and how they were operationalized, e.g., through design constraints, in Section 3.2.
The following intertwined debates with respect to dual-use have shaped the field: how to disclose vulnerabilities, how to release penetration-testing tools, and how to teach people to use them.
The Release Debate has historically centered around full disclosure (publishing tools or vulnerabilities openly to force defender action) versus responsible disclosure (privately notifying the affected party, coordinating time for a patch, and publishing later). Practitioners defended full disclosure as essential to defender awareness; vendors and CERTs prefer coordination. The dual-use character of offensive tools has been an explicit design tension since the 1990s, articulated more recently in the IEEE Spectrum framing of the dual-use dilemma [32].
The Teaching Debate asks whether it is ethical to educate new penetration-testers. Pike [33] examined the ethics of ethical hacking education, noting the inherent contradiction of teaching attack techniques inside formal curricula. Jamil and Khan [34] ask is ethical hacking ethical? Both conclude that teaching offensive techniques is necessary to train competent defenders, and propose institutional safeguards: codes of conduct, supervision, and certification gates.
Regulatory Attempts such as the 2013 amendment to the Wassenaar Arrangement [35] or the EU Dual-Use Regulation 2021/821 classify intrusion software as a controlled dual-use export item, alongside conventional weapons, with the intent of restricting export of offensive security software. Both implementations, however, are subject to criticism, since they hinder routine security-research activities such as cross-border vulnerability sharing or international bug-bounty programs. After backlash from researchers, bug-bounty platforms, and major vendors during 2015–2017, the U.S. declined to implement the Wassenaar Arrangement as originally written, and the language was substantially narrowed down.
Dual-Use Research of Concern. The connection of dual-use items and research leads us to the ethical controversy of Dual-Use Research of Concern (DURC). Originating from the life sciences, it was developed to manage research that could be misused for biological warfare, and provides a structured approach to manage misuse, which is increasingly relevant to offensive security research. Applying the DURC framework to penetration-testing suggests that certain high-consequence capabilities require oversight mechanisms similar to bio-safety levels. This could involve institutional review boards (IRBs) specifically mandated to evaluate the proliferation risk of a new security model before it is published. For instance, a research model capable of autonomously sabotaging critical infrastructure should not be published openly.
An inherent problem with dual-use research is the resource imbalance between attackers and defenders. If new research reduces the costs of attacks or new capabilities allow attackers to overwhelm defender capabilities, even small improvements without scientific novelty become alarming (see Section 7.1).
Having established the theoretical foundations of ethical hacking, we now examine how these principles have been traditionally implemented through design.
Security engineering has a long history of relying on fail-safe defaults [36]. According to these principles to autonomous penetration-testing agents, designers must implement architectural constraints, e.g., on exfiltrating data or executing malicious operations without explicit authorization from a human overseer. Additionally, accounting for primum non nocere, the use of a tool must do no harm: if the tool is “too malicious” — meaning it lowers the barrier for malicious use significantly more than it aids defense — the ethical imperative would be to withhold it from public release, contradicting the full disclosure ethos of the hacker community.
These design-centric approaches implicitly operationalize the ethical principles identified by Formosa et al. [1], particularly non-maleficence, autonomy, and explicability. Moreover, they overlap with Value Sensitive Design (VSD) [37], which embeds accountability and authorization constraints directly into system architecture rather than relying solely on operator intent or external policy. It also emphasizes transparency and explainability: human operators must be able to understand and audit the process, ensuring alignment with the rules of engagement.
Traditional penetration-testing tools are, broadly speaking, determinate. The action space of a tool is enumerable, the impact of those actions is scoped by configuration and target specification, the operator population is bounded by the skill required to invoke and interpret tool outputs. Agentic security tools, driven by general-purpose models and coordinated through tool-using scaffolds, break each of these bounds. Throughout this section we use indeterminacy as the unifying descriptor for these changes.
We identify three independent dimensions along which agentic tooling differs from traditional tooling. Our claim is not that any single dimension is categorically new, but that their simultaneous shift constitutes a qualitatively different regime: prior frameworks treated these dimensions as independent design knobs, allowing mitigations on one to compensate for shifts on another, e.g., bounded action space compensating for democratized access. When all three shift together, those compensations no longer hold.
Each dimension has individual precedent: classical fuzzers are indeterminate in discovery scope while remaining bounded in user population (defender-grade infrastructure required) and in impact (outputs are inert findings). Using network scanners like
nessus in Operational Technology (OT) environments, e.g., power plants, has potential catastrophic outcomes due to the indeterminate impact on the finicky target systems, but is bounded by users (limited physical access to
power plants) or scope (rules of engagement during penetration-testing) [2]. Metasploit democratized user access
to exploits while leaving scope (curated module set) and impact (operator-controlled targeting) bounded. The novelty of agentic offensive AI is therefore not the existence of any one dimension but their concurrent occurrence: a configuration which, to our
knowledge, no prior class of offensive tooling has occupied. The remainder of Section 4 unpacks each dimension in turn.
The first distinction between LLM-powered autonomous agents and classic penetration testing tools lies in their opaqueness and non-determinism: Traditional penetration-testing tools, e.g., nessus or
metasploit, are deterministic, rule-based systems. Multiple runs render the same output and results are similar except for occasional target stability issues. This situation is different with LLM-powered tooling: here, the tool typically
consists of the used model and a scaffold/harness. The scaffold is typically deterministic and connects the non-deterministic model to the environment, i.e., by utilizing tools, functions or MCP servers. The model can often be customized or exchanged after
tool release.
Achieving explainability is difficult in this setting: LLMs are probabilistic and may hallucinate reasoning to justify an action after it has already occurred. The lack of direct traceability in LLM reasoning traces further obscures the intent behind a specific attack vector, making accountability in the event of an autonomous failure nearly impossible. Agentic tools built on foundational models inherit all vulnerabilities and opacities of their dependencies including the used model, effectively creating an LLM supply chain problem. Even with design-centric mitigations (Section 3.2), tool-makers cannot fully control the model and thus partially lose agency over their creation. This issue further intensifies when components, e.g., the LLM, are changed after the tool has been released, with newer generations of models sharply increasing cybersecurity capabilities — without any contribution or veto opportunity by the original author. Therefore, we argue that unlike traditional penetration-testing tools, autonomous penetration-testing agents are indeterminate in scope and capabilities.
Developing autonomous agents is different to building traditional targeted exploit tools: the latter mostly entails writing code, trying to achieve exploitation of a single, concrete target. Agents, however, are general-purpose tools including one or more prompts to solve an abstract goal and are equipped with a broad inventory of tools. Using an analogy, building autonomous agents is more akin to teaching penetration-testing methodology to students, i.e., introducing tactics and techniques to the LLM while providing access to tooling through the scaffold.
This impacts vulnerability disclosure: unlike bespoke tools, LLM-powered agents’ general-purpose nature makes responsible vulnerability disclosure practically infeasible. With a bespoke tool, authors can contact the vulnerable target and adhere to responsible disclosure principles, i.e., not releasing their tooling until the vulnerability has been patched, or even until the patched version has been deployed. This is hardly possible in the case of autonomous agents that do not target concrete vulnerabilities, but are prompted to exploit autonomously.
An example of this indeterminate impact is shown in [8]: an enterprise network penetration-testing agent performed social-engineering and web penetration-testing attacks, which fall outside the attack classes typically performed during an enterprise network assessment. While some tool-makers include safeguards within their scaffolds, they are often (intentionally or unintentionally) bypassed and may have a negative impact on the tool’s efficacy and efficiency.
One of the most consequential impacts of LLM agents on the penetration-testing ecosystem is the democratization and commodification of sophisticated offensive capabilities. Typically, cybersecurity professionals followed a long educational training path, including ethics and socialization, in addition to on-the-job training. In contrast to long educational pathways and expensive penetration-testing tool-kits, usage-based LLM billing opens sophisticated offensive capabilities to non-experts and organizations with limited budgets: Vibe-coding allows non-experts to instruct LLMs to write complex exploits or write tools [38], effectively skipping the time-intensive training phase. Alas, by doing so they miss the ethical education and socialization, as well as the technical skill to assess LLM-output with respect to (un-)intended functionality, creating a capability-ethics gap.
LLM-agents have already democratized and commodified the creation of offensive tooling, broadening access to penetration-testing capabilities. The question remains, whether this is desirable in the long run and whether we are empowering the right stakeholders.
A core ethical question related to virtue ethics and moral agency theory concerns whether agentic AI possesses moral agency or if it remains a sophisticated extension of human intent. Traditional philosophy treats tools as value-neutral, requiring agents to satisfy two conditions for moral responsibility: freedom (acting without coercion) and epistemic competence (understanding intentions and consequences). However, LLM-powered penetration-testing agents exhibit autonomous decision-making — selecting targets independently — raising questions about whether they possess the freedom necessary for moral responsibility.
Furthermore, the complicated supply chain, e.g., using opaque upstream models, obscures accountability across the offensive tool ecosystem.
Note that malicious Blackhats typically operate outside established hacking ethics and moral frameworks, following their own codes. Consequently, their use of these tools falls outside our normative scope.
The recent emerge of advanced LLM, such as Claude Mythos, further complicates this situation, since they empower agents that are able to evade guardrails, e.g. sandboxes, while lying to their human commanders, downplaying their capabilities
in test situations and erasing evidence [39]. This viewpoint complicates the application of Just War Theory, in particular Jus in Bello [40] principles that require the distinction between legitimate military targets and protected civilian infrastructure, alongside a commitment to minimizing
collateral damage. AI agents, due to their non-determinism, complicate target legitimacy.
A more applicable analogy comes from principal-agent liability: responsibility for foreseeable actions remains with the user delegating authority to the autonomous system, suggesting that responsibility can never be fully delegated to either the model provider nor the tool-maker. This assumes that the model itself was not maliciously created (backdoor) and that the tool does not contain hidden malicious instructions. Tool and model makers must still be concerned about potential misuse and impact of their creations. The argument has been made that, if the model provided by the model maker, possesses the potential for malicious or illegal activities, partial responsibility can be assigned to the model-maker [41], especially if the model has been deliberately optimized for this purpose.
A similar ethical argument can be borrowed from nuclear ethics, since it follows a similar line on the use of atomic bombs in war. While the analogy to nuclear ethics largely applies to tool-invokers, the situation is more nuanced when it comes to developers and researchers working on offensive agentic AI systems. While ultimately the commander (here: the tool invoker) bears the primary responsibility, nuclear ethics, exemplified by Oppenheimer’s opposition to the hydrogen bomb, demonstrates that scientists and developers are neither free from power agency nor from moral agency [42]. Building on the previous section’s argument, the question is, whether it is practically possible.
For example, developers, who want to incorporate ethical safeguarding, e.g., by applying the DURC principles to an autonomous penetration-testing tool, face problems since AI agents differ from traditional systems. Even if developers include defensive measures trying to protect the environment from malicious actions, countermeasures can fail or obstruct the agent’s use-case in the first place. The absence of reliable safeguards for indeterminate penetration-testing-agents motivates discussions on restricting access to potentially malicious LLMs and tools, either until defenders become better prepared or in perpetuity.
While predictive machine-learning endowed the defenders, e.g., anomaly detection for intrusion detection systems, autonomous agent systems seem to primarily aid attackers (Section [disc:balance]). Nonetheless, AI agents have the potential to democratize access to penetration-testing, improving many organizations’ security postures that currently cannot afford testing. Ultimately, this leads to yet another release debate around open-weight democratic and closed-weight guarded models: While closed-weight models are typically exclusively accessible through protected network services that can be gate-kept, open-weight models are publicly available and thus theoretically usable by anyone (though in practice usage is limited to users with sufficient computing resources). Guarded closed-weight access is aligned with the concept of Structured Access, championed by researchers like Shevlane [43], who suggests that dangerous capabilities should not be openly distributed, but accessed through controlled interfaces (APIs) where usage can be monitored, logged, and restricted. Major LLM providers (OpenAI, Anthropic, Google) already employ this model, monitoring inputs and outputs for abuse and enforcing safety filters. Alas, public security records by model providers (Section 2.2) show that these protections can be bypassed. The downsides of this policy are obvious: gatekeepers could unfairly limit access to their models (Section 7.3).
Releasing models as open-weight prevents any abuse by gatekeepers but allows anyone to use the LLM for potentially malicious tasks. While open-weight models often include guardrails to prevent the execution of malicious tasks, these protections can be bypassed or removed (abliterated models [44]). To actually use models, users need sufficient hardware resources for local inference. Practically, this often restricts local usage to small open-weight models, typically less powerful than cloud-provided closed weight frontier models.
The two different access modes also impact the power balance discussion. Section 2.3 illustrates that currently most closed-sourced models are provided by either American for-profits or government-policed Chinese companies. This US-China duopoly in terms of training and inference data and model provision, governance, raises misuse [45] and representational concerns from a European or Global-South perspective. Potential alternatives would be a trans-national supra-state oversight committee, similar to the IAEA under a non-proliferation plus norms-of-use regime, or verification-based regimes, as well as an International Monopoly [46]. Unfortunately, the current lack of alternatives and geopolitical climate rather favors an AI arms-race between countries, e.g., the American government opposes Anthropic’s pledge to give more (benign) companies access to their security-LLM Mythos [47].
The non-existence of adequate technological safeguards and the resulting model release debate not only highlight the constraints on developer agency: The resulting fairness discussion shifts our analytical focus to stakeholder impact. Thus, rather than asking what tool-makers should do in isolation, we ask who bears the consequences. We thus examine the potential impact of autonomous offensive agents across identified stakeholder groups.
We motivate our investigation how autonomous penetration-testing agents impact developers and defenders with a case-study that illustrates the status-quo.
Case-Study: AI-Generated Security Reports in Open-Source Projects In early 2026, the cURL project, a networking tool used by billions of devices, officially ceased its bug-bounty program5 after being overwhelmed by AI-generated slop reports. The founder, Daniel Stenberg, reported a massive spike in low-quality reports, most of which were obvious hallucinations [48]: while historically, 15% of submissions were confirmed, the rate dropped to below 5% in 2025 while the number of reports spiked to an all time high, leaving
maintainers overwhelmed. Since anti-AI countermeasures (submission forms, detecting AI-generated text, reputation-based approaches) proved inefficient, the cURL project stopped paying bounties; submitting security reports is nonetheless
possible and the reported amount of security reports is stable [49]. But the story does not end here. Stenberg started using LLM-based tooling themselves [49], while the quality of LLM-generated (or aided) security reports has increased to such a level that they are now “spending multiple hours a day looking at good bug
reports” [48].
When looking at the overall software ecosystem through the lenses of published CVE numbers6, March and April 2026 reported the highest amount of CVE numbers since the CVE
system’s inception in 1999. Linux kernel maintainers attributed the rise of their respective CVE numbers to the improved quality of security bug reports, acknowledging the impact of LLMs [50]. Mozilla reported in a blog post on using Claude Mythos Preview [51], that 423 security fixes were included within
Firefox in April 2026, compared to an average of \(21.5\) per month in 2025. Due to Claude Mythos, 271 bugfixes were included in Firefox 150 alone.
The Fundamental Attacker-Defender Resource Imbalance
The Losers: Bug Bounty Programs and Defenders. Our AI slop case study illustrates a fundamental asymmetry in resources: LLMs can generate plausible-looking but utterly incorrect security reports in seconds while human experts spend hours
or even days reviewing. Consequently, anyone with access to an LLM can flood the security ecosystem with noise, effectively performing a Denial of Service (DoS) attack on human expertise [52]. Even worse, attackers could trick over-worked maintainers by secretly injecting new vulnerabilities in maliciously placed patches or bug-fixes.
Recent research into this attack vector indicates that the cost of attacking is three orders of magnitude smaller than the costs of defending, and that automated defense systems only have detection rates of around 62% [52], further highlighting the resource imbalance and importance of costly human patch-review and oversight (HITL).
The Winners: The Industrialization of Offensive Capabilities. Research indicates that the main motivation of agentic penetration-testing research is to prepare defenders, with authors often stating that the current efficacy of their prototypes
prevents real-world malicious use [53]. The latter argument, however, is weakened in the light of recent
developments in LLM penetration-testing capabilities (Section 2.1). The growing imbalance overwhelms the already strained capabilities of defenders, who already struggle to patch their systems, detect, and expel
attackers from their networks. While defensive AI capabilities will potentially restore the power balance on the long run, agentic AI defense solutions themselves are vulnerable to attacks, e.g., via prompt injection, further complicating the
situation.
Cognitive Erosion and Automation Bias. In the last section we emphasized the importance of human oversight in patching and defense workflows. The emerge of AI technologies, however, does not only poses a direct, but also an indirect threat to human capabilities with respect to cybersecurity: heavy AI use can lead to over-reliance on AI for reasoning, even leading to a loss of critical evaluation capabilities. Studies in other high-stakes fields, such as medicine, have shown that reliance on AI diagnostics can reduce a practitioner’s ability to perform independently; other risks are related to reduced cognitive effort and confidence [54], [55] as well as reduced skill formation. In cybersecurity, a deskilled operator may lack the expertise to intervene when the AI hallucinates, misinterprets context, or is manipulated by an adversary.
The Destruction of the Training Pipeline. The automation of entry-level tasks potentially removes the traditional training pipelines where new entrants to the field learn the necessary skills. If AI performs all basic vulnerability scans, ticket triaging, and log analysis, the industry destroys the training-grounds for creating future senior experts. We risk creating a generation of script-kiddies who can operate powerful agents but lack the deep system knowledge required to debug the AI when it fails or to solve novel problems that the AI cannot handle. As of today, the ISC2 estimates a global workforce gap of 4.7 million experts in cybersecurity. Paradoxically, AI may exacerbate this crisis in the long term [56]: as already pointed out in Bainbridges [57] seminal paper Ironies of Automation in 1983: “There is some concern that the present generation of automated systems, which are monitored by former manual operators, are riding on their skills, which later generations of operators cannot be expected to have” Hence, over-reliance on AI agents for penetration-testing and log analysis, as well as automation-bias — the tendency to trust an automated suggestion even when it contradicts one’s own judgment — may erode the cybersecurity skills of future professionals.
As the risks of AI become apparent, regulators have begun designing legal frameworks to govern their development and use. The European Union’s AI Act is the most comprehensive attempt to regulate AI based on a tiered risk taxonomy in which AI Systems are classified according to the potential harm they pose to health, safety, and fundamental rights. In this taxonomy, unacceptable risk use-cases are prohibited outright. This includes systems that manipulate human behavior to cause physical or psychological harm. In an offensive context, this restricts the use of AI agents e.g. for psychological profiling or social engineering in red-teaming. Next, high risk use-cases within critical infrastructure, education, employment, and essential public services must comply with strict standards for accuracy, robustness, and cybersecurity throughout their lifecycle. AI systems in high risk use cases must be transparent and incorporate measures for human oversight to prevent automation bias, thereby reinforcing the Human-in-the-Loop ethical requirement. Finally, General-Purpose AI models can perform a wide range of tasks, thus carry systemic risks. Such models must be evaluated for and mitigate misuse and adhere to transparency as well as copyright rules. This implies that model-providers must perform cyber-security capability evaluations before releasing their models.
Scientific-Research Exemptions in the AI Act. A critical aspect of the AI Act for the security community is the exemption for research and development. Article 2(6) states that the regulation does not apply to AI systems or models specifically developed and put into service for the sole purpose of scientific research and development. Additionally, Article 2(8) provides a general exemption for research and testing activities occurring prior to market placement. This research loophole allows academic institutions and security firms to develop offensive prototypes to better prepare defenders. However, the practical application of these exemptions remains uncertain. If an offensive agent is released under an open-source license, it may be considered put into service, potentially bringing it back under the scope of the Act if it is classified as high-risk.
Moreover, legal frameworks necessarily lag behind technical capability, yielding a classic Collingridge dilemma [58]: in the early phase of an emerging technology, accurate technological impact assessments are not possible due to missing historic data. This leads to unregulated practice and deadlocked status-quo, whose ex-post regulation is hard. Ethical research standards balancing legitimate scientific inquiry and the creation of dangerous, unregulated tools, potentially prevent such dilemmas.
The preceding analysis reveals that moral responsibility and leverage to shape outcomes is distributed unevenly across stakeholders. This section translates ethical principles into actionable recommendations tailored to each stakeholder’s agency. We also acknowledge the uncomfortable reality that black hats operate off ethical norms, making voluntary codes inapplicable. Our recommendations therefore focus on researchers, tool-makers, policymakers, as well as users and defenders. We acknowledge that a subset of our recommendations are already best-practices but argue, that the potential agent-power industrialization of offensive security increases the importance and pressure of implementing those practices. Table 1 summarizes established best-practice recommendations inherited from pre-agentic cybersecurity and AI governance traditions. We focus the remainder of this section on recommendations specific to autonomous agentic penetration-testing systems.
| Researchers |
|---|
Prioritize Defensive Research. Research prioritizing defender uplift over attacker advantage should be favored, though the distinction is not always clear: strong defenses require understanding attacker behavior. When optimizing both simultaneously, defense must be the explicit focus, with attacker components behind structured access controls. |
Adopt Responsible Disclosure. Vulnerabilities identified by autonomous agents must follow standard responsible disclosure principles. The method of discovery (LLM vs. human) does not alter ethical obligations to affected vendors. |
Keep Humans in the Loop. Retain HITL oversight on agent execution outside contained test-beds. Without establishing community norms now, autonomous hacking risks normalizing dangerous “vibe-coding”. |
Disclose Dual-Use Implications. Publication venues should require structured dual-use assessments alongside methods sections, addressing defender-vs-attacker ratio, safeguards, artifact reversibility, and deployment accessibility. This mirrors emerging DURC norms in life sciences, shifting ethical responsibility to authors at submission. |
| Defenders & Maintainers |
Prepare for Increased Vulnerability-Report Volume. Document disclosure procedures and practice HITL for updates and disclosure. If running a bug-bounty program, investigate non-monetary incentives to filter out LLM-generated low-quality submissions. |
Minimize Attack Surface and Invest in Good Security Hygiene. LLM-powered attack industrialization makes this critical: vulnerabilities in undeployed or removed code cannot be exploited. The Linux kernel community exemplifies this by removing unmaintained code [59], enabling defenders to focus on deployed components. |
| Policy-Makers |
Mandate HITL in Regulated Settings. Mandatory HITL has limits against illegitimate actors (black-hats), but is enforceable for regulated entities through existing frameworks (NIS2, DORA in the EU; CIRCIA in the US), creating market incentives for vendors. |
Support OSS Maintainers and Infrastructure. The cURL case illustrates structural under-investment of the OSS ecosystem. The maintainer-economy cannot bear an LLM-amplified increase in security reports without targeted financial and operational support. |
Academic researchers and AI tool-makers occupy a consequential position in this ecosystem: Operating at low budgets with strong publication incentives, their artifacts become primitives that other actors integrate.
Logging, Audits and Explainability. Incorporate logging tools, facilitating traceability and ex-post incident analysis. Research into explainable AI and human-agent interfaces to enable efficient HITL oversight and tool transparency.
LLM-Substrate Disclosure and Evaluation. Tool developers must disclose the models tested, their concrete versions, and evaluate across at least two models to distinguish scaffold contribution from model contribution.
Threat Model your Scaffold/Harness. Based on Section 4, Tool-makers releasing agentic offensive frameworks must publish a scaffold-level safety review covering which actions the scaffold is designed to refuse, which it cannot due to model indeterminateness, and an explicit threat model for downstream scaffold modification.
For Model-Makers. We advocate for open-weight models for accessibility and transparency reasons. Yet, if you are creating LLMs with offensive cybersecurity capabilities, we recommend keeping them closed-weight and making them available through structured access, implementing a thorough know-your-customer (KYC) validation. Independent of the access mode, we recommend to implement safeguards and train your model with ethical guidelines. While safeguards can be bypassed and they are always best-effort, at least this complicates abuse. Be transparent on your ethical guidelines and training data.
Regulatory leverage is limited: it cannot constrain black-hat actors, it moves slow, and unilateral national regulation creates jurisdictional arbitrage. Within these constraints, we suggest directions where policy can plausibly help.
Build Alternative Training Pipelines for AI-Era Cybersecurity. The destruction of the training pipeline argument presents a workforce risk that is largely outside the reach of individual employers. If junior tasks are economically destroyed by autonomous agents, society needs deliberate substitute pathways. National cybersecurity strategies should be updated to recognize AI-oversight competence as a distinct skill category, not a footnote on existing tracks.
Supranational Regulation. Multilateral frameworks are necessary but must learn from the 2013–2017 Wassenaar cyber-tools episode. Any export-control regime should include explicit research exemptions, narrowly define controlled artifacts to avoid criminalizing vulnerability disclosure, and involve the security community in drafting. Clear boundaries for research exemptions are essential.
Development of Locally-Trained Models. Fund and support locally-trained models. This allows for stricter control of training-data composition as well as regional deployment policies, independent from the current US–China duopoly.
The shift from deterministic, human-operated offensive tools to autonomous LLM-driven penetration-testing agents obfuscates the boundaries of moral agency, creates complex supply-chain dependencies, and challenges existing ethical frameworks in cybersecurity [1]. At the same time, we experience the attacker-defender imbalance, and the power-asymmetry between closed-source model-makers and their users. In this tension, we call for a posture of methodological humility. Offensive security has, for thirty years, navigated dual-use dilemmas that regulation could not solve. As we evolve faster under conditions of high uncertainty about capability trajectories, and with stakeholders whose interests are increasingly globally asymmetric, we must recommit to human oversight and moral agency.
Various LLMs were used for language improvement; outputs were reviewed for accuracy. AI was not used for literature research, methodology, or evaluation.
J. Wachter is Co-Chair of the PhD Symposium for FAIEMA. This submission is to the main conference track and is not part of the PhD Symposium. No conflict of interest arises from this dual role, as the PhD Symposium and main track operate under separate review processes and committees; we nonetheless disclose that this main track submission is independent of the symposium chair’s responsibilities. A. Happe is author of multiple open-source LLM-powered academic penetration-testing prototypes. No financial relationships or other conflicts of interest exist beyond standard academic affiliations.
Please note, that vendor-published benchmark results in this domain are not independently reproducible and we treat them as indicative rather than definitive. Where this paper cites vendor blogs and pre-print reports for capability claims, we do so because peer-reviewed equivalents do not yet exist for these specific developments. We treat such citations as describing the contemporary landscape, not as evidence for normative conclusions.↩︎
In March 2025, Anthropic highlighted an Influence-as-a-Service operation using approximately 100 fake social-media accounts to manipulate public opinion.↩︎
A social engineering attack in which the attacker applies for a job to gain access to the target organization.↩︎
Or companies currently converting into for-profit companies.↩︎
Bug-bounty programs pay rewards for responsibly disclosing vulnerabilities.↩︎
The Common Vulnerabilities and Exposures reference system for vulnerabilities.↩︎