It is not enough to give your moderation rules to ChatGPT


Abstract

Content moderation practices and governance paradigms are changing rapidly, as fewer human moderators are deployed as ‘experts’ by social media companies in a centralized manner. Instead, the companies are focusing more on community approaches, relying on volunteers to provide accurate information and make correct decisions. In decentralized moderation, communities have always relied on volunteers, updated community guidelines, and internal discussions thereof. For both content moderation paradigms, Artificial Intelligence (AI) seems like it could help ease moderation burdens of time, mental health, and accuracy. One possible way to operationalize AI in content moderation is a “policy-as-prompt” approach, where the policy is formulated as a natural-language prompt and then passed to a large language model (LLM). This model then aids in moderation tasks. In this paper, we briefly lay out the technical and governance properties of this approach, and argue that its limitations lead to specific risks and harms that have to be addressed. Towards alleviating them, we lay out multiple considerations towards more effective prompt governance, but ultimately find that writing prompts alone is not appropriate for ensuring meaningful community governance.

1 The (New) Content Moderation Landscape↩︎

Artificial intelligence has had a large impact on the content moderation landscape. Machine learning classifiers are now the first line of moderation on almost all major social media platforms, such as TikTok, Facebook, Instagram, or X/Twitter, while human moderation has been scaled back or removed from many of them. In modern moderation, humans mostly act as reviewers of edge cases, appeals reviewers, and (increasingly) training data labellers. [1][4] This is in part due to a shift of trust and safety policies being politically contested [5], as X/Twitter and Facebook/Instagram (Meta), for example, have reduced moderation enforcement in the name of free speech [6], [7].

Large Language Models and Large Multimodal Models (LLMs & MLLMs) are increasingly proposed and trialed for tasks related to content moderation, promising to alleviate pressures on time and mental health while outperforming human moderators [8], [9]. For example, Meta’s Llama Guard [10] takes a safety taxonomy as an input and then labels posts as ‘safe’ or ‘unsafe’. More recently, OpenAI published an open-weight model gpt-oss-safeguard [11], where developers supply a written policy and the model reasons over it to produce a labelled and explained decision on a post.

Due to scaling (and monetary hurdles), biases, and performance drops in non-English text, these models are not yet widely independently deployed [9], [12][14]. As models get better, these problems are expected to become negligible. LLMs would then swiftly replace multiple tasks in content moderation workflows. Still, one underlying problem would persist: simply giving an LLM a policy or a set of rules as a prompt alone cannot govern its behavior absolutely [15], and will have implications on content moderation practice and governance. Therefore, in this paper, we explore the implications of setting moderation policies through natural language prompts given to AI systems on community governance.

2 Content Moderation as a Governance Problem↩︎

Content moderation has never been just about the application of a set of textual rules: it is also a community governance practice [5], [16], [17]. Moderators deliberate and discuss edge cases; formalized appeals processes determine how community members can seek recourse. Additionally, mechanisms to ensure meaningful transparency of implemented rules are vital [18], [19]. The process of moderation also needs to be able to adapt to the evolution of norms in the moderated space, leaving room for changing definitions, contexts, and understandings to emerge [12]. These cannot be described by one centralized way of knowing, as the same words (can) carry different meaning across communities and contexts: “Words change depending on who speaks them; there is no cure”. [20]

These sense-making operations have historically been facilitated through human processes, but with the advent of generative AI systems paradigms in moderation, practices like deliberation by multiple moderators over specific edge cases could be outsourced. Therefore, depending on how LLM-based moderation is implemented, it risks disempowering communities by destructing their influence and recourse options on moderation practices.

3 Moderation through Policy-by-Prompt↩︎

The seemingly most accessible approach to integrating LLMs into moderation is called “policy-as-prompt”. The moderation policy itself is encoded as a natural-language instruction  [8]. This prompt is directly supplied as a system instruction [21] to a general-purpose LLM, which is asked to apply it to a given piece of content. Policy-as-prompt is appealing, as moderation could be reconfigured by simply editing a sentence in the prompt instead of retraining a model. Thus, it promises rapid adaptation to emerging harms, new policies by organizations or jurisdictions, or a shift in platform or community norms. Yet, the swiftness of adaptation is dictated by interdisciplinary writing processes of the new prompts as well as their evaluations and trials before deployment. However, as outlined in previous research [15], system instructions are not reliable enough to give governance guarantees towards goals such as alignment, performance, or robustness of the system.

An additional limiting factor of these approaches is the presence of what are sometimes described as “prompt stacks” [21], [22]. Prompts are processed according to a hierarchical order that prioritizes those set by foundation-model developers before downstream developers, API users, or interface end-users. This prioritization means that moderation instructions (e.g., asking the LLM to apply rules to a specific piece of content) ought to only be processed if they align with the input from prompt layers before. Higher priority prompts should not be able to be overridden by lower priority ones [23]. Therefore, what looks like writing hard rules is, in practice, adding one instruction to a possibly unaligned hierarchy.

For generative AI approaches in content moderation, the mixture of technological limitations [24], [25], well-known evasion or circumvention tactics (e.g., prompt injections) [26], [27], insufficient guardrail effectiveness [28], [29], and other behavior stemming from scaffolding around the foundation model [30], [31], ultimately means that policy-as-prompt approaches on their own are not stable enough to deliver on the envisioned guarantees. As prompt governance unfolds further, prompts have to be supplemented with robust evaluations and other governance measures to have a chance at easing content moderation burdens. Still, the seeming ease of implementation and oversight incentivizes the application of policy-as-prompt.

4 Possible Impacts of Policy-as-Prompt↩︎

In addition to prompt-based moderation having specific technical risk and harm possibilities, the removal of human expertise and accountability from moderation processes will be of consequence. We trace both kinds of impacts across two moderation paradigms.

4.1 Centralized Moderation↩︎

On platforms where moderation is centralized, e.g., done by a governing body, expertise plays a vital role. Policy specialists draft guidelines that then get tested and adjusted with the help of moderators and domain experts [8], [32], ideally in dialogue with community members [33]. This process shapes understanding of the policy that moderators then have to apply in practice. With policy-as-prompt regimes, a deliberated policy is compressed into a single artifact translated not for a human moderator but for an AI system. As [8] put it: “Writing for machines relies on machine interpretation”. New AI policy translation is therefore not focused on translating between organizations and communities, but between organizations and AI. The AI prompt now carries all of the nuance once passed through levels of judgment and competing incentives [34]. This translation additionally might not work as intended due to limitations of language operationalization of stochastic systems in general [8], [15].

4.2 Decentralized Moderation↩︎

Self-governing communities rely heavily on volunteers and a contextual understanding of their community. Volunteer moderators write their own guidelines, deliberate over them in (mostly) transparent ways, and local norms evolve with the community and their discussions [35]. A prompted moderator AI could change community dynamics: the rules could still be written, revised, and understood by the community, then defined in a (possibly non-transparent) prompt set and applied by AI. Some moderators may not want to disclose instructions, because of worries about ‘easier circumvention’ [22]. This also changes rule enforcement dynamics. When a moderator from the community applies a rule, this is based on context and knowledge of the community. By no longer performing this task, the community members lose vital expertise. This would impact the feedback loop in self-governance and potentially lead to less effective and situated moderation. The rules may even be changed in their formulation to account for misinterpretations by AI [36], [37]. The community guidelines become what the generative AI system can operationalize and change in parallel to the development of LLMs [15], [38]. Change in community norms will then (in part) be governed by the affordances of technological developments.

5 Considerations on Policy-as-Prompt Practices↩︎

Policy-as-prompt approaches already exist, and are likely to increase across organizations due to perceived ease of intervention and accessibility [15]. As such, some limitations and implications should be taken into consideration when designing such systems.

First, given the technical limitations and idiosyncrasies of generative AI models, with the possibility of prompt injection and (emergent) circumvention methods, they should generally be handled with caution in content moderation [25], [39], [40]. As such, if policy-as-prompt is implemented, it needs to go beyond specifying words to go into system instructions. Writing a prompt and giving it to a general-purpose model is simply not robust enough at steering system behavior [15]. Prompts need to be evaluated on their impact on model behavior and judgment, as well as their effects across the prompt stack, e.g., through sensitivity analyses.

Second, technical evaluation of policy-as-prompt will still not give governance guarantees. The approach needs to be embedded in robust governance structures that can deliver meaningful disclosure, agency, and accountability [18], [22], [41] to the governed communities.

Third, the inability of generative AI systems to take responsibility for their outputs limits the content moderation tasks they can be deployed for. Moderation actions should be contestable and/or appealable. As it stands, LLMs cannot take the required accountability for their (moderation) actions. Therefore, they should only assist in human decisions, or supplement human deliberation processes.

6 Conclusion↩︎

Although AI might have its place in content moderation, using policy-as-prompt on its own is not enough. Moderation is highly dependent on the fluidity of language, and, thus, on context, time, and intention. As such, content moderation necessitates various sense-making operations, which include deliberation, appeals, and contextual interpretation of broader community values. Outsourcing these mechanisms (even partially) to LLMs has a profound impact. Not only because they take some of these sense-making operations out of human hands, but also because LLM “decisions” are based on a warped mirror of past language and its context, instead of current use of language within the specified context. As such, looking forward requires acknowledging that AI can only look backward, which both justifies and necessitates human-driven processes in content moderation.

References↩︎

[1]
A. Hunsberger, LLM Content Moderation: Implementation Guide for Trust & Safety Teams.” Accessed: Jun. 23, 2026. [Online]. Available: https://musubilabs.ai/post/the-top-challenges-of-using-llms-for-content-moderation-and-how-to-overcome-them.
[2]
T. Kuo, A. Hernani, and J. Grossklags, “The Unsung Heroes of Facebook Groups Moderation: A Case Study of Moderation Practices and Tools,” Proceedings of the ACM on Human-Computer Interaction, vol. 7, no. CSCW1, pp. 1–38, Apr. 2023, doi: 10.1145/3579530.
[3]
C. Newton, “The secret lives of Facebook moderators in America.”
[4]
“Scroll. Click. Suffer: The Hidden Human Cost of Content Moderation and Data LabellingEquidem.” Accessed: Jun. 23, 2026. [Online]. Available: https://equidem.org/reports/scroll-click-suffer-the-hidden-human-cost-of-content-moderation-and-data-labelling/.
[5]
M. Alizadeh, F. Gilardi, E. Hoes, K. J. Klüser, M. Kubli, and N. Marchal, “Content Moderation As a Political Issue: The Twitter Discourse Around Trump’s Ban,” Journal of Quantitative Description: Digital Media, vol. 2, Oct. 2022, doi: 10.51685/jqd.2022.023.
[6]
Z. McMahon Liv andZoe Kleinman and C. Subramanian, [Accessed 22-06-2026]Meta to replace ’biased’ fact-checkers with moderation by users — bbc.com.” 2025, [Online]. Available: https://www.bbc.com/news/articles/cly74mpy8klo.
[7]
J. van de Kerkhof, [Accessed 22-06-2026]Musk, Techbrocracy, and Free Speech — verfassungsblog.de.” 2025, [Online]. Available: https://verfassungsblog.de/musk-techbrocracy-and-free-speech/.
[8]
K. Palla et al., arXiv:2502.18695 [cs]“Policy-as-Prompt: Rethinking Content Moderation in the Age of Large Language Models.” arXiv, Feb. 2025, doi: 10.48550/arXiv.2502.18695.
[9]
W. C. Yew et al., “Dynamic Content Moderation in Livestreams: Combining Supervised Classification with MLLM-Boosted Similarity Matching,” in Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.1, Apr. 2026, pp. 2528–2538, doi: 10.1145/3770854.3783936.
[10]
H. Inan et al., “Llama guard: LLM-based input-output safeguard for human-AI conversations.” 2023, [Online]. Available: https://arxiv.org/abs/2312.06674.
[11]
OpenAI, “Introducing gpt-oss-safeguard.” 2025, [Online]. Available: https://openai.com/index/introducing-gpt-oss-safeguard/.
[12]
M. Franco, O. Gaggi, and C. E. Palazzi, “Integrating Content Moderation Systems with Large Language Models,” ACM Transactions on the Web, vol. 19, no. 2, pp. 1–21, May 2025, doi: 10.1145/3700789.
[13]
J. F. Gomez, C. Machado, L. M. Paes, and F. Calmon, “Algorithmic Arbitrariness in Content Moderation,” in The 2024 ACM Conference on Fairness, Accountability, and Transparency, Jun. 2024, pp. 2234–2253, doi: 10.1145/3630106.3659036.
[14]
F. Shahid, M. Elswah, and A. Vashistha, arXiv:2501.13836 [cs.CL]“Think Outside the Data: Colonial Biases and Systemic Issues in Automated Moderation Pipelines for Low-Resource Languages.” arXiv, Aug. 2025, doi: 10.48550/arXiv.2501.13836.
[15]
A. Neumann, H. Sargeant, and J. Singh, “Prompt governance? On governing technologies governed by natural language,” in The 2026 ACM conference on fairness, accountability, and transparency (FAccT ’26), 2026, doi: 10.1145/3805689.3806763.
[16]
R. Gorwa, R. Binns, and C. Katzenbach, “Algorithmic content moderation: Technical and political challenges in the automation of platform governance.” 2020, Accessed: Feb. 17, 2025. [Online]. Available: https://journals.sagepub.com/doi/full/10.1177/2053951719897945.
[17]
K. Klonick and K. Klonick, “The New Governors: The People, Rules, and Processes Governing Online Speech,” Harvard Law Review, Mar. 2017, Accessed: Jun. 22, 2026. [Online]. Available: https://www.semanticscholar.org/paper/The-New-Governors%3A-The-People%2C-Rules%2C-and-Processes-Klonick-Klonick/cb52e32499d15ad3624a228a926416a3db14deb7.
[18]
C. Norval, K. Cornelius, J. Cobbe, and J. Singh, “Disclosure by Design: Designing information disclosures to support meaningful transparency and accountability,” in Proceedings of the 2022 ACM Conference on Fairness, Accountability, and Transparency, Jun. 2022, pp. 679–690, doi: 10.1145/3531146.3533133.
[19]
N. P. Suzor, S. M. West, A. Quodling, and J. York, “What Do We Mean When We Talk About Transparency? Toward Meaningful Transparency in Commercial Content Moderation,” International Journal of Communication, vol. 13, pp. 18–18, Mar. 2019, Accessed: Jun. 22, 2026. [Online]. Available: https://ijoc.org/index.php/ijoc/article/view/9736.
[20]
M. Nelson, The argonauts. Graywolf Press, 2015.
[21]
E. Wallace, K. Xiao, R. Leike, L. Weng, J. Heidecke, and A. Beutel, arXiv:2404.13208 [cs]“The Instruction Hierarchy: Training LLMs to Prioritize Privileged Instructions.” arXiv, Apr. 2024, doi: 10.48550/arXiv.2404.13208.
[22]
A. Neumann, Y. Pi, and J. Singh, arXiv:2603.00089 [cs]“Who Controls the Conversation? User Perspectives On Generative AI (LLM) System Prompts.” Feb. 2026, doi: 10.1145/3772318.3791726.
[23]
A. Neumann, E. Kirsten, M. B. Zafar, and J. Singh, “Position is Power: System Prompts as a Mechanism of Bias in Large Language Models (LLMs),” in Proceedings of the 2025 ACM Conference on Fairness, Accountability, and Transparency, Jun. 2025, pp. 573–598, doi: 10.1145/3715275.3732038.
[24]
Y. Salini and J. HariKiran, “Sarcasm Detection: A Systematic Review of Methods and Approaches,” in 2023 3rd International Conference on Smart Data Intelligence (ICSMDI), Mar. 2023, pp. 15–22, doi: 10.1109/ICSMDI57622.2023.00012.
[25]
D. Kumar, Y. AbuHashem, and Z. Durumeric, Version Number: 2“Watch Your Language: Investigating Content Moderation with Large Language Models.” arXiv, 2023, doi: 10.48550/ARXIV.2309.14517.
[26]
F. Perez and I. Ribeiro, arXiv:2211.09527 [cs.CL]“Ignore Previous Prompt: Attack Techniques For Language Models.” arXiv, Nov. 2022, doi: 10.48550/arXiv.2211.09527.
[27]
K. Greshake, S. Abdelnabi, S. Mishra, C. Endres, T. Holz, and M. Fritz, “Not What You’ve Signed Up For: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection,” in Proceedings of the 16th ACM Workshop on Artificial Intelligence and Security, Nov. 2023, pp. 79–90, doi: 10.1145/3605764.3623985.
[28]
Y. Dong et al., “Safeguarding large language models: A survey,” Artificial Intelligence Review, vol. 58, no. 12, p. 382, Oct. 2025, doi: 10.1007/s10462-025-11389-2.
[29]
G. Bertollo, N. Bodemir, and J. Burgess, arXiv:2510.16005 [cs.CR]“Breaking Guardrails, Facing Walls: Insights on Adversarial AI for Defenders & Researchers.” arXiv, Oct. 2025, doi: 10.48550/arXiv.2510.16005.
[30]
A. Neumann and J. Singh, arXiv:2603.00088 [cs]AI Safety Evaluations Need To Consider Cascading Effects.” arXiv, Feb. 2026, doi: 10.48550/arXiv.2603.00088.
[31]
Q. Zhan, Z. Liang, Z. Ying, and D. Kang, arXiv:2403.02691 [cs.CL]InjecAgent: Benchmarking Indirect Prompt Injections in Tool-Integrated Large Language Model Agents.” arXiv, Aug. 2024, doi: 10.48550/arXiv.2403.02691.
[32]
M. Ruckenstein and L. L. M. Turunen, “Re-humanizing the platform: Content moderators and the logic of care,” New Media & Society, vol. 22, no. 6, pp. 1026–1042, Jun. 2020, doi: 10.1177/1461444819875990.
[33]
B. Schaffner et al., “"Community Guidelines Make this the Best Party on the Internet": An In-Depth Study of Online PlatformsContent Moderation Policies,” in Proceedings of the CHI Conference on Human Factors in Computing Systems, May 2024, pp. 1–16, doi: 10.1145/3613904.3642333.
[34]
S. M. West, “Raging Against the Machine: Network Gatekeeping and Collective Action on Social Media Platforms,” Media and Communication, vol. 5, no. 3, pp. 28–36, Sep. 2017, doi: 10.17645/mac.v5i3.989.
[35]
C. L. Cook, A. Patel, and D. Y. Wohn, “Commercial Versus Volunteer: Comparing User Perceptions of Toxicity and Transparency in Content Moderation Across Social Media Platforms,” Frontiers in Human Dynamics, vol. 3, p. 626409, Feb. 2021, doi: 10.3389/fhumd.2021.626409.
[36]
H. Yakura et al., arXiv:2409.01754 [cs.CY]“Empirical evidence of Large Language Model’s influence on human spoken communication.” arXiv, Jul. 2025, doi: 10.48550/arXiv.2409.01754.
[37]
M. Adkins, arXiv:2401.06382 [cs.HC]“What should I say? – Interacting with AI and Natural Language Interfaces.” arXiv, Jan. 2024, doi: 10.48550/arXiv.2401.06382.
[38]
H. Kim, H. Yi, J. Bae, and Y. Kim, arXiv:2602.22790 [cs.CL]“Natural Language Declarative Prompting (NLD-P): A Modular Governance Method for Prompt Design Under Model Drift.” arXiv, Feb. 2026, doi: 10.48550/arXiv.2602.22790.
[39]
J. Zeng, Q. Guan, A. Matamoros-Fernández, and X. Liu, “How do multi-modal large language models understand non-English visual hate? Insights from studying hate speech in Chinese-speaking communities on Instagram,” Platforms & Society, vol. 2, p. 29768624251383735, Sep. 2025, doi: 10.1177/29768624251383735.
[40]
D. Hartmann, A. Oueslati, D. Staufer, L. Pohlmann, S. Munzert, and H. Heuer, “Lost in Moderation: How Commercial Content Moderation APIs Over- and Under-Moderate Group-Targeted Hate Speech and Linguistic Variations,” in Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems, Apr. 2025, pp. 1–26, doi: 10.1145/3706598.3713998.
[41]
K. Lewicki, M. S. A. Lee, J. Cobbe, and J. Singh, “Out of Context: Investigating the Bias and Fairness Concerns of Artificial Intelligence as a Service,” in Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems, Apr. 2023, pp. 1–17, doi: 10.1145/3544548.3581463.