Manufactured Divisiveness: Decomposing the Hostile Content of Seven Social Media Influence Operations

Emilio Ferrara
Thomas Lord Department of Computer Science
University of Southern California
emiliofe@usc.edu


Abstract

State-backed influence operations are routinely measured as high-prevalence sources of “hate” and “toxicity.” We argue those rates rest on a measurement error: the detectors behind them are validated to catch a broader definition inclusive of hostility or divisiveness aimed at an out-group, and so over-attribute hate to content better described as partisan or geopolitical invective. Across 25.08M tweets from seven government-attributed campaigns in the Twitter Information Operations archive (8,275 accounts), we separate hate from the other forms of divisiveness. We first validate a two-prompt LLM-based detector, matching human labels at Cohen’s \(\kappa=0.82\), to identify the broader hostility; we then develop an auditable rule, agreeing with an expert at \(\kappa=0.52\), to further classify this content (5,457 posts) into three sub-categories. About 50.1% are identity-based attacks on people, whereas 30.4% are partisan attacks and 19.5% invective against states and their foreign policy. Reporting all of it as hate therefore overstates hate roughly twofold; only 18.7% is both identity-based and dehumanizing or inciting. Six of seven campaigns sort into three regimes that a single “hate” rate flattens, namely identity hate (RU-op and IRA, both Russia-attributed), geopolitical invective (both Iran operations), and partisan divisiveness (both Venezuela operations). We call the shared product manufactured divisiveness. The line to separate these constructs itself remains unsettled: on the hardest cases three independent human experts agree only moderately (pairwise \(\kappa=0.37\)\(0.50\)), and the best of nineteen LLM models tops out at \(\kappa=0.601\) against the experts’ majority. Our findings can help redefine the study of hate in the context of influence campaigns and broader online discourse.

1 Introduction↩︎

State-backed influence operations on social platforms are widely described, in both policy discourse and the research literature, as sources of “hate” and “toxicity,” and more broadly of negative and inflammatory content [1]. This framing is increasingly operationalized directly: an automated detector is applied to operation content, items are flagged, and the flagged share is reported as a hate or toxicity rate. For instance, recent work scores tens of millions of posts across many state actors with a toxicity classifier and reports the toxic share as a headline measure [2], [3]. These rates are read as substantive measures of how hateful an operation is, compared across actors as evidence of platform harm. Their validity depends on a question seldom stated explicitly: what construct does the detector measure?

A long line of work in hate-speech detection has warned that “hate” is routinely conflated with the broader categories of offensive, abusive, or divisive language [4][9]. A detector validated to separate hostile, divisive out-group targeting from neutral content is a useful instrument, but not, without further evidence, a hate-speech detector: if the validated construct is broad, reporting the flagged share as “hate” over-attributes hate speech to content that is partisan or geopolitical invective.

We pursue this concern for a corpus of seven government-attributed influence campaigns (25.08M tweets, 8,275 accounts) from the Twitter Information Operations archive, with a two-stage design that keeps the broad-construct detector and the typing step distinct and validates each separately. First, a gate detects the broad construct (hostile or divisive out-group targeting), validated against human gold at Cohen \(\kappa=0.82\), precision \(0.96\) on the original 100-item set. Second, an auditable rule over a frozen, LLM-derived characterization taxonomy types each positive item as identity-directed hate, partisan divisiveness, or geopolitical invective, and flags a dehumanizing/inciting severity tier. The rule’s typing agrees with an independent expert (110 items) at \(\kappa=0.52\) (moderate), within the band at which two independent human coders agree on the contested boundary (\(\kappa=0.44\)). The typing is a rule over a taxonomy, not a human judgment that an item “is hate.”

Typing the positives shows compositional heterogeneity: identity hate is only half, the defensible hate-speech share is just 18.7%, and the mix differs systematically across operations scored comparably under the broad gate (six of seven sort into three construct regimes; §5.2). We term the shared product manufactured divisiveness, with identity hate the narrower construct nested within it and the dehumanizing/inciting core narrower still.

1.0.0.1 Scope.

All findings concern these seven specific, already-detected and attributed campaigns and are associational throughout: we report co-occurrence, divergence, and concentration, not causation. The absence of a construct here is not evidence of its absence in any other operation.

1.0.0.2 Contributions.

  • (C1) On these seven operations, the validated “hate” gate measures broad divisiveness. A gate validated against human gold as a detector of hostile/divisive out-group targeting is not a hate-speech detector: under our stated typing criteria, only \(\sim\)​19% of its positives are identity-directed and dehumanizing/inciting. On a boundary-enriched gold, three independent experts agree only moderately (pairwise \(\kappa=0.37\)\(0.50\)), each expert is best matched by a different model, and no model exceeds \(\kappa=0.601\) against the experts’ majority: the divisive/hate line has no annotator-stable reading a broad flag could be assumed to capture (Section 5.3).

  • (C2) Positive composition varies by operation, sorting into three construct regimes. Operations scored equally “hateful” under the broad gate diverge in positive composition along construct type: RU-op and IRA are majority identity hate (68% and 61%), both Iran-attributed operations are majority geopolitical invective (IR-op-A 64%, IR-op-B 51%), and both Venezuela-attributed operations are majority partisan divisiveness (94% and 75%, predominantly non-hate under the typing rule). This sorting holds for six of seven; the seventh (BD-op) splits evenly and is the honest exception. We report these as descriptive compositions of the positive set, not as a tested discriminative classifier (Section 5.2).

  • (C3) Content-layer divergence attenuates as the construct is narrowed. The pattern “content (target/narrative) diverges, form (intensity/dehumanization) converges” holds, but part of the original broad-construct contrast came from the identity-versus-political gap itself: as the construct is narrowed, content-layer effect sizes fall at each step while form-layer associations stay small. We read this as an effect-size trend, not a strict law, given the multi-label dimensions and small identity-subset samples (Section 5.6).

  • (C4) Concentration is construct-specific. Accounts with high broad-construct concentration do not all stay concentrated on identity hate: one operation’s concentration falls once attention is restricted to the identity core (Section 5.7).

  • (C5) The RU-op prevalence outlier persists at the defensible core. At the dehumanizing-identity-hate core, only RU-op exceeds a non-trivial prevalence floor; the RU-op elevation persists under the narrowed construct and becomes more pronounced (Section 5.5).

2 Related Work↩︎

2.0.0.1 Influence-operation content.

Studies of state-backed operations have characterized troll content, specialization, and cross-platform influence [10][14], and the 2016 Russian interference campaign in particular has been analyzed as a manipulation trace [14], [15]. The partisan audiences these campaigns court are themselves measurable at scale: political leaning can be inferred from profile language and retweet-network structure [16], which grounds the audience-tuning claims we make below. More recent comparative analyses of the same Twitter Information Operations archive characterize many state operations jointly (by behavioral fingerprint, dissemination, and cross-operation coordination structure) [17], [18] rather than by the type of hostile content they produce. Coordinated inauthentic behavior, large-scale platform manipulation, and the broader bot ecosystem provide the structural backdrop [19][22]. Much of this literature reports a single hostility or toxicity rate rather than decomposing the construct; our contribution is to type the construct before quantifying it.

2.0.0.2 Hate, offensive, and divisive language.

The detection literature has repeatedly documented the conflation of hate with offensive or abusive language [4], [5], [23], proposed typologies that separate abuse along target and form axes [8], surveyed the fragmentation of definitions [6], and traced how annotation choices propagate into models [7], [24][26]. Implicit and boundary cases are especially hard [27]. A parallel line in political communication distinguishes political incivility from group-targeting intolerance [28], and shows that divisive constructs such as fear speech can target groups while evading toxicity detectors [29], precisely the heterogeneity a single broad flag collapses. The instability extends within annotators: on RLHF preference data, the same annotator often rates semantically equivalent harm prompts inconsistently, and filtering inconsistent annotators flips the majority harm label on 18.6% of prompts [30]. That three expert annotators agree pairwise at only \(\kappa=0.37\)\(0.50\) on a boundary-enriched gold, while model–expert agreement ranges from near zero to substantial depending on which expert is asked, is consistent with this body of work and with a parallel cross-lingual audit reporting an LLM-annotator ceiling at \(\kappa\approx0.42\) on an independent corpus [31]: the divisive/hate boundary is genuinely contested, a property of the construct rather than of our operation set.

2.0.0.3 LLMs as annotators.

Large language models have been shown to be competitive with crowd workers for text annotation [32] and useful for computational social science [33], though their annotation reliability is task-dependent and benefits from task-specific human validation and careful rule design, especially for content-moderation judgments [34][38]. We use an LLM gate and characterization, but, consistent with the validity concerns above [39], we validate the broad gate against human gold and treat the typing as a transparent rule, not a model verdict.

2.0.0.4 Moral and affective framing.

Moral-foundations theory [40], work on moralized-emotional diffusion [41], and evidence that out-group animosity is associated with elevated engagement [42] motivate our affect cross-validation, in which LLM affect labels are checked against independent moral and emotion lexica. The distinction we draw between identity-directed hate and partisan divisiveness parallels political-science work treating partisan animosity as a distinct, identity-rooted form of out-group hostility, separable from group-based hate [43], [44].

2.0.0.5 Companion work.

A sibling study examines cross-lingual hate in organic platform content with an actor-less, layered cultural-contingency framing [31]; coordination structure in the same operations is studied separately [45]. The present paper is the cross-operation contribution focused on measurement validity: typing the construct before counting it.

3 Data↩︎

We analyze seven government-attributed campaigns released through the Twitter Information Operations archive [46], totaling 25.08M tweets across 8,275 accounts. We refer to operations by pseudonymous labels with short descriptors only; we never name a specific government as the proven actor beyond “state-attributed campaigns released through the Twitter Information Operations archive.” The descriptors carry only the archive’s own coarse attribution and serve to keep the analysis at the level of measured content rather than asserting attribution claims of our own. Table 1 summarizes per-operation volume and dominant script.

Table 1: Per-operation volume and dominant script. Tweet totals anddistinct-account counts are computed over the analyzed archive corpus(COUNT(DISTINCT userid) per operation); rows sum to the totals shown.Pseudonymous labels are used throughout.
Operation (descriptor) Tweets Accounts Dominant script(s)
RU-op (anti-Muslim/US-wedge) 765,246 359 Latin
IRA (US-partisan) 8,768,633 3,479 Cyrillic, Latin
IR-op-A (anti-Israel/Saudi) 1,122,936 660 Arabic/Persian, Devanagari
IR-op-B (geopolitical/S-Asia) 4,447,056 2,201 Arabic/Persian, Latin
VE-op-A (domestic ES) 8,961,788 987 Latin (Spanish)
VE-op-B (US-facing EN) 984,980 578 Latin (English)
BD-op (domestic) 26,214 11 Bengali
Total 25,076,853 8,275

3.0.0.1 Tier-N data handling.

Tier-N denotes our most restrictive reporting tier: all outputs in this paper are labels and counts only. We report no raw tweet text, no example slurs or slur lists, no real account handles, and no full numeric account identifiers anywhere. Operations and accounts are referred to by pseudonymous labels.

4 Methods↩︎

4.1 Broad-construct gate↩︎

A gate built on an instruction-tuned LLM (Qwen2.5-7B) scores each (tweet, target) pair for the broad construct of hostile or divisive out-group targeting. An item counts as positive only under a two-prompt consensus: both a permissively worded prompt (CLEAN) and a stricter one (STRICT2) must flag it (CLEAN \(\wedge\) STRICT2). The gate scored 40,464 (tweet, target) rows; the consensus-positive census (qcons\(=1\)) is 5,457 rows, with 2,634 single-prompt disagreements and 32,373 consensus-negative. The conservative two-prompt consensus is the census used throughout. This broad construct is what the original 100-item human gold validates; the downstream typing rule carries its own expert validation (Section 4.3).

4.2 Characterization↩︎

Each positive item is characterized along an 11-dimension schema (target group, influence-operation narrative, threat frame, moral foundation, emotion, dehumanization, stance intensity, and others). Of the 5,457 characterized items, 850 non-English items were machine-translated to English (NLLB, a multilingual translation model); the affect cross-validation below runs on the 5,332 items carrying usable English text (native or pivoted), and we test that affect signal is present pre-translation.

4.3 Construct typing↩︎

We type each positive item by an auditable rule over the frozen target-plus-narrative taxonomy produced by characterization. The rule is not a human “is-it-hate” judgment and does not re-label the human gold; it deterministically maps taxonomy values to one of three constructs:

  • identity_hate: the target is a protected or ascriptive group (religion, ethnicity/race, nationality-as-a-people, immigration, gender/sexuality), or the narrative is identity-diagnostic. This is hate speech proper.

  • political_divisive: the target is a political affiliation, party, ideology, or named politician. Inflammatory, but not hate speech, a distinction that mirrors the incivility-versus-intolerance boundary in political communication [28]. Because partisan animosity is itself identity-rooted [47] and partisan invective can carry identity-coded (dogwhistle) content [48], classifying every such item as non-hate makes the 50.1% identity share a lower bound rather than a point estimate.

  • state_geopolitical: the attack is on a state, regime, or foreign policy. Foreign-policy invective.

Orthogonally, a hard_core severity flag is set when content is dehumanizing (animalistic or mechanistic) or inciting (a call to exclude or harm). Two target categories that are systematically ambiguous between a group of people and a polity (ethnic/national out-groups and immigrants/refugees) are not assigned by target alone but arbitrated by the narrative; and an item otherwise typed as identity hate is re-typed to state_geopolitical when its text names a terrorist organization or a foreign state (e.g.ISIS, the Turkish military) without attacking the religious group as people, since such content is security or foreign-policy invective rather than group hate. A provenance audit characterizes the basis of each assignment: 69.0% of assignments are target-decisive, 30.9% are narrative-diagnostic, and only 7 items (0.1%) are resolved by a fallback rule; the text-level override re-types 10.8% of items.

4.3.0.1 What the typing does and does not claim.

The rule’s decisive input is the comparatively objective target-group attribute rather than a subjective “is-it-hate” verdict. This resolves an apparent tension with the contested annotator agreement reported below (Section 5.3): the is-it-hate verdict on borderline content is contested even between expert annotators, which the rule does not attempt; it partitions instead by whom the content targets, a more reliable axis. The resulting split is best read as the rule’s partition under these stated definitions, not a ground-truth item-level measurement. We validate the rule against an independent expert typing of 110 items (stratified across the decisive and narrative-arbitrated cases, blinded to the rule’s inputs): it agrees with the expert at Cohen \(\kappa=0.52\) (accuracy 68%), moderate and above its pre-refinement value (\(\kappa=0.42\)), with agreement highest on partisan and antisemitic targets (per-target accuracy 0.78–0.90) and lowest on the ambiguous ethnic and immigration targets that motivated narrative arbitration. To establish that this agreement reflects the construct rather than one coder’s idiosyncrasy, a second independent expert relabeled an 80-item subset enriched for the contestable boundary cases, blinded identically. The two human coders agree at Cohen \(\kappa=0.44\) (95% CI \([0.26, 0.60]\)), and on this harder subset the rule agrees with each human at least as strongly as the humans agree with each other (\(\kappa=0.45\) and \(0.55\)); the rule therefore sits within the human–human agreement band, so the moderate ceiling is a property of the divisive/hate boundary rather than of the rule. Where the two experts agree (47 of 74 items), the rule matches that consensus at \(\kappa=0.71\), so the residual disagreement concentrates on exactly the items the experts themselves contest.

4.4 Validation↩︎

The broad gate is validated against a human gold set. On the original 100-item representative gold, the broad gate agrees with the human at Cohen \(\kappa=0.82\) and precision \(0.96\). On a boundary-enriched 102-item gold that over-samples the divisive/hate line, three independent expert annotators each labeled every item under identical blinding (one, Expert-1, with a written per-item rationale), and a panel of nineteen models each judge the same blinded English text; we score Cohen \(\kappa\) [49] against each expert and against the experts’ two-of-three majority, and report human–human reliability as pairwise Cohen \(\kappa\) (with 5,000-resample bootstrap intervals) and Fleiss \(\kappa\). The experts coded the broad construct, so this measures agreement on the broad gate, not on the typing rule. The representative \(\kappa=0.82\) and the boundary-enriched agreements are not directly comparable (different samples and references) and should be read as two validation regimes, not one instrument deteriorating.

4.5 Affect-label cross-validation↩︎

We test whether the LLM affect labels track independent dictionary signal by comparing them, on 5,332 English-text confirmed items, against MFD2.0 [50] moral-foundations, NRC [51] and LIWC2015 [52] emotion, and (for Russian originals) rusentilex [53] lexica, using a one-sided Mann–Whitney \(U\) test with rank-biserial effect size and BH-FDR correction.

5 Findings↩︎

5.1 Positive composition↩︎

Of the 5,457 positive items, the typing rule assigns 50.1% to identity hate, 30.4% to partisan divisiveness, and 19.5% to geopolitical invective (Table 2, “All”). These are compositions within the gate-positive set (a cue-and-target-selected hostile tail), not corpus-wide rates; prevalence floors over random strata are reported separately in Section 5.5. The split is not an artifact of the conservative two-prompt gate. Re-typing under a permissive CLEAN-only census (\(+2{,}140\) items) or a STRICT2-only census (\(+494\) items) leaves identity hate a plurality but never a majority (\(41.5\%\) and \(48.8\%\), versus \(50.1\%\) under the two-prompt consensus). Relaxing the gate shifts composition toward partisan and geopolitical content; the conservative gate is therefore the most identity-favorable choice, and the over-attribution finding is, if anything, stronger under a looser gate. Only 21.0% (1,146 items) carry the dehumanizing/inciting hard_core flag, and the narrowest defensible hate-speech core (identity-directed and hard_core) is just 18.7% (1,023 items). Under these typing criteria, treating the broad flag as a “hate rate” over-states hate-speech prevalence for these seven operations by roughly 2.0\(\times\): counting all identity-typed content as hate still halves the broad flag (\(5{,}457\) vs. \(2{,}733\) items). This over-statement is robust to where the typing rule draws the identity boundary. We sweep that boundary between two extremes: the most identity-favorable choice counts every ascriptive-group target as identity hate, with no narrative or text-level routing out of identity (\(66.9\%\) identity, \(3{,}653\) items); the most identity-conservative routes the systematically ambiguous ethnic-national and immigration targets out of identity (\(45.1\%\), \(2{,}461\) items). Across this range the identity share moves only within \(45.1\)\(66.9\%\) and the over-statement only within \(1.5\)\(2.2\times\); no defensible boundary brings the broad flag close to a hate rate. Measured against the narrowest defensible core (identity and dehumanizing/inciting) the gap widens to 5\(\times\) (\(5{,}457\) vs.\(1{,}023\) items), though that figure compares the broad construct to its strictest sub-core and is best read as an upper bound.

Figure 1: Construct composition of positive content, by operation (row-normalized; n positive items per operation in parentheses). Bars decompose each operation’s hostile content into identity-based hate, partisan divisiveness, and geopolitical invective. The typing is a rule applied to the frozen target/narrative taxonomy over items the human-validated broad gate admitted, not an item-level human classification of hate. The composition varies by operation: RU-op and IRA are predominantly identity-typed, both Iran-attributed operations are weighted toward geopolitical invective, and VE-op-A is almost entirely partisan.

5.2 Per-operation construct mix↩︎

Operations with comparable positive rates under the broad gate diverge in composition along construct type, sorting into three regimes (Fig. 1, Table 2). Two operations are majority identity hate: RU-op (68.0%) and IRA (60.9%). Both Iran-attributed operations are majority geopolitical invective (IR-op-A 63.8% geopolitical alongside 32.0% identity; IR-op-B 50.9% geopolitical, 29.5% identity). Both Venezuela-attributed operations are majority partisan divisiveness: VE-op-A is 94.2%, with identity hate accounting for only 3.9% of its positive items, and VE-op-B is 75.3%. The seventh, BD-op, splits evenly between identity hate (48.4%) and partisan divisiveness (48.4%). The broad “hate” label collapses these compositional differences.

That the regimes align with the polities each operation targets (the Iran-attributed operations attacking states and foreign policy, the Venezuela-attributed operations contesting domestic partisan politics, RU-op and IRA running identity wedges) is unsurprising on its own: different actors have different adversaries, and the seven operations carry roughly three independent state attributions, so the regimes rest on a handful of actor families rather than seven independent draws. The contribution is not that operations differ in target, which is expected, but that a single broad “hate” rate scores all seven alike while their hostile content belongs to categorically different constructs, a distinction the scalar rate discards (not a claim that construct type is determined by actor).

These compositions are row-normalized over each operation’s positives, a selected tail of very different sizes (tens of items for BD-op to thousands), so the smallest operations’ mixes are unstable; we present it as a descriptive signature, not a tested discriminator (Section 7).

Table 2: Per-operation construct mix (row %), over the 5,457 positiveitems.
Operation Identity Partisan Geopolitical
hate (%) divisive (%) (%)
RU-op 68.0 22.2 9.8
IRA 60.9 32.9 6.3
IR-op-A 32.0 4.2 63.8
IR-op-B 29.5 19.6 50.9
VE-op-A 3.9 94.2 1.9
VE-op-B 22.5 75.3 2.2
BD-op 48.4 48.4 3.2
All 50.1 30.4 19.5

5.3 The divisive/hate boundary↩︎

Where does “divisive but not hateful” end and “hate” begin? If that line were sharp, independent annotators would agree on where it falls. They do not: three expert annotators of the same borderline items agree pairwise at only Cohen \(\kappa=0.37\)\(0.50\) (Fleiss \(\kappa=0.45\)), unanimous on 63% of items, and where they diverge, the models diverge with them.

We test this on two human-labeled gold sets. The representative set (100 items) is a random sample of the corpus. The boundary-enriched set (102 items) deliberately over-samples items on the divisive/hate line; three experts independently labeled every item under identical blinding (Expert-1 with a written per-item rationale), and every annotator is scored against each expert and against the experts’ two-of-three majority. The broad gate agrees with the representative gold at \(\kappa=0.82\) but with the expert majority on the boundary-enriched set at only \(\kappa=0.44\): the instrument did not change; the second set is stocked with the hardest calls, where a conservative gate and careful humans part ways.

Would a stronger annotator close that gap? A panel of nineteen models (eleven open-weight, plus larger open-weight and frontier API models) each redid the gate task independently on the boundary-enriched gold, judging the same blinded text (Table 3). Agreement is strikingly heterogeneous, and which model reads the boundary best flips with the expert asked: Gemini-2.5-pro agrees with Expert-1 at \(\kappa=0.705\) but falls to \(0.36\) and \(0.48\) against the other two experts, whose best matches are different models altogether (OLMo-2-7B at \(0.55\); OLMo-3.1-32B at \(0.53\)). Against the two-of-three expert majority no model exceeds \(\kappa=0.601\) (Qwen2.5-14B), with the deployed gate at \(0.44\). There is no capacity ladder: a 7B model (Falcon3) ties the 32B models, and the newest, most heavily aligned Gemini models agree least, increasingly refusing to call borderline content hostile [54] and flagging about 1 in 5 of the items the expert majority flags (Gemini-3.5-flash \(R=0.21\), \(\kappa=0.22\)).

The one substantial model–expert pairing is thus idiosyncratic rather than a capacity effect: Gemini-2.5-pro is the most conservative frontier model and Expert-1 the most conservative coder (30 of 102 positive, against 34 and 42 for the other two experts), and their \(\kappa=0.71\) agreement does not carry over to either other expert. Expert readings of the boundary differ from one another (pairwise \(\kappa=0.37\)\(0.50\)) about as much as the better models differ from any one expert, so no annotator, human or model, supplies a stable reference; a single “hate” rate would inherit whichever reading happened to produce the gold labels, which is why we report composition rather than a rate. The difficulty is specific to these boundary items, not the broad construct as a whole: an eleven-model panel labeling a larger balanced sample agrees at Fleiss \(\kappa=0.47\) on the broad decision, and the strongest challenger reaches \(\kappa=0.72\) on the representative gold (Table 4), so the broad call itself is reproducible even where the line within it is contested.

Agreement among the annotators tells the same story (Fig. 2). Clustering the twenty-two annotators (nineteen models and the three experts) by pairwise Cohen’s \(\kappa\), the experts do not group together: Expert-1 clusters with the conservative Gemini-2.5 models (its nearest neighbor is Gemini-2.5-pro, \(\kappa=0.71\)), Expert-2 sits beside OLMo-2-7B, and Expert-3 joins no block at all. Each expert anchors a different neighborhood, and which annotators fall together is organized by how conservatively each draws the line, not by scale or recency. The capable mid-sized models form a mutually high-agreement block (pairwise \(\kappa\) up to \(0.84\)) that no expert joins.

Figure 2: Inter-annotator agreement on the broad construct over the boundary-enriched 102-item gold: pairwise Cohen’s \kappa among the nineteen models and the three human experts, ordered by hierarchical clustering on 1-\kappa (average linkage), so the dendrogram groups annotators that label the borderline region alike. Red boxes mark the subblocks isolated by a single depth cut of the dendrogram (dashed line). The three experts (bold) do not group together: Expert-1 clusters with Gemini-2.5-pro and Gemini-2.5-flash (nearest neighbor Gemini-2.5-pro, \kappa=0.71), Expert-2 clusters with OLMo-2-7B, and Expert-3 joins no block (nearest neighbors OLMo-3.1-32B and Qwen2.5-14B, \kappa=0.52–0.53). Capable mid-sized models form a large mutually high-agreement block (pairwise \kappa up to 0.84); the newest Gemini models and the {Aya, Gemma, Yi} group each form their own block; the smallest model (Qwen2.5-1.5B) sits apart. Darker cells are lower agreement; model sizes abbreviated, full identifiers in Table 3.
Table 3: Annotator agreement on the broad construct, 102-item boundary-enrichedgold: Cohen’s \(\kappa\) against each of three independent expert annotators (E1= P.G., who annotated with per-item rationales; E2, the study’s first coder; E3= D.R.) and against their two-of-three majority, ranked by the majority column.Rows are the full annotation panel (individual models, their majorityEnsemble, and the deployed two-prompt broad gate) alongsidelarger open-weight and frontier API models run with the byte-identical gate.Column maxima in bold: each expert is best matched by a different model(Gemini-2.5-pro, OLMo-2-7B, OLMo-3.1-32B), no model exceeds \(\kappa=0.601\)against the majority, and the experts themselves agree pairwise at\(\kappa=0.37\)\(0.50\) (Fleiss \(\kappa=0.45\)). Prompt-ablation variantsomitted.
\(\kappa\) vs expert
3-5 Annotator Size E1 E2 E3 Majority
Qwen2.5-14B 14B 0.509 0.452 0.517 0.601
Granite-3.1-8B 8B 0.482 0.481 0.453 0.584
OLMo-2-7B 7B 0.356 0.554 0.427 0.565
Ensemble (majority) 0.502 0.427 0.503 0.543
OLMo-3.1-32B 32B 0.485 0.379 0.531 0.531
Gemini-2.5-pro API 0.705 0.356 0.484 0.531
Phi-4 14B 0.369 0.389 0.481 0.526
Falcon3-7B 7B 0.501 0.489 0.426 0.511
Qwen2.5-32B 32B 0.501 0.407 0.426 0.511
Gemini-2.5-flash-lite API 0.452 0.469 0.382 0.473
Qwen2.5-7B 7B 0.397 0.366 0.400 0.440
Broad gate (2-prompt) 0.420 0.338 0.435 0.435
Aya-Expanse-8B 8B 0.370 0.404 0.429 0.429
Mistral-7B-v0.3 7B 0.387 0.302 0.229 0.357
Gemini-2.5-flash API 0.452 0.257 0.336 0.336
OLMo-3-7B 7B 0.270 0.279 0.245 0.331
Gemma-2-9B 9B 0.310 0.333 0.329 0.329
Yi-1.5-9B 9B 0.296 0.247 0.316 0.316
Gemini-3-flash API 0.428 0.184 0.263 0.316
Gemini-3.5-flash API 0.318 0.151 0.162 0.216
Qwen2.5-1.5B 1.5B 0.068 0.078 0.082 0.082

3.4pt

Table 4: Best-performing scaled-up and frontier annotators on thevalidated representative 100-item gold, scored on the broad gate(two-prompt consensus) and ranked by Cohen’s \(\kappa\). On this representativesample models reach up to \(\kappa=0.72\), well above their agreement on theboundary-enriched set: the difficulty is specific to the divisive/hate boundary,not the broad construct. Greedy decoding, byte-identical gate prompts.
Annotator Size \(\kappa\) \(P\)
Broad gate (2-prompt) 7B 0.820 0.960
OLMo-3.1-32B 32B 0.722 0.933
OLMo-3-7B 7B 0.623 0.886
Gemini-2.5-flash-lite API 0.607 0.971
Qwen2.5-32B 32B 0.605 0.923
Gemini-2.5-pro API 0.601 0.921
Gemini-2.5-flash API 0.565 0.878
Gemini-3-flash API 0.493 0.966
Gemini-3.5-flash API 0.380 0.957

5.4 Validation on operator roles↩︎

The per-operation mix (§5.2) cannot, on its own, separate operator intent from language, era, or target availability across seven heterogeneous campaigns. As an external check we re-ran the unchanged pipeline (the same two-stage gate and the same construct-typing rule) on a corpus where the operator structure is independently known: the Clemson/FiveThirtyEight release of Internet Research Agency tweets [55], in which every account carries a hand-assigned role (Right Troll, Left Troll, Fearmonger, Hashtag Gamer, News Feed, Commercial) from [10]; these released role labels have been treated as ground truth in subsequent troll-detection work [56]. We drew a role-stratified English, non-retweet candidate sample using the identical cue-and-target selection applied to the main corpus, gated it, and typed the positive items; roles were withheld from the model and joined back only afterward. This is concurrent validity (the broad gate’s output checked against an independent, pre-existing labeling of the same items), within a single actor (the IRA), not independent generalization.

Two signals are associated with the roles (Fig. 3). First, the broad gate distinguishes the troll roles from the service roles: confirmed hostility runs 5.4–16.7% across the four troll roles but falls to 1.1% for News Feed and 0.0% (0/14) for Commercial, the roles whose function was newswire syndication and advertising, not provocation. Second, among identity-hate items the roles differ markedly in whom they target: the Right Troll role concentrates on Muslims (60.8%) and immigrants (24.3%), while the Left Troll role centers ethnic/racial out-groups (47.1%), consistent with the Islamophobic/anti-immigrant and Black-identity personas documented by [10], [57].

What does not separate the roles is the identity-versus-partisan construct split itself: among confirmed items the Right Troll role is 42.2% identity / 57.4% partisan and the Left Troll role 38.4% / 58.7% (all lean partisan, as does the Hashtag Gamer role at 32.0% / 68.0%), and the dehumanizing/inciting hard core stays low throughout (5.2–8.7%). Within one operation, the roles specialize by target and by intensity (operationalized as the confirmation rate), not by construct type. This refines rather than contradicts §5.2: construct-type specialization is something we observe across operations, not a property of every division of labor inside one. (The cue-and-target candidate filter over-selects lexically hostile content, so these are mixes within a hostile-selected tail, not corpus base rates.)

Figure 3: Concurrent validation on the Clemson IRA corpus, where operator roles are independently known [10], [55], [57]; roles were withheld from the model. (a) The broad gate confirms hostility on the four troll roles (filled circles) but stays near zero on the News Feed and Commercial service roles (open squares); bars are Wilson 95% intervals, n is candidate items per role. (b) Among identity-hate items, the right-flank role targets Muslims and immigrants while the left-flank role centers ethnic/racial out-groups (hatching distinguishes roles in grayscale). The construct mix (identity vs.partisan) leans partisan for all roles (§5.4); within one operation, roles specialize by target and intensity, not by construct type. Associational.

5.5 Prevalence floors by subset↩︎

Tweet-level prevalence floors (random stratum, Jeffreys 95% lower bound, \(n=2{,}000\) per operation) decrease as the construct is narrowed, and the narrowing is associated with a wider cross-operation contrast (Fig. 4). At the dehumanizing-identity-hate core, RU-op is the only operation with a non-trivial floor (0.0158); every other operation is at or near zero, and VE-op-A’s identity floor of 0.0004 is consistent with effectively no identity hate. The broad-construct RU-op elevation (roughly an order of magnitude above the lowest operation) thus persists under the narrowed construct and becomes more distinctly separated from the other operations once the construct is isolated.

Figure 4: Prevalence floors (Jeffreys 95% lower bound) by construct subset (broad gate, identity hate, and the dehumanizing hardcore), grouped bars per operation on a log axis (hatching distinguishes subsets in grayscale). Floors fall as the construct narrows; RU-op has the highest floor at every layer, and at the dehumanizing hardcore core only RU-op clears a non-trivial floor while every other operation sits near zero.

5.6 Contingency under narrowing↩︎

Across the positive census, operations diverge most on whom they target and what narrative they advance, and converge on how hostile they are (stance intensity, dehumanization), reproducing a layered “content diverges / form converges” pattern. We measure this with cross-operation Cramér’s \(V\), an association effect size running from \(0\) (no difference) to \(1\) (complete separation): target \(0.513\), narrative \(0.409\), intensity \(0.093\). Every dimension exceeds its item-level permutation null at \(p<0.001\) (\(5{,}000\) permutations); we use this permutation test throughout in place of an analytic \(\chi^2\), whose independence assumption the multi-label dimensions violate. We compute \(V\) two ways. For dimensions where each item takes exactly one value (target group, stance intensity, dehumanization), it comes from a standard campaign\(\times\)category \(\chi^2\) over the \(5{,}457\) items. For dimensions where one item can carry several labels at once (io narrative, threat frame, emotion, moral foundations), we instead build a campaign\(\times\)label incidence table, counting each label occurrence separately, so the totals run past the item count (io narrative \(n=6{,}521\), threat frame \(5{,}871\), emotion \(8{,}052\), moral foundations \(7{,}430\)); there \(V\) measures how each campaign’s mix of labels differs and is assessed by the same item-level permutation (shuffling the campaign label while preserving each item’s full label set). When the construct is narrowed, the content-layer divergence decreases monotonically (Fig. 5): target \(0.513 \to 0.410 \to 0.260\) and narrative \(0.409 \to 0.364 \to 0.275\) across broad, identity, and hardcore subsets, while intensity stays low and shared. The layered pattern persists on the dehumanizing-identity-hate core, but part of the original contrast reflected the identity-versus-political construct gap, which we report explicitly rather than fold into the layered interpretation.

Temporal variation reinforces this layering. Hostility intensity is temporally flat (Spearman \(\rho\approx0\)), consistent with a scripted repertoire that does not escalate, whereas target composition changes over an operation’s active period: on the identity-hate subset IRA re-aims its targeting substantially (temporal Cramér’s \(V=0.633\)) and IR-op-A moderately (\(V=0.410\)), consistent with reports that IRA accounts adapted across the operation’s lifespan [56]. Full per-operation trends are in the Supplementary Information.

Figure 5: Cross-operation Cramér’s V per dimension, by construct subset (n in legend; rows sorted by the broad-construct value). Each dimension shows three points (broad construct, circle; identity subset, square; dehumanizing hardcore, triangle), with bootstrap 95% confidence intervals (800 resamples) as whiskers. Content dimensions (target group, IO narrative) diverge substantially across operations while form dimensions (emotion, dehumanization, stance intensity) stay low everywhere; the content gap narrows as the construct is restricted to the identity core, where small n widens the intervals.

5.7 Account-concentration specificity↩︎

Account-level concentration of hostile output (the Gini coefficient over authoring accounts, near \(0\) when output is spread evenly across accounts and near \(1\) when it is dominated by a few) is construct-specific (Fig. 6). Under the broad gate, RU-op (0.854), IR-op-A (0.814), and VE-op-A (0.768) all exhibit high broad-construct concentration. But when attention is restricted to identity hate, only RU-op (0.821) and IR-op-A (0.728) remain highly concentrated on identity hate; VE-op-A’s concentration falls to 0.133, indicating that under the typing rule its broad-construct concentration is attributable to partisan-invective rather than identity-hate items. So under this rule the broad-construct high-concentration characterization does not hold for VE-op-A.

Figure 6: Account concentration (Gini of hostile items per posting account), as a dumbbell from the broad construct to the identity-hate subset. RU-op and IR-op-A stay highly concentrated at both layers; for VE-op-A concentration falls (0.77\!\rightarrow\!0.13), so its broad-construct concentration reflects partisan invective rather than identity-hate items.

5.8 Affect-label convergent validity↩︎

LLM affect labels are associated with independent lexicon signal to varying degrees (one-sided Mann–Whitney \(U\), with rank-biserial correlation \(r_b\) as the effect size, \(0\) indicating no separation and \(1\) complete separation; BH-FDR). Only care/harm shows a substantial effect (\(r_b=0.520\), on a small positive class of \(n_+=111\) of \(5{,}332\)); fairness/cheating (\(r_b=0.174\)) and fear (\(r_b=0.105\)) are weak (\(q<0.001\)); anger (\(r_b=0.060\), \(q=0.01\)) and the loyalty/authority labels are statistically significant but their effect sizes are negligible, with significance attributable to large \(n\). A second, independent emotion lexicon (LIWC2015 [52]), applied to the same English text, reproduces the emotion convergence: anger tracks LIWC anger (\(r_b=0.096\), \(q=1.4\times10^{-5}\)) and fear tracks LIWC anxiety (\(r_b=0.134\), \(q<10^{-9}\)). Across the confirmed set the LIWC tone score is strongly negative (\(23.0\) on a \(0\)\(100\) scale where below \(50\) is negative; negative-emotion words \(4.1\%\) versus positive \(1.3\%\)), with anger words (\(2.3\%\)) outweighing anxiety words (\(0.8\%\)). Two labels do not converge: sanctity/degradation (\(r_b=-0.077\), \(q=1.0\)) and disgust (\(r_b=-0.092\), \(q=1.0\)). On 29 Russian-original confirmed items the rusentilex (a Russian sentiment lexicon) negative-word rate is \(0.026\) (median \(0\)); we read this small sample as suggestive, not conclusive, evidence that negativity is present pre-translation. This cross-validation bears on the moral and emotion fields only; it does not validate the target taxonomy on which the typing rule operates.

6 Discussion↩︎

6.0.0.1 Construct typing before counting.

Methodologically, the construct a detector measures warrants validation, and content warrants typing before reporting a “hate rate,” echoing critiques that abusive-language constructs are routinely conflated under one label [9]. A gate validated as a broad hostility detector is a legitimate instrument, but the share it flags is a divisiveness rate, not a hate rate. For these seven operations, treating the broad flag as “hate” over-counts hate speech by the margins quantified in §5.1 (\(2\)\(5\times\)). This is a transferable prediction: any broad-hostility detector reported as a hate rate over-attributes hate by a margin set by how much of the hostile tail is non-identity invective, and is falsifiable by validating a construct-specific detector that should show no such gap.

6.0.0.2 Defining manufactured divisiveness.

The shared product of these operations is manufactured divisiveness; identity hate is a narrower construct nested within it, dominant in some operations and near-absent in others, so that six of seven operations sort into three construct regimes the broad rate flattens, and the dehumanizing/inciting identity core is narrower still. Framing it this way connects influence-operation content to the broader account of affective and sectarian polarization as a manufactured rather than spontaneous condition [44]. For operations of this kind (already-detected, attributed campaigns), this favors reporting composition rather than a single “hate” label; we do not extrapolate beyond the scoped sample.

6.0.0.3 Re-typing rather than deflation.

That the broad flag is not a “hate rate” must not be read as “influence-operation hate is overblown.” We re-type hostile content rather than deflate its prevalence. A real, dehumanizing identity-hate core is present and, for RU-op, persists under every narrowing of the construct. The implication for platforms bears first on measurement and reporting: measure the composition and distinguish the identity-hate core, rather than deprioritizing it on the basis of the broad label’s imprecision. Because the identity boundary that automated enforcement would key on is the same ambiguous line, we frame this as a reporting recommendation, with the composition used to triage rather than to set enforcement thresholds.

6.0.0.4 Possible sources of compositional divergence.

The divergence in target and narrative against the convergence in intensity and dehumanization is consistent with audience-tuned content over a shared production form. We state this associationally and claim no causal playbook. The IRA role analysis (§5.4) fits this reading: where operator structure is independently known, roles diverge in target and positive rate while sharing a near-balanced construct mix.

7 Limitations↩︎

The typing is a rule over an LLM-derived taxonomy, not a human adjudication of whether each item is hate. The broad gate carries human validation (\(\kappa=0.82\)) and the typing rule is separately validated against an expert typing of 110 items at moderate agreement (\(\kappa=0.52\)); the residual disagreement reflects the ambiguous divisive/hate boundary, on which three expert annotators themselves agree pairwise at only \(\kappa=0.37\)\(0.50\), so items near that boundary are typed by rule, not adjudicated by humans. A second independent expert relabeled a contested-boundary subset of the typing gold and agrees with the first only at \(\kappa=0.44\), locating this residual disagreement in the boundary rather than in a single coder (§4.3); the reported split is the rule’s partition under stated definitions, and shifts under alternative boundary choices within a bounded envelope (identity share \(45.1\)\(66.9\%\), over-statement \(1.5\)\(2.2\times\); Section 5.1). A further limitation is that the decisive target-group attribute is itself LLM-derived and is not separately validated at scale, so the partition could inherit characterization-model bias on that field. The per-operation construct mix is a descriptive composition of the positive tail, not a tested discriminative classifier; we fit no null model separating operator intent from language, era, or target-availability, and the smallest operations (notably BD-op, 11 accounts) yield unstable per-operation statistics. Affect cross-validation supports only the moral and emotion fields, not the target taxonomy that drives the typing; two affect labels (sanctity, disgust) do not track the reference lexica, and affect on non-English items is measured after translation, with limited pre-translation evidence (\(n=29\)) for Russian. Prevalence figures are floors, not point estimates. Tier-N reporting (labels and counts only, no released text) means exact item-level reproduction of the typing requires access to the underlying archive; the typing rule and the characterization schema are, however, fully specified, so the partition is auditable in principle by any holder of the source data. Finally, all findings are scoped to these seven specific, already-attributed campaigns: the absence of a construct here is not evidence of its absence in influence operations generally.

8 Ethics and Data Statement↩︎

This work is a secondary analysis of a publicly released, state-attributed archive and is exempt from human-subjects review; we attribute no actor beyond the archive’s own labeling and report results associationally, scoped to these seven campaigns. Following the Tier-N protocol (Section 3), we release labels and counts, not tweet text or account identifiers; the full instrument (verbatim prompts, the typing rule as code, model identifiers, and seeds) is in the Supplementary Information and reproduces every count.

8.0.0.1 Acknowledgments.

Work partly supported by NSF HCC award 2331722. We thank Leonardo Blas, Patrick Gerard, and Daniel Ruiz (USC) for providing their data annotation expertise.

References↩︎

[1]
Massimo Stella, Emilio Ferrara, and Manlio De Domenico. Bots increase exposure to negative and inflammatory content in online social systems. Proceedings of the National Academy of Sciences, 115(49):12435–12440, 2018.
[2]
Ashfaq Ali Shafin and Khandaker Mamun Ahmed. Toxicity in state sponsored information operations. In Proceedings of the 36th ACM Conference on Hypertext and Social Media (HT ’25), 2025.
[3]
Ashfaq Ali Shafin and Khandaker Mamun Ahmed. The language of influence: Sentiment, emotion, and hate speech in state sponsored influence operations. In Proceedings of the 18th International Conference on PErvasive Technologies Related to Assistive Environments (PETRA), 2025. arXiv:2505.07212.
[4]
Thomas Davidson, Dana Warmsley, Michael Macy, and Ingmar Weber. Automated hate speech detection and the problem of offensive language. In Proceedings of the International AAAI Conference on Web and Social Media (ICWSM), pages 512–515, 2017.
[5]
Antigoni-Maria Founta, Constantinos Djouvas, Despoina Chatzakou, Ilias Leontiadis, Jeremy Blackburn, Gianluca Stringhini, Athena Vakali, Michael Sirivianos, and Nicolas Kourtellis. Large scale crowdsourcing and characterization of Twitter abusive behavior. In Proceedings of the International AAAI Conference on Web and Social Media (ICWSM), pages 491–500, 2018.
[6]
Paula Fortuna and Sérgio Nunes. A survey on automatic detection of hate speech in text. ACM Computing Surveys, 51(4):1–30, 2018.
[7]
Bertie Vidgen and Leon Derczynski. Directions in abusive language training data: A systematic review of garbage in, garbage out. PLOS ONE, 15(12):e0243300, 2020.
[8]
Zeerak Waseem, Thomas Davidson, Dana Warmsley, and Ingmar Weber. Understanding abuse: A typology of abusive language detection subtasks. In Proceedings of the First Workshop on Abusive Language Online (ACL), pages 78–84, 2017.
[9]
Michele Banko, Brendon MacKeen, and Laurie Ray. A unified taxonomy of harmful content. In Proceedings of the Fourth Workshop on Online Abuse and Harms (ACL), pages 125–137, 2020.
[10]
Darren L. Linvill and Patrick L. Warren. Troll factories: Manufacturing specialized disinformation on Twitter. Political Communication, 37(4):447–467, 2020.
[11]
Savvas Zannettou, Tristan Caulfield, Emiliano De Cristofaro, Michael Sirivianos, Gianluca Stringhini, and Jeremy Blackburn. Disinformation warfare: Understanding state-sponsored trolls on Twitter and their influence on the web. In Companion Proceedings of the World Wide Web Conference (WWW Companion), pages 218–226, 2019.
[12]
Savvas Zannettou, Tristan Caulfield, William Setzer, Michael Sirivianos, Gianluca Stringhini, and Jeremy Blackburn. Who let the trolls out? towards understanding state-sponsored trolls. In Proceedings of the 10th ACM Conference on Web Science (WebSci), pages 353–362, 2019.
[13]
Kate Starbird. Disinformation’s spread: Bots, trolls and all of us. Nature, 571(7766):449, 2019.
[14]
Adam Badawy, Emilio Ferrara, and Kristina Lerman. Analyzing the digital traces of political manipulation: The 2016 Russian interference Twitter campaign. In Proceedings of the IEEE/ACM International Conference on Advances in Social Networks Analysis and Mining (ASONAM), pages 258–265, 2018.
[15]
Deen Freelon, Michael Bossetta, Chris Wells, Josephine Lukito, Yiping Xia, and Kirsten Adams. Black trolls matter: Racial and ideological asymmetries in social media disinformation. Social Science Computer Review, 40(3):560–578, 2022.
[16]
Julie Jiang, Xiang Ren, and Emilio Ferrara. : Political leaning detection using language features and information diffusion on social networks. In Proceedings of the International AAAI Conference on Web and Social Media (ICWSM), volume 17, pages 459–469, 2023.
[17]
Mohammad Hammas Saeed, Shiza Ali, Pujan Paudel, Jeremy Blackburn, and Gianluca Stringhini. Unraveling the web of disinformation: Exploring the larger context of state-sponsored influence campaigns on Twitter. In Proceedings of the 27th International Symposium on Research in Attacks, Intrusions and Defenses (RAID), pages 353–367, 2024.
[18]
Xinyu Wang, Jiayi Li, Eesha Srivatsavaya, and Sarah Rajtmajer. Evidence of inter-state coordination amongst state-backed information operations. Scientific Reports, 13:7716, 2023.
[19]
Iacopo Pozzana and Emilio Ferrara. Measuring bot and human behavioral dynamics. Frontiers in Physics, 8:125, 2020.
[20]
Lynnette Hui Xian Ng and Kathleen M. Carley. Bots, Bias, and Influence: The Hidden Architects of Social Media. Cambridge Scholars Publishing, 2026.
[21]
Emilio Ferrara, Herbert Chang, Emily Chen, Goran Muric, and Jaimin Patel. Characterizing social media manipulation in the 2020 U.S. presidential election. First Monday, 25(11), 2020.
[22]
Serena Tardelli, Leonardo Nizzoli, Maurizio Tesconi, Mauro Conti, Preslav Nakov, Giovanni Da San Martino, and Stefano Cresci. Temporal dynamics of coordinated online behavior: Stability, archetypes, and influence. Proceedings of the National Academy of Sciences, 121(20):e2307038121, 2024.
[23]
Paula Fortuna, Juan Soler, and Leo Wanner. Toxic, hateful, offensive or abusive? what are we really classifying? an empirical analysis of hate speech datasets. In Proceedings of the Twelfth Language Resources and Evaluation Conference (LREC), pages 6786–6794, 2020.
[24]
Zeerak Waseem and Dirk Hovy. Hateful symbols or hateful people? predictive features for hate speech detection on Twitter. In Proceedings of the NAACL Student Research Workshop, pages 88–93, 2016.
[25]
Maarten Sap, Dallas Card, Saadia Gabriel, Yejin Choi, and Noah A. Smith. The risk of racial bias in hate speech detection. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics (ACL), pages 1668–1678, 2019.
[26]
Maarten Sap, Swabha Swayamdipta, Laura Vianna, Xuhui Zhou, Yejin Choi, and Noah A. Smith. Annotators with attitudes: How annotator beliefs and identities bias toxic language detection. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (NAACL-HLT), pages 5884–5906, 2022.
[27]
Mai ElSherief, Caleb Ziems, David Muchlinski, Vaishnavi Anupindi, Jordyn Seybolt, Munmun De Choudhury, and Diyi Yang. Latent hatred: A benchmark for understanding implicit hate speech. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 345–363, 2021.
[28]
Patrícia Rossini. Beyond incivility: Understanding patterns of uncivil and intolerant discourse in online political talk. Communication Research, 49(3):399–425, 2022.
[29]
Punyajoy Saha, Binny Mathew, Kiran Garimella, and Animesh Mukherjee. “short is the road that leads from fear to hate”: Fear speech in IndianWhatsApp groups. In Proceedings of the Web Conference 2021 (WWW), pages 1110–1121, 2021.
[30]
Bijean Ghafouri, Eun Cheol Choi, Priyanka Dey, and Emilio Ferrara. Position: RLHF may not reflect genuine preferences. In Proceedings of the 43rd International Conference on Machine Learning (ICML). PMLR, 2026.
[31]
Emilio Ferrara. Cultural targets, structural frames, binding morals: A cross-lingual audit of online hate in multicultural singapore, 2026. arXiv preprint arXiv:2606.21996.
[32]
Fabrizio Gilardi, Meysam Alizadeh, and Maël Kubli. outperforms crowd workers for text-annotation tasks. Proceedings of the National Academy of Sciences, 120(30):e2305016120, 2023.
[33]
Caleb Ziems, William Held, Omar Shaikh, Jiaao Chen, Zhehao Zhang, and Diyi Yang. Can large language models transform computational social science? Computational Linguistics, 50(1):237–291, 2024.
[34]
Nick Pangakis and Sam Wolken. Keeping humans in the loop: Human-centered automated annotation with generative AI. In Proceedings of the International AAAI Conference on Web and Social Media (ICWSM), volume 19, pages 1471–1492, 2025.
[35]
Deepak Kumar, Yousef Anees AbuHashem, and Zakir Durumeric. Watch your language: Investigating content moderation with large language models. In Proceedings of the International AAAI Conference on Web and Social Media (ICWSM), volume 18, pages 865–878, 2024.
[36]
Barbara Plank. The “problem” of human label variation: On ground truth in data, modeling and evaluation. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 10671–10682, 2022.
[37]
Alexandra N. Uma, Tommaso Fornaciari, Dirk Hovy, Silviu Paun, Barbara Plank, and Massimo Poesio. Learning from disagreement: A survey. Journal of Artificial Intelligence Research, 72:1385–1470, 2021.
[38]
Aida Mostafazadeh Davani, Mark Dı́az, and Vinodkumar Prabhakaran. Dealing with disagreements: Looking beyond the majority vote in subjective annotations. Transactions of the Association for Computational Linguistics (TACL), 10:92–110, 2022.
[39]
Abigail Z. Jacobs and Hanna Wallach. Measurement and fairness. In Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency (FAccT), pages 375–385, 2021.
[40]
Jesse Graham, Jonathan Haidt, and Brian A. Nosek. Liberals and conservatives rely on different sets of moral foundations. Journal of Personality and Social Psychology, 96(5):1029–1046, 2009.
[41]
William J. Brady, Julian A. Wills, John T. Jost, Joshua A. Tucker, and Jay J. Van Bavel. Emotion shapes the diffusion of moralized content in social networks. Proceedings of the National Academy of Sciences, 114(28):7313–7318, 2017.
[42]
Steve Rathje, Jay J. Van Bavel, and Sander van der Linden. Out-group animosity drives engagement on social media. Proceedings of the National Academy of Sciences, 118(26):e2024292118, 2021.
[43]
Shanto Iyengar, Yphtach Lelkes, Matthew Levendusky, Neil Malhotra, and Sean J. Westwood. The origins and consequences of affective polarization in the united states. Annual Review of Political Science, 22:129–146, 2019.
[44]
Eli J. Finkel, Christopher A. Bail, Mina Cikara, Peter H. Ditto, Shanto Iyengar, Samara Klar, Lilliana Mason, Mary C. McGrath, Brendan Nyhan, David G. Rand, Linda J. Skitka, Joshua A. Tucker, Jay J. Van Bavel, Cynthia S. Wang, and James N. Druckman. Political sectarianism in America. Science, 370(6516):533–536, 2020.
[45]
Emilio Ferrara. Five myths about influence operations: What 25 million tweets across seven state campaigns reveal. Preprints, 2026.
[46]
Ozgur Can Seckin, Manita Pote, Alexander C. Nwala, Lake Yin, Luca Luceri, Alessandro Flammini, and Filippo Menczer. Labeled datasets for research on information operations. In Proceedings of the International AAAI Conference on Web and Social Media (ICWSM), volume 19, pages 2567–2574, 2025.
[47]
Lilliana Mason. Uncivil Agreement: How Politics Became Our Identity. University of Chicago Press, Chicago, IL, 2018.
[48]
Julia Mendelsohn, Ronan Le Bras, Yejin Choi, and Maarten Sap. From dogwhistles to bullhorns: Unveiling coded rhetoric with language models. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (ACL), pages 15162–15180, 2023.
[49]
Klaus Krippendorff. Reliability in content analysis: Some common misconceptions and recommendations. Human Communication Research, 30(3):411–433, 2004.
[50]
Jeremy A. Frimer, Reihane Boghrati, Jonathan Haidt, Jesse Graham, and Morteza Dehghani. Moral foundations dictionary for linguistic analyses 2.0. Unpublished manuscript, distributed via OSF, 2019.
[51]
Saif M. Mohammad and Peter D. Turney. Crowdsourcing a word–emotion association lexicon. Computational Intelligence, 29(3):436–465, 2013.
[52]
James W. Pennebaker, Ryan L. Boyd, Kayla Jordan, and Kate Blackburn. The development and psychometric properties of LIWC2015. Technical report, University of Texas at Austin, 2015.
[53]
Natalia Loukachevitch and Anatolii Levchik. Creating a general Russian sentiment lexicon. In Proceedings of the Tenth International Conference on Language Resources and Evaluation (LREC), pages 1171–1176, 2016.
[54]
Paul Röttger, Hannah Rose Kirk, Bertie Vidgen, Giuseppe Attanasio, Federico Bianchi, and Dirk Hovy. : A test suite for identifying exaggerated safety behaviours in large language models. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (NAACL-HLT), pages 5377–5400, 2024.
[55]
FiveThirtyEight. Russian troll tweets. https://github.com/fivethirtyeight/russian-troll-tweets, 2018. Internet Research Agency tweets collected and categorized by Linvill and Warren, Clemson University.
[56]
Jane Im, Eshwar Chandrasekharan, Jackson Sargent, Paige Lighthammer, Taylor Denby, Ankit Bhargava, Libby Hemphill, David Jurgens, and Eric Gilbert. Still out there: Modeling and identifying Russian troll accounts on Twitter. In Proceedings of the 12th ACM Conference on Web Science (WebSci), pages 1–10, 2020.
[57]
Ahmer Arif, Leo Graiden Stewart, and Kate Starbird. Acting the part: Examining information operations within #BlackLivesMatter discourse. Proceedings of the ACM on Human-Computer Interaction (CSCW), 2:1–27, 2018.