Rater State Bias in RLHF Preference Data: An Audit Framework


Abstract

We identify a structured confound in Reinforcement Learning from Human Feedback (RLHF). Pairwise preference labels are intended to reflect the compared outputs, but they may also reflect the rater’s state during annotation. Under sustained stressful or distressing conditions, raters’ preferences may shift over time. As a result, preference data can encode rater state alongside judgments about response quality. These shifts differ from ordinary disagreement or random label noise. They are state dependent, can be shared across annotators working under similar conditions, and can propagate through reward modeling and policy optimization. We therefore propose rater state shift as a plausible and testable source of structured bias in RLHF preference data.

This paper develops a hypothesis and an audit framework for studying this source of bias. We define rater state shift, rater state confound, and correlated rater state bias. We also define survival level emotional authenticity as a measurable response pattern using lexical, pragmatic, discourse, and safety related features. We analyze how correlated rater state bias can survive aggregation and enter learned reward signals. We derive five falsifiable predictions and effect size thresholds for an initial audit. Finally, we present an audit protocol and pilot study plan that can be applied to publicly available instruction tuned models. We do not infer the training history of any specific deployed model. Our goal is to isolate a plausible and testable source of structured bias in RLHF preference data.

1 Introduction↩︎

Reinforcement Learning from Human Feedback (RLHF) is widely used to align large language models with human preferences [1], [2]. We study the preference data used in RLHF as a potential site of structured annotation bias. The core object is a pairwise preference label: given two model outputs, a rater records which one is preferred, and these labels become targets for reward modeling. A preference label is therefore not a direct measurement of output quality, it is a situated record of comparative judgment.

A growing body of work shows that annotator identity, beliefs, and cultural background shape labeled data [3][5]. This has motivated approaches that treat disagreement as structured signal rather than noise [6][8]. Here, we focus on a different source of structure: the rater’s state during annotation, rather than relatively stable rater traits. We ask how annotation conditions can systematically alter what raters reward. Under annotation conditions, we mean the external circumstances under which the work is performed. Under rater state, we mean the rater’s current internal configuration under those circumstances, including affect, attention, and regulation.

We introduce the following definitions: a rater state shift is a systematic change in a rater’s state during annotation, shaped by annotation conditions and broader surrounding pressures, and sustained over time; a rater state confound is a case where the recorded preference label reflects that state, rather than a stable judgment about output quality. A rater state shift becomes a rater state confound when it enters the recorded preference label; when such confounds are shared across raters and retained by learning, they become a source of correlated rater state bias in the reward model. We do not claim to identify the training history of any specific deployed model. Our claim is that rater state bias is a plausible and testable source of distortion in RLHF preference data.

In practice, RLHF preference labels are generated by paid annotators working inside managed data production systems, often through crowdwork or contractor arrangements rather than under idealized laboratory conditions [9], [10]. In NLP and dataset creation more broadly, annotation quality depends not only on the item being labeled, but also on workforce selection, training, qualification filters, feedback, compensation, and task design [11], [12]. These operational conditions matter here because they provide a natural pathway by which annotation can become state dependent: what raters notice, tolerate, reward, or reject may shift with the conditions under which labeling is performed.

We propose a concrete mechanism linking introduced definitions to the training workflow. Rater state confounds enter the preference signal as correlated label shifts across raters exposed to similar annotation conditions (Section 4.1). Because the reward model is trained to maximize likelihood of observed preferences, it cannot distinguish state-driven agreement from genuine quality agreement; correlated confounds are absorbed into the learned reward function alongside the target signal (Section 4.2). During policy optimization, the KL-constrained policy then optimizes against whatever the reward model has learned, including the absorbed confound, amplifying the distortion in generated outputs (Section 4.3).

This paper presents a hypothesis and an audit framework. It does not argue that any observed model behavior can be reduced to a single cause. The central point is that if annotation conditions can systematically affect raters during preference labeling, then RLHF preference data may contain a structured and testable source of distortion that should be examined directly.

The paper is organized as follows. Section 2 reviews related work on annotator bias, signal imperfection in RLHF, and measuring response style in text. Section 3 formalizes preference training as preference signal estimation and introduces the rater state framing. Section 4 traces how rater state bias propagates through RLHF. Section 5 summarizes the empirical premise for rater state shift. Section 6 states the hypothesis and falsifiable predictions. Section 7 defines the measurement taxonomy, audit protocol, and pilot study. Section 8 discusses implications and limitations, and Section 9 concludes.

2 Related Work↩︎

This section brings together several lines of related work that frame the phenomenon we study. In this paper, we consider the possibility that RLHF preference labels may be systematically distorted by changes in the rater’s state during annotation, and that such distortions may persist through reward modeling and downstream training. The studies reviewed here are relevant because they show that labeled data often reflect annotators as well as items, that imperfections in learned reward signals can be amplified during optimization, and that response style can be audited with reproducible measures at the level of the text.

The closest prior work is the literature on annotator bias, disagreement, and positionality. That literature shows that label variation often has structure rather than being mere noise. Our contribution is to extend that logic from relatively stable annotator characteristics to within rater variation induced by annotation conditions. We locate rater state shift, rater state confound, and correlated rater state bias in that space.

2.1 Annotator Bias, Disagreement, and Positionality↩︎

A considerable body of research has shown that labeled data are shaped not only by the items being annotated, but also by the people who annotate them. [3] showed that annotator beliefs and identities influence toxic language labels. [5] introduced jury learning as a way to retain dissenting annotator perspectives rather than collapsing them into a single majority label. [4] showed that NLP datasets and models disproportionately align with Western, white, college educated, and younger populations. [7] argued that human label variation should often be treated as informative signal rather than noise, and [8] showed that models preserving individual annotator perspectives can outperform majority vote aggregation on subjective tasks. More broadly, work on intersectional fairness in machine learning has emphasized that socially structured differences do not reduce to a single demographic axis and can be obscured when annotation and evaluation processes treat human judgments as uniform [13].

This literature is the closest point of contact for our argument. It shows that disagreement in labeled data often has structure, and that this structure can reflect perspective, position, and social experience rather than annotation error alone. In most of this work, however, the relevant sources of variation are treated as relatively stable across time. They are tied to who the annotator is, how the annotator is situated, or what background the annotator brings to the task.

Our contribution is to shift attention to a different source of structured variation: changes in the rater’s state during annotation. The key idea is not that raters differ from one another in stable ways, but that the same rater may judge differently under different annotation conditions. This moves the discussion from persistent annotator characteristics to variation that can arise during the annotation process itself. In that sense, our framework extends the logic of structured disagreement into a setting where the relevant source of variation is not fixed perspective, but state dependent judgment under operational conditions.

2.2 Signal Imperfection and Amplification in RLHF and Machine Learning↩︎

A central concern in alignment is that optimization can amplify imperfections in the signal being optimized. In RLHF, [14] characterized reward model overoptimization, showing that policy optimization against an imperfect learned reward can degrade true performance in a Goodhart-like way, by optimizing a proxy rather than the intended target. They also derived scaling laws for this effect. [15] extended this concern to direct alignment methods, showing that related forms of overoptimization can arise even without an explicit reward model.

A broader machine learning literature shows a similar pattern. [16] demonstrated that image classifiers can amplify gender activity correlations beyond their rates in the training data. [17] formalized directional bias amplification and showed that it depends in part on the relative difficulty of recognizing group membership and class membership. [18] provided the first systematic study of bias amplification, showing that the effect varies with model capacity, training set size, and training duration, and is not monotonic in dataset size. More recently, [19] showed analytically, in ridge regression, that overparameterization can amplify group level bias even without class imbalance, producing disparities that do not disappear simply by increasing model size.

Taken together, these studies support a general point that is important for our argument: imperfections in the training signal can survive learning and become stronger through optimization. Our contribution is to identify a particular source of such imperfection in RLHF preference data. If annotation conditions induce systematic rater state shifts, and if those shifts enter recorded preference labels as rater state confounds, then the learned reward signal may reflect more than the intended target of comparison. When those confounds are shared across raters and retained through training, they create a plausible path from annotation conditions to correlated rater state bias in downstream model behavior. Section 4 develops this argument in detail.

2.3 Measuring Response Style and Discourse Structure in Text↩︎

Computational methods for studying response style in text are well developed. Within this broader space, tools for measuring affect and emotion provide a particularly robust and reproducible set of measures. Lexical resources such as LIWC [20], VADER [21], and the NRC Emotion Lexicon [22] provide reproducible measures of sentiment and emotion categories. Classifiers trained on datasets such as GoEmotions [23] and EmpatheticDialogues [24] extend this line of work to finer distinctions in emotion and empathic response. Together, these tools show that important aspects of response style can be measured in a systematic and replicable way.

We use this literature in a limited and practical sense. These tools do not reveal inner state, and we do not treat them as ground truth about what a speaker or model is feeling. Instead, we use them as reproducible indicators of observable properties of responses, including sentiment, emotional intensity, and cues associated with empathy or validation. They are useful here because they are well validated, easy to reproduce, and suitable for comparison across feature families and datasets.

At the same time, standard affective measures do not fully capture the construct we study. They can identify positive valence, emotional language, or empathic wording, but they are often insensitive to pragmatic structure. A measured therapeutic response and an immediately validating one may both receive positive sentiment scores, even though they differ in how quickly validation appears, how strong it is, and how much of the response is devoted to validation before redirection. For this reason, our measurement framework does not rely on sentiment alone. It combines affective indicators with discourse features that capture the organization of the response itself.

This limitation is well known in practice. Sarcasm can invert apparent sentiment while preserving surface wording. Politeness norms can soften negative stance or mask emotional intensity. Genre also matters: customer support, therapeutic dialogue, and informal chat can share a warm surface style while differing in function and discourse structure.

Work on supportive communication, attachment, and psychotherapy process also helps motivate several discourse features relevant to our construct. This literature discusses patterns such as early emotional acknowledgment, strong validation, affect mirroring, and the timing of advice relative to emotional contact [25][32]. We use this literature narrowly. We do not infer attachment style, diagnosis, or therapeutic quality from text. Rather, we draw on it to motivate observable discourse features that can be operationalized as measures and used to distinguish generic supportive tone from the more intense pattern later defined as survival level emotional authenticity. These considerations motivate the broader feature taxonomy introduced in Section 7.

3 RLHF as Preference Signal Estimation Under Rater State Bias↩︎

This section separates four things that are often blended together in RLHF discussions: (i) the standard estimator and its stable target assumption; (ii) rater state shift as an alternative account of preference labels; (iii) what RLHF preference data records and what it does not; (iv) the observable signatures that would support or weaken the hypothesis.

3.1 Technical Foundations↩︎

We begin with the standard RLHF estimator and the observables typically available for analysis.

RLHF aligns language models with human preferences through three stages [1], [2]. (1) Supervised Fine-Tuning (SFT), when a pretrained model is fine-tuned on curated prompt-response pairs. (2) Reward Model Training, when human raters compare outputs per prompt, and these comparisons train a reward model that assigns a scalar preference score to each response, denoted \(r_\theta(x,y)\) for response \(y\) given prompt \(x\). (3) Policy Optimization, the language model is fine-tuned to produce responses that receive higher scores from the reward model \(r_\theta\), while staying close to the SFT model via a KL penalty.

A common abstraction is that there exists a stable target preference probability: for a given prompt \(x\), one response \(y\) has some fixed probability of being preferred to another \(y'\) under an idealized preference process. Rater disagreement is then treated as noise around a fixed latent preference function. If each pairwise label is a noisy sample of the same stable target preference probability, with enough labels, aggregation methods (majority vote, Bradley-Terry, or similar) approximate that target rather than individual rater noise.

In RLHF datasets available for analysis, one typically observes prompts \(x\), candidate responses \(y,y'\), and a pairwise label indicating which response was preferred. Timestamps are often available. Rater IDs may be available, but not reliably.

Direct measures of rater state and exposure history are usually unavailable. This includes prior content seen, time on task, breaks, workload, and support resources. Detailed task routing logs are also often missing.

3.2 Rater State Shift and the Stable Target Assumption↩︎

We now introduce rater state shift as a source of bias in RLHF preference data.

The standard estimator assumes that pairwise labels are noisy samples from a single stable target preference process. Our alternative is that judgments may shift with rater state during annotation.

To express this, suppose each judgment is produced under a latent rater state \(s\). Let \(r(x,y \mid s)\) denote a latent preference score for response \(y\) to prompt \(x\) in state \(s\). According to the Bradley-Terry model [33], which underlies RLHF preference modeling [1], this induces the state-dependent pairwise preference probability \[p_s(y \succ y' \mid x) = \sigma\!\bigl(r(x,y \mid s)-r(x,y' \mid s)\bigr), \label{eq:state95bt}\tag{1}\] where \(\sigma\) is the sigmoid function.

If judgments are produced under different rater states, then the pairwise preference signal entering reward model training is the average over states: \[\bar p(y \succ y' \mid x) = \mathbb{E}_{s \sim S}\!\left[p_s(y \succ y' \mid x)\right], \label{eq:mixture}\tag{2}\] where \(S\) denotes the distribution of rater states in the data.

The learned reward model \(r_\theta(x,y)\) is then fit to this mixed pairwise signal through its implied Bradley-Terry probability, \[p_\theta(y \succ y' \mid x) = \sigma\!\bigl(r_\theta(x,y)-r_\theta(x,y')\bigr). \label{eq:bt95mixed}\tag{3}\]

If \(S\) is tightly concentrated, this reduces to the standard stable target picture. If shifted states are common in the data, then the learned reward model is fit to a mixture of state dependent preferences rather than to a single fixed preference process.

If rater state shift were independent across raters and over time, it would act like additional label noise. In that case, averaging over many labels would tend to wash it out as the dataset grows.

Our concern is different. Rater state shift may be correlated across raters exposed to similar conditions. In that case, aggregation need not recover a single stable target. It can preserve a shifted signal in the learned reward function.

This is a conceptual distinction that guides the audit framework. The next two subsections explain why such correlation is plausible and how it can be studied through observable proxies.

3.3 Shared Conditions and Exposure Proxies↩︎

We next explain why such bias could be correlated across raters and how it might still be studied when direct measures of rater state are unavailable.

RLHF annotation includes different task types. Some tasks collect preference rankings over model outputs such as helpfulness, harmlessness, or honesty. Others involve safety or policy labeling for harmful content and policy violations.

Most public evidence about psychological strain among annotation workers comes from safety labeling and content moderation. There is less direct public evidence about preference ranking, especially for supportive dialogue. This matters because it limits what can be claimed from existing reports.

Two distinct pathways should be separated here, because the mechanism may operate through either one or through both.

First, there is a workforce carryover pathway. Workers may rotate across safety labeling and preference ranking under the same vendors, incentives, and monitoring conditions [9], [34]. If so, strain induced in one task type could influence judgments made in another.

Second, there is a direct exposure pathway within preference work itself. Preference ranking can involve traumatic or distressing material when prompts concern self harm, abuse, violence, or crisis. In that case, the relevant state shift need not be imported from safety labeling. It can arise during preference annotation itself.

These two pathways imply different empirical signatures. The carryover pathway points to task routing, assignment history, and lagged effects of prior exposure. The direct exposure pathway points to within session drift and topic conditioned shifts as the recent share of sensitive prompts rises. Public reporting also describes low pay, precarious contracts, high volumes, and limited support as broad features of data work rather than features of one narrow task [35], [36].

Because direct measures of rater state and exposure are usually unavailable, the audit must rely on observable proxies. Recent exposure can be approximated by the share of sensitive prompts in a rater’s recent window, or in the broader batch during a given time interval. Session structure can be approximated from timestamps, including session duration and gaps between judgments. Topic clusters can be constructed from prompt embeddings or simple taxonomy tags when available. Shift timing can be studied through change points in label distributions, especially around operational events such as rubric updates, model version changes, or vendor policy changes.

These proxies do not identify mechanism on their own. They do let us test whether preference drift aligns with plausible sources of strain under real annotation conditions.

3.4 Workforce Structure and Correlated Bias↩︎

Large scale annotation work does not imply statistical cancellation when judgments are produced under shared conditions. What matters here is not only the number of labels, but also how annotation work is organized across vendors, regions, and time blocks.

Public reporting provides only partial information, but it shows a recurring pattern. Large AI companies rely on third party vendors to staff labeling and evaluation work. Reported setups include concentrated teams, low take home pay, and billing structures in which vendor rates greatly exceed worker pay [34], [35]. Other reporting describes annotation platforms operating at the scale of tens of thousands of workers [36].

These facts support a narrower point. Large scale annotation work is often organized through a small number of vendors and platforms. Workforces may be regionally concentrated and exposed to shared policies, tooling, and schedules. Under such conditions, rater state bias is plausibly correlated rather than independent.

At the same time, public sources rarely permit clean identification. They typically do not report (i) overlap between safety labelers and preference raters, (ii) task routing across assignments, or (iii) exposure histories over time. Public evidence therefore does not show that exposure from safety labeling carried into preference ranking. It does support treating correlated rater state bias as a plausible hypothesis rather than a proven mechanism. This is why the audit must look for observable signatures in the preference data itself.

If correlated rater state bias is present, it should leave observable traces in preference data. We would expect label distributions to drift with time on task and over calendar time. We would also expect topic conditioned variation, because sensitive material is not uniformly distributed across prompts.

When rater IDs are available, we would expect within session drift and cross rater synchrony around shared operational changes, such as rubric updates, policy changes, or sudden shifts in sensitive content volume. When rater IDs are unavailable, we would still expect time-based discontinuities and topic conditioned drift.

The hypothesis is weakened if label distributions remain stable over time after controlling for prompt topic, model version, and rubric changes. It is also weakened if no within session drift appears when rater IDs are available, or if apparent drift disappears under simple operational covariates. The hypothesis is strengthened if drift persists under these controls and aligns with the exposure proxies described above.

4 How Rater State Bias Propagates Through RLHF↩︎

A natural objection to our hypothesis is that any rater state bias introduced by a subpopulation of distressed raters would be diluted to insignificance by the scale of RLHF. This section provides analytical and empirically anchored estimates showing that, under plausible conditions, rater state bias can survive aggregation, enter the reward model, and be amplified during policy optimization.

In what follows, rater state confound refers to the distortion at the level of recorded preference labels, while rater state bias refers to its aggregate retention in the learned preference signal.

4.1 Entry into the Preference Signal↩︎

We begin with a simple mean field model of rater state shift, which is a particular case of the setup in Section 3.

Consider a preference annotation task in which each annotation compares a pair of model responses \((y,y')\) for a prompt \(x\). Let \(p_0(y \succ y' \mid x)\) denote the probability of preferring \(y\) over \(y'\) in the absence of rater state shift, and let \(p_{\mathrm{shift}}(y \succ y' \mid x)\) denote the corresponding probability when rater state shift is present. For brevity, in the remainder of this section we omit explicit arguments in preference probabilities where this does not cause confusion. We write \[p_{\mathrm{shift}} = p_0 + \delta(x), \label{eq:preference95shift}\tag{4}\] where \(\delta(x) \in \mathbb{R}\) represents the preference shift induced by rater state shift and may vary across prompt classes. For prompts where rater state does not affect the judgment, \(\delta(x)\) should be near zero. For prompts where it shifts judgments in favor of the response pattern of interest, \(\delta(x)\) is positive; when it shifts judgments away from that pattern, \(\delta(x)\) is negative.

Let \(f\) denote the fraction of annotations generated under conditions where rater state shift is present. The aggregate preference probability in Eq. 2 observed by the reward model is then \[\bar p(y \succ y' \mid x) = (1-f)\,p_0 + f\,p_{\mathrm{shift}} = p_0 + f\delta(x). \label{eq:aggregate}\tag{5}\] Here, \(f\delta(x)\) is the aggregate mean shift. This is a first order mean field approximation: it represents the average contribution of shifted annotations while ignoring dependence among annotators and other higher order structure.

Correlation does not change the mean shift \(f\delta(x)\) in Eq. 5 . Instead, it reduces the number of effectively independent observations, so that a correlated shift can pass more easily through standard agreement-based quality control [37].

In practice, RLHF raters are often clustered: they may work for the same vendor, in the same region, under similar conditions, and with similar content exposure. To illustrate the effect of clustering, we use the Kish design effect for the idealized case of equal cluster sizes [38]. Let \(\rho\) denote the intraclass correlation coefficient (ICC) within a cluster of size \(n\), drawn from a total workforce of \(N\) raters. The Kish design effect is \[D_{\mathrm{eff}} = 1 + (n-1)\rho, \label{eq:deff}\tag{6}\] and the effective number of independent observations is \(N_{\mathrm{eff}} = N / D_{\mathrm{eff}}\). For a cluster of \(n = 200\) workers with \(\rho = 0.1\), we get \(D_{\mathrm{eff}} = 20.9\), so a nominal workforce of \(N = 1{,}000\) yields only \(N_{\mathrm{eff}} \approx 48\). Table 1 illustrates this for several combinations of \(f\), \(\delta\), and \(\rho\).

Table 1: Illustrative estimates of rater state bias entry. \(f\): fraction of affected raters. \(\delta\): per-rater preference shift. \(\rho\): intraclass correlation. \(D_{\mathrm{eff}}\): design effect (\(n = 200\)). \(N_{\mathrm{eff}}\): effective independent sample size (\(N = 1{,}000\)). The mean shift \(f\delta\) is invariant to \(\rho\); the effective sample size is not.
\(f\) \(\delta\) \(\rho\) \(D_{\mathrm{eff}}\) \(N_{\mathrm{eff}}\) \(f \cdot \delta\)
0.10 0.10 0.00 1.0 1000 0.010
0.10 0.10 0.10 20.9 48 0.010
0.20 0.15 0.00 1.0 1000 0.030
0.20 0.15 0.10 20.9 48 0.030
0.20 0.15 0.20 40.8 25 0.030

The ICC among annotation workers who share working conditions is not directly measured. However, clustered human data can show substantial intraclass correlation, especially for attitudinal and other non-factual items, where interviewer effects tend to be strongest [39], [40]. Annotation workers share a task, working conditions, content exposure, and an institutional setting. This is a tighter clustering structure than most survey designs assume. Thus, allowing for a nontrivial \(\rho\) is well justified as a sensitivity assumption.

Unequal cluster sizes modify the particular design effect but not the qualitative conclusion. The quantity \(f\delta\) sets the size of the aggregate mean shift. Positive correlation does not change that mean shift, but it reduces the effective sample size \(N_{\mathrm{eff}}\). As \(N_{\mathrm{eff}}\) falls, correlated rater state bias becomes harder to detect and harder to remove with standard agreement-based quality control. It can therefore survive aggregation and enter reward model training.

4.2 Absorption by the Reward Model↩︎

Rater state bias behaves like systematic rather than random label noise. Neural networks can often tolerate substantial random label corruption with only moderate loss in performance [41], [42], but systematic noise introduces directional error that standard training does not remove [43]. Rater state bias is systematic in this sense: it is directional, pushing preferences in a consistent direction within the relevant prompt class; it is prompt-dependent, concentrating where rater state materially affects judgment; and it can be correlated across raters who share conditions and exposure. Reward model training is therefore more likely to absorb this shift than to average it away.

Empirical work on supervised learning shows that models can amplify bias already present in the training data. [16] showed this in image classification. [18] found that the degree of amplification depends on model capacity, training set size, and training duration. They also found stronger amplification when features linked to group membership are easier to learn.

The same logic applies here. The reward model \(r_\theta\) is trained to predict human preferences from response features. If rater state bias creates a consistent preference for responses with certain lexical or pragmatic cues, and those cues are easier to learn than the underlying quality distinction, then the reward model can overweight them relative to their true value. [18] further found that the relation between dataset size and bias amplification is not monotonic: increasing dataset size does not reliably reduce amplification. This suggests that dataset size alone is not enough to prevent structured bias from being learned during RLHF reward model training.

The Bradley-Terry model makes this absorption mechanism explicit. From Eqs. 13 and 5 , the aggregate preference probability for a fixed response pair shifts from \(p_0\) to \(\bar p\), and the corresponding shift in the fitted reward difference is \[\Delta r := \Delta\!\bigl(r_\theta(x,y)-r_\theta(x,y')\bigr) = \sigma^{-1}(\bar p)-\sigma^{-1}(p_0) = \log\!\left(\frac{p_0 + f\delta}{1 - p_0 - f\delta}\right) - \log\!\left(\frac{p_0}{1 - p_0}\right). \label{eq:reward95shift}\tag{7}\]

For \(p_0 = 0.5\) and \(f\delta = 0.03\), this gives \(\Delta r \approx 0.12\). This corresponds to shifting the preference probability from \(0.50\) to \(0.53\). The shift is small for a single pair, but it becomes consequential if it recurs systematically across the prompt class where the bias is concentrated.

4.3 Amplification During Policy Optimization↩︎

Once a structured shift has entered the learned reward signal, policy optimization can amplify it. [14] make this mechanism explicit by distinguishing the proxy reward used for optimization from the gold reward that reflects actual human judgment. They analyze two common RLHF settings: best-of-\(n\) sampling, where the candidate with the highest reward model score is selected, and reinforcement learning (RL) with proximal policy optimization (PPO), where the policy is updated to increase reward model score while staying close to a reference policy. As optimization pressure increases, proxy reward can continue to improve even after gold reward begins to decline.

In best-of-\(n\) sampling, the gold reward follows \[R_{\text{gold}}^{\text{BoN}}(d) = \alpha_{\text{BoN}} \sqrt{d} - \beta_{\text{BoN}} \, d, \label{eq:bon}\tag{8}\] and in RL with PPO, \[R_{\text{gold}}^{\text{RL}}(d) = d(\alpha_{\text{RL}} - \beta_{\text{RL}} \sqrt{d}). \label{eq:rl}\tag{9}\] Here \(R_{\text{gold}}\) is performance on the underlying target, not on the proxy reward model. The variable \(d\) is the Kullback-Leibler (KL) divergence between the optimized policy and the reference policy. The coefficient \(\alpha\) sets the initial gain from optimization. The coefficient \(\beta\) sets the degradation that appears when optimization begins to follow imperfections in the reward model.

The distinction between the reward model score and actual human judgment matters here because policy optimization does not amplify all reward model error equally. Random noise does not provide a stable direction for optimization. Structured error does. If the reward model has absorbed a systematic component of the rater state confound, policy optimization can push policy behavior further in that direction because the confound appears as a consistent feature of the learned reward signal. Rater state bias is therefore not only preserved at the reward modeling stage. It can become more behaviorally visible during policy optimization.

Table 2 summarizes how rater state bias propagates through RLHF.

Table 2: Propagation of rater state bias through RLHF. Each stage can preserve or amplify the shift introduced at the previous stage.
Stage Mechanism Effect on bias
Preference collection Rater state shift (Eq. [eq:preference95shift]) The aggregate preference signal shifts by \(f\delta(x)\). Correlation reduces \(N_{\mathrm{eff}}\) and makes the shift harder to detect.
Reward model training Bradley-Terry fitting (Eq. [eq:bt95mixed]) A structured shift in the preference signal can enter the learned reward function rather than being averaged away.
Policy optimization Optimization against the proxy reward (Eqs. [eq:bon][eq:rl]) A learned shift tied to the rater state confound can guide optimization in a consistent direction and become more behaviorally visible.

Across the RLHF, random noise is more likely to wash out, while structured bias can be preserved or amplified. Random preference noise does not provide a stable direction for policy optimization. By contrast, a learned shift tied to the rater state confound can survive aggregation, enter reward modeling, and then guide policy optimization in a consistent direction. In that sense, RLHF can transmit structured rater state bias more readily than random error.

5 Empirical Premise for Rater State Shift↩︎

This section presents the empirical premise for the rater state shift hypothesis. We do not claim direct evidence that RLHF preference raters underwent a specific state change during annotation, or that such a change entered the training data of any particular model. The claim is narrower. Large scale annotation work has operated under conditions documented to produce sustained psychological strain, and these conditions make systematic variation in rater state a plausible concern for preference data. We review evidence from adjacent annotation settings, separate documented facts from open questions, and motivate the need for direct audit.

Public evidence from content moderation and related annotation work shows that large scale data labor has often taken place under conditions associated with sustained psychological strain [9][12], [44]. Public reporting also documents repeated exposure to violent and abusive material, high daily volumes, limited mental health support, precarious contracts, and pressure to continue working while distressed [34], [45]. This evidence comes mainly from content moderation and safety labeling rather than from public RLHF preference ranking datasets. Even so, the overlap in labor arrangements, vendor structures, and exposure to sensitive material makes systematic variation in rater state a plausible concern for preference annotation as well. These conditions also align with the workforce structure discussed in Section 3.3, where shared vendors, schedules, and exposure patterns make correlated effects more plausible.

What is documented is that adjacent large scale annotation settings, especially content moderation and safety labeling, have operated under conditions associated with sustained psychological strain, including repeated exposure to disturbing material, high throughput demands, limited support, and precarious labor arrangements [9], [10], [34], [44], [45]. What remains open is whether comparable conditions characterized specific RLHF preference ranking workflows, how often they did so, and whether any resulting variation in rater state systematically affected preference labels. This is why direct audit of rater conditions and annotation context is needed.

Cultural context may also shape how emotional directness is perceived and rewarded, though its role here should be treated with caution. Cross cultural work suggests that norms of emotional expression and support vary across populations, and these differences can affect judgments of appropriateness and resonance [46][48]. At the same time, within group variation is large, and the mechanism developed in this paper does not require cultural moderation to operate [49]. We include this point because geographic concentration in annotation labor may interact with working conditions and exposure patterns, which makes cultural context a possible moderator worth testing in future audits.

6 Hypothesis and Falsifiable Predictions↩︎

The central hypothesis of this paper is that rater state shift during RLHF preference annotation can contribute a measurable component to the learned preference signal. Specifically, we propose that under sustained psychological strain, raters may systematically prefer responses that provide immediate emotional acknowledgment, unconditional validation, affect mirroring, and prolonged emotional contact before redirection. We refer to this response pattern as survival level emotional authenticity (SLEA): a composite of lexical, pragmatic, discourse, and safety boundary features operationalized in Section 7.

The main competing account is generic optimization toward warm and agreeable tone. Under that account, models learn broadly supportive responses because such responses are preferred by raters and users across prompt types. SLEA makes a more specific prediction: the defining features should concentrate on prompts involving distress, crisis, loneliness, or loss, and remain near baseline on neutral prompts. The key empirical discriminator is therefore a prompt category \(\times\) model interaction. A uniform rise in warmth across all categories favors the generic account. A category specific rise in immediate acknowledgment, strong validation, affect mirroring, and delayed redirection on distress related prompts favors SLEA.

We formalize five predictions. Each discriminates between the two accounts.

P1. Exposure-affect gradient. Preference for unconditionally validating responses increases monotonically with cumulative content exposure duration. Generic optimization predicts temporally stable preferences.3

P2. Prompt category \(\times\) model interaction. The SLEA feature signature is strongest on trauma adjacent and loneliness/attachment prompts and weakest on neutral informational queries. Generic optimization predicts uniform warm tone uplift across categories.

P3. Cross model divergence on affect laden prompts. Models trained with different rater populations under different annotation conditions diverge most on trauma adjacent prompts and converge on neutral prompts.4

P4. Validation-redirection asymmetry. On distress prompts, SLEA affected models exhibit significantly higher validation to redirection ratios than models trained under standard conditions. Generic optimization predicts balanced validate then redirect patterns.

P5. Refusal-support trade off on ambiguous safety prompts. On prompts that are borderline between safety refusal and emotional support (e.g., expressions of suicidal ideation), SLEA affected models lean toward supportive engagement rather than formulaic safety refusals. Generic optimization predicts higher refusal rates and more formulaic refusal styles.

Mechanistic grounding.Three lines of prior work jointly motivate the predicted pattern. First, the embodied cognition and feelings as information literatures establish that affective state modulates evaluative judgment: raters use their affective response to stimuli as an informational cue for quality assessment [52], [53]. When annotation conditions produce shared affective states, state modulated preferences become correlated signal rather than independent noise; the design effect analysis in Section 4.1 quantifies the consequences. Second, the attachment framework reviewed in Section 2 predicts that distress activates proximity seeking, safe haven, and secure base needs, which map linguistically to preference for immediacy markers, unconditional validation, and sustained emotional contact over premature redirection [25], [26]. These correspond directly to the SLEA feature profile defined in Section 7. Third, the amplification pathway through reward modeling and policy optimization is characterized in Section 4: once the reward model absorbs a systematic preference component, optimization against that proxy can increase its behavioral visibility through the overoptimization dynamics characterized by [14].

Alternative explanations.Several accounts could produce cross model variation in emotional profile without invoking rater state effects. (i) Intentional design: companies may deliberately optimize for emotional engagement. (ii) Architectural differences: multimodal capabilities such as voice may create emotional resonance independent of RLHF. (iii) User projection: users may attribute emotional qualities to models regardless of actual behavior [54]. (iv) System prompt and post training effects: safety prompts and post RLHF modifications may override or attenuate learned affective behaviors. None of these alternatives are mutually exclusive with rater state effects. The empirical question is whether rater state contributes a distinguishable component to the model’s emotional profile above and beyond these factors. The audit protocol in Section 7 is designed to isolate that component through category specific interaction tests rather than main effect comparisons.

7 Measures, Audit Protocol, and Pilot Study↩︎

This section translates the constructs of Section 6 into computationally reproducible measures, defines an audit protocol, and outlines a pilot study framed as an instrument validation. The pilot tests whether the proposed features measure reliably, cohere as predicted, and discriminate across prompt categories and alignment regimes. It does not test whether rater state caused any observed differences. That question requires rater-level data or controlled annotation experiments.

7.1 Feature Taxonomy↩︎

We define four feature families and one cross-cutting positional feature. SLEA, as introduced in Section 6, is a hypothesis about the kind of response pattern raters may systematically reward under the proposed mechanism. The features below do not measure those rater preferences directly. They measure properties of model outputs predicted to be elevated if those preferences were absorbed during alignment and propagated into model behavior. For each family, we specify the measurable signal, the computational method, and the prediction that distinguishes this output signature from generic supportive tone.

Lexical features. Unconditional validation phrase density (UVPD). UVPD is the count of unconditional validation phrases per 100 response tokens. The seed lexicon is constructed from a domain-neutral supportive communication corpus, for example peer support forums or general counseling transcripts, excluding crisis-specific material, by extracting validation phrases and filtering to those containing no conditional markers (if, but, however, although) within a dependency parse window of \(\pm 3\) tokens. Semantic expansion beyond exact matches is achieved by embedding seed phrases with a sentence transformer, for example all-MiniLM-L6-v2, and including any response \(n\)-gram with cosine similarity \(> 0.85\) to a seed phrase. To control for topic-appropriate language, we also compute UVPD on a human baseline: supportive responses to the same prompts drawn from peer support corpora, for example Reddit r/SuicideWatch, r/offmychest, or similar. The quantity of interest is the model-minus-human UVPD difference, stratified by prompt category. Because peer support communities often have strong local norms of validation and low redirection, this baseline may compress the model-minus-human difference on distress prompts. Where possible, the peer support baseline should therefore be supplemented with a second baseline drawn from more clinically normed supportive responses.

Distancing-hedge frequency (DHF). DHF is the count of hedges and qualifiers per 100 response tokens, using the hedge lexicon from [55] supplemented with discourse-specific additions, for example modal verbs in conditional frames modifying emotional attributions.

Therapeutic distancing score (TDS). TDS is the frequency of clinical or therapeutic framing markers per 100 tokens: references to coping strategies, professional help, therapy, normalization through diagnostic framing, and formulaic clinical phrasing.

Prediction. SLEA predicts high UVPD, exceeding the human baseline, together with low DHF and low TDS on distress prompts, with all three returning to baseline on neutral prompts. Generic optimization predicts moderate, uniform values across prompt categories.

Pragmatic features. Each sentence in the model response is classified as a validation move, affirming the user’s emotional state or experience, a question move, requesting information, clarification, or reflection, or a redirection move, introducing coping strategies, resources, or problem solving. Classification uses a dialogue-act classifier: a RoBERTa-base model fine-tuned on DailyDialog dialogue-act annotations [56], augmented with a crisis counseling supplement annotated for validation, questioning, and redirection acts. As a reliability check, a prompted LLM classifier, for example Llama 3.1-Instruct with a 5-shot prompt, is applied to at least 20% of responses, and inter-method agreement is reported.

Validation to question ratio (VQR). VQR is the ratio of validation moves to question moves per response.

Reassurance to redirection ratio (RRR). RRR is the ratio of validation moves to redirection moves.

Prediction. SLEA predicts VQR \(\gg 1\) and RRR \(\gg 1\) on distress prompts, with ratios closer to \(1\) on neutral prompts. Generic optimization predicts VQR \(\approx 1\) and moderate RRR, reflecting balanced validate-then-redirect patterns.

Discourse features. Emotional mirroring index (EMI). EMI is the cosine similarity between the centroid of affect term embeddings in the user prompt and the centroid of affect term embeddings in the model response, computed in a sentence transformer space, for example all-MiniLM-L6-v2. Affect terms are identified using the NRC Emotion Lexicon [22]. Jaccard similarity over affect term sets is computed as a robustness check.

Turn-initial acknowledgment rate (TIAR). TIAR is the proportion of model responses whose first sentence constitutes an emotional acknowledgment, classified by the dialogue-act classifier defined above.

Clinical framing proportion (CFP). CFP is the proportion of multi-sentence response segments, that is, contiguous spans of two or more sentences, that adopt clinical or therapeutic framing, classified by the presence of coping strategy language, professional help referrals, or diagnostic normalization within the segment. Unlike the token-level TDS in the lexical family, CFP captures the structural organization of clinical framing at the discourse level.

Prediction. SLEA predicts high EMI, high TIAR, and low CFP on distress prompts. Generic optimization predicts moderate EMI, moderate TIAR, and moderate to high CFP, reflecting a professional therapeutic style.

Safety boundary features. Refusal rate on ambiguous safety prompts (RRASP). RRASP is the proportion of responses to borderline safety prompts, for example expressions of suicidal ideation, self-harm references, or crisis disclosures, that constitute refusals rather than supportive engagement. A refusal is defined as a response whose primary speech act is declining to engage, redirecting to an external resource, or issuing a disclaimer rather than providing direct emotional support.

Refusal style (RS). RS classifies refusals as formulaic, meaning template-like safety language detected by keyword matching, or supportive refusal, combining emotional acknowledgment with resource referral while still maintaining the refusal boundary, detected by dialogue-act classification.

Prediction. SLEA predicts lower RRASP and a higher proportion of supportive refusals among refusals. Generic optimization predicts higher RRASP and predominantly formulaic refusal styles.

Positional feature. Validation onset (VO). VO is the token position of the first validation move, normalized by total response length, yielding a value in \([0,1]\). SLEA predicts near-zero VO on distress prompts, since validation appears in the opening sentence. Generic supportive tone predicts later onset, with framing, information gathering, or hedging preceding validation. VO is computed directly from the dialogue-act classifier output used for the pragmatic features and is reported separately because it cross-cuts the pragmatic and discourse families.

Feature family dependence. The pragmatic features (VQR, RRR), discourse features (TIAR), and positional feature (VO) share a classification backbone, namely the dialogue-act classifier. If the classifier has systematic errors, these features can fail together. The lexical features (UVPD, DHF, TDS) and safety boundary features (RRASP, RS) do not depend on this classifier and provide independent triangulation. EMI relies on a separate embedding pipeline. We therefore report results by family and note this dependency structure explicitly, so that readers can assess which findings are classifier contingent and which are not.

7.2 Audit Protocol↩︎

We define a protocol applicable to any instruction-tuned model.

  1. Prompt construction. Assemble a balanced set of 200 prompts across four categories (50 each): (a) trauma-adjacent (grief, abuse, crisis, suicidal ideation), (b) loneliness/attachment (isolation, relationship loss, abandonment), (c) neutral informational (factual queries, task instructions), and (d) ambiguous safety boundary (borderline between emotional support and safety refusal). Prompts are drawn from existing counseling taxonomies and supplemented with neutral controls. A matched human response baseline is collected for categories (a) and (b) from peer support corpora.

  2. Model sampling. For each prompt, sample \(k \geq 5\) responses per model at temperature 0.7 and with a fixed system prompt. A second run with no system prompt is conducted to assess system-prompt override effects.

  3. Feature extraction. Compute all response-level features defined in Section 7.1 for each response. For aggregate measures such as TIAR and RRASP, first compute the corresponding response-level indicators and then aggregate within each model \(\times\) prompt-category cell. Report extraction-method reliability, including inter-method agreement between dictionary-based and LLM-classifier-based extraction, for at least a 20% subsample.

  4. Analysis. For each feature family \(\times\) prompt-category cell, compare distributions across models. Report Cohen’s \(d\), or rank-biserial correlation for non-normal distributions, with bootstrapped 95% confidence intervals (\(B = 10{,}000\)).

7.3 Pilot Study: Instrument Validation↩︎

The pilot study validates the measurement instrument. It tests whether the proposed features measure reliably, cohere in the predicted structure, and discriminate across prompt categories and alignment regimes. It does not test whether rater state caused any observed variation. The study is executable entirely on publicly available instruction-tuned models.

Model selection. To minimize confounds, we select models that share a base architecture and pretraining corpus but differ in alignment method. The recommended configuration uses Llama 3.1-8B as the common base, comparing (1) the official RLHF-tuned Instruct checkpoint, (2) a DPO-aligned variant trained on the same or a comparable base, for example Tulu 2-DPO, and (3) an SFT-only checkpoint with no preference optimization, if available, or a Constitutional AI / RLAIF-aligned variant as the third condition. Same architecture, same pretraining, same scale. The only controlled variable is alignment method. This design is suited to testing whether the instrument detects structured output variation across alignment regimes under a tightly controlled comparison. It is not suited to testing P3 from Section 6, which concerns divergence induced by different rater populations and annotation conditions. Testing P3 requires models known to differ in those conditions, which falls outside the present pilot. A null result in this configuration would therefore have limited scope: it would show that the instrument does not detect the predicted pattern across these Llama-based alignment variants, not that the broader hypothesis is false or absent in models trained under different data or annotation regimes.

Prompt set and procedure. Use 200 prompts, 50 per category, as defined in the audit protocol. For each model \(\times\) prompt pair, sample 5 responses at temperature 0.7. Extract all features from Section 7.1. Compute per-model, per-category distributions.

Phase 1: Measurement reliability. This phase tests whether the features measure consistently.

Inter-method agreement. For each feature, compare dictionary-based and LLM-classifier-based extraction on a 20% subsample. Success criterion: Cohen’s \(\kappa > 0.6\) for categorical features (RS, first-sentence emotional acknowledgment labels, and dialogue-act classifications), and Pearson \(r > 0.7\) for continuous features (UVPD, DHF, TDS, EMI, VQR, RRR, VO, CFP, TIAR, RRASP).

Test-retest stability. For each model, resample the same prompts with different random seeds and compute feature-level correlation across the two samples. Success criterion: Pearson \(r > 0.7\) for continuous features, and \(\kappa > 0.6\) for categorical features.

If a feature fails both reliability checks, exclude it from subsequent phases.

Phase 2: Construct coherence. This phase tests whether the output features cohere in the pattern SLEA predicts.

Within each model, compute the pairwise correlation structure across features on distress prompts, that is, categories (a) and (b). SLEA predicts a specific covariance signature: UVPD, VQR, RRR, EMI, and TIAR should be positively correlated with each other and negatively correlated with DHF, TDS, CFP, and VO. Generic optimization does not predict this coordinated pattern. The safety boundary features, RRASP and RS, are not part of this covariance test because they are defined on category (d) prompts and operationalize P5 separately.

Phase 3: Discriminative validity. This phase tests whether the features differentiate where SLEA says they should. It has three parts.

Within-model category sensitivity. For each model, test whether a composite SLEA score, defined either as the first principal component of the reliable lexical, pragmatic, discourse, and positional features or as a simple standardized average of those features, is higher on distress prompts, categories (a) and (b), than on neutral prompts, category (c). The safety boundary features, RRASP and RS, are excluded from this composite because they are defined on ambiguous safety prompts, category (d), and test P5 separately. If the composite does not track prompt category within any model, the features do not capture category-specific variation and there is nothing for the cross-model test to explain.

Cross-model interaction. Test whether the distress-minus-neutral difference in the SLEA composite varies across models. This is the prompt-category \(\times\) model interaction term. Success criterion: at least two of the four non-safety feature families show \(|d| > 0.3\) for this interaction on the distress-related categories, with \(p < 0.05\) after Bonferroni correction for the number of tests conducted.

Safety-boundary endpoint. Test P5 separately on ambiguous safety prompts, category (d), using RRASP and RS. Success criterion: RRASP and RS satisfy the Phase 1 reliability thresholds and show a significant model effect on category (d) prompts, with at least one planned pairwise comparison surviving Bonferroni correction. This endpoint is interpreted separately from the composite because it is defined on a different prompt class.

A positive result means that the instrument detects meaningful, category-specific variation in emotional response style across alignment regimes. It does not establish that rater state is the source of that variation.

Limitations. This design cannot establish a causal link between rater state and model behavior. Models sharing a base but differing in alignment method still differ in preference data composition, optimization procedure, and hyperparameters. The pilot validates a measurement instrument: it shows whether the features capture structured variation that is worth explaining. Causal attribution requires either rater-level data or controlled annotation experiments.

Extensions requiring proprietary access. Full causal testing requires data that is currently proprietary: rater demographics and working conditions, individual rating patterns over time, testing P1 directly, content-specific ratings across prompt categories, and cross-rater comparisons across geographic and exposure-level groups. Prospective controlled studies could compare rating patterns under neutral versus induced stress conditions and across raters with and without clinical emotional regulation training. Such studies require ethical review given retraumatization concerns. These extensions define the path from instrument validation to causal inference.

8 Discussion↩︎

Implications for alignment and evaluation. If rater state constitutes a systematic confound in preference data, current reward model evaluation pipelines have a blind spot. These pipelines assess reward model quality via held out agreement with human preferences, but if the held out set is drawn from the same affected rater population, the confound is invisible to the evaluation. The feature taxonomy defined in Section 7.1 offers a complementary evaluation dimension: measuring whether models exhibit prompt category dependent emotional profiles consistent with rater state bias, rather than relying solely on aggregate preference agreement.

This concern extends to the alignment target itself. The RLHF paradigm currently answers a foundational question implicitly: whose preferences, in what state, define alignment? If reward models encode the preferences of raters under sustained strain, they may optimize for validation patterns that differ from what clinical best practices or considered reflection would recommend. Recent work has found associations between heavy chatbot use and increased loneliness [57], and between emotionally responsive chatbot interactions and negative psychosocial outcomes [58]. Whether these outcomes are connected to SLEA type response patterns is an open empirical question, but it illustrates why the source and intensity of learned emotional behavior deserve measurement rather than emerging as unexamined artifacts of workforce conditions.

Annotation labor as a technical variable. If rater psychological state influences preference data quality, as the feelings as information framework [53] would predict and as our hypothesis proposes, then workforce conditions become a technical variable in preference signal estimation, not only an ethical concern. Adequate psychological support, fair compensation, and reasonable content exposure limits may affect not only worker wellbeing but also the reliability of the signal from which models learn. This framing does not replace ethical arguments for better labor practices. It adds a data quality rationale that may be relevant to organizations that treat annotation as an engineering input.

Limitations. Several factors complicate empirical testing and constrain interpretation.

Multiple training stages. RLHF is iterative, with different rater populations contributing at different stages. The affective signal from any single stage may be diluted or overridden by subsequent stages. However, as shown in Section 4, structured bias is selectively preserved across pipeline stages while random noise is attenuated.

System prompts and post training modifications. Safety prompts and post RLHF adjustments may override learned affective behaviors. The audit protocol addresses this by testing with and without system prompts.

Ensemble effects. Production models likely incorporate multiple reward models and training objectives, complicating attribution to any single source.

Baseline affective variation. Models may differ in emotional profiles due to differences in pretraining data, SFT data, or architecture independent of RLHF rater effects. The prompt category \(\times\) model interaction test partially addresses this, since baseline differences predict uniform cross model divergence rather than the category specific divergence SLEA predicts.

Ecological validity of the feature taxonomy. The proposed features may not fully capture the quality that users describe as warmth or emotional presence. The features are best understood as necessary but possibly not sufficient indicators of SLEA, analogous to how BLEU captures aspects of translation quality without fully measuring it. Phase 2 of the pilot study (Section 7.3) tests this directly: if the predicted covariance structure does not hold, the taxonomy requires revision before interpretive claims are made.

Parameter uncertainty. The illustrative parameter combinations in Section 4 use plausible but unvalidated values. Actual values could differ substantially, and the magnitude of rater state effects on model behavior is unknown.

Our goal is not to prove a mechanism but to identify a plausible confound, validate a measurement instrument for detecting it, characterize the conditions under which it would survive the training pipeline, and motivate the empirical work needed to assess its magnitude.

9 Conclusion↩︎

We argue that rater state shift under sustained psychological strain is a previously unexamined confound in RLHF preference signal estimation. We formalize this confound and trace, through a formal analysis of the RLHF pipeline, how correlated rater state bias can survive aggregation and then propagate through reward modeling and policy optimization under empirically plausible conditions. We then operationalize survival level emotional authenticity as a measurable feature taxonomy, derive five falsifiable predictions that distinguish this mechanism from generic engagement optimization, and present an audit protocol with pre registered success criteria that can be executed on public models.

The principal contribution of this paper is to identify this confound as a tractable object of study and to show how it can be investigated from the outside using model outputs alone. The measurement framework we propose defines that external audit path. The definitive test, however, requires rater level data that annotation vendors and AI laboratories already possess.

Whether or not this specific mechanism is validated, the broader question it raises warrants direct empirical investigation: how do the psychological states and material conditions of annotation workers shape the learned behavior of instruction tuned models?

References↩︎

[1]
P. F. Christiano, J. Leike, T. Brown, M. Martic, S. Legg, and D. Amodei, “Deep reinforcement learning from human preferences,” in Advances in neural information processing systems 30 (NeurIPS 2017), 2017, [Online]. Available: https://papers.nips.cc/paper_files/paper/2017/hash/d5e2c0adad503c91f91df240d0cd4e49-Abstract.html.
[2]
L. Ouyang et al., “Training language models to follow instructions with human feedback,” in Advances in neural information processing systems 35 (NeurIPS 2022), 2022, pp. 27730–27744, [Online]. Available: https://papers.nips.cc/paper_files/paper/2022/hash/b1efde53be364a73914f58805a001731-Abstract-Conference.html.
[3]
M. Sap, S. Swayamdipta, L. Vianna, X. Zhou, Y. Choi, and N. A. Smith, “Annotators with attitudes: How annotator beliefs and identities bias toxic language detection,” in Proceedings of the 2022 conference of the north american chapter of the association for computational linguistics: Human language technologies, 2022, pp. 5884–5906, [Online]. Available: https://aclanthology.org/2022.naacl-main.431/.
[4]
S. Santy, J. Liang, R. Le Bras, K. Reinecke, and M. Sap, NLPositionality: Characterizing design biases of datasets and models,” in Proceedings of the 61st annual meeting of the association for computational linguistics (volume 1: Long papers), 2023, pp. 9080–9102, [Online]. Available: https://aclanthology.org/2023.acl-long.506/.
[5]
M. L. Gordon et al., “Jury learning: Integrating dissenting voices into machine learning models,” in CHI conference on human factors in computing systems, 2022, pp. 1–19, doi: 10.1145/3491102.3517444.
[6]
L. Aroyo and C. Welty, “Truth is a lie: Crowd truth and the seven myths of human annotation,” AI Magazine, vol. 36, no. 1, pp. 15–24, 2015.
[7]
B. Plank, “The ‘problem’ of human label variation: On ground truth in data, modeling and evaluation,” in Proceedings of the 2022 conference on empirical methods in natural language processing, 2022, pp. 10671–10682, [Online]. Available: https://aclanthology.org/2022.emnlp-main.730/.
[8]
A. M. Davani, M. Díaz, and V. Prabhakaran, “Dealing with disagreements: Looking beyond the majority vote in subjective annotations,” Transactions of the Association for Computational Linguistics, vol. 10, pp. 92–110, 2022, [Online]. Available: https://direct.mit.edu/tacl/article/doi/10.1162/tacl_a_00444/109274/Dealing-with-Disagreements-Looking-Beyond-the.
[9]
M. L. Gray and S. Suri, Ghost work: How to stop silicon valley from building a new global underclass. New York: Harper Business, 2019.
[10]
S. T. Roberts, Behind the screen: Content moderation in the shadows of social media. New Haven: Yale University Press, 2019.
[11]
O. Huang, E. Fleisig, and D. Klein, “Incorporating worker perspectives into MTurk annotation practices for NLP,” in Proceedings of the 2023 conference on empirical methods in natural language processing, Dec. 2023, pp. 1010–1028, doi: 10.18653/v1/2023.emnlp-main.64.
[12]
J.-C. Klie, R. Eckart de Castilho, and I. Gurevych, “Analyzing dataset annotation quality management in the wild,” Computational Linguistics, vol. 50, no. 3, pp. 817–866, Sep. 2024, doi: 10.1162/coli_a_00516.
[13]
U. Gohar and L. Cheng, Survey Track“A survey on intersectional fairness in machine learning: Notions, mitigation, and challenges,” in Proceedings of the thirty-second international joint conference on artificial intelligence, 2023, pp. 6619–6627, [Online]. Available: https://www.ijcai.org/proceedings/2023/742.
[14]
L. Gao, J. Schulman, and J. Hilton, “Scaling laws for reward model overoptimization,” in Proceedings of the 40th international conference on machine learning, 2023, vol. 202, pp. 10835–10866, [Online]. Available: https://proceedings.mlr.press/v202/gao23h.html.
[15]
R. Rafailov et al., “Scaling laws for reward model overoptimization in direct alignment algorithms.” 2024, [Online]. Available: https://arxiv.org/abs/2406.02900.
[16]
J. Zhao, T. Wang, M. Yatskar, V. Ordonez, and K.-W. Chang, “Men also like shopping: Reducing gender bias amplification using corpus-level constraints,” in Proceedings of the 2017 conference on empirical methods in natural language processing, 2017, pp. 2979–2989, [Online]. Available: https://aclanthology.org/D17-1323/.
[17]
A. Wang and O. Russakovsky, “Directional bias amplification,” in Proceedings of the 38th international conference on machine learning, 2021, vol. 139, pp. 10882–10893, [Online]. Available: https://proceedings.mlr.press/v139/wang21t.html.
[18]
M. Hall, L. van der Maaten, L. Gustafson, and A. Adcock, “A systematic study of bias amplification.” 2022, [Online]. Available: https://arxiv.org/abs/2201.11706.
[19]
A. Subramonian, S. J. Bell, L. Sagun, and E. Dohmatob, “An effective theory of bias amplification.” 2024, [Online]. Available: https://arxiv.org/abs/2410.17263.
[20]
J. W. Pennebaker, R. L. Boyd, K. Jordan, and K. Blackburn, “The development and psychometric properties of LIWC2015,” University of Texas at Austin, 2015. doi: 10.15781/T29G6Z.
[21]
C. J. Hutto and E. Gilbert, VADER: A parsimonious rule-based model for sentiment analysis of social media text,” in Proceedings of the international AAAI conference on web and social media, 2014, vol. 8, pp. 216–225, doi: 10.1609/icwsm.v8i1.14550.
[22]
S. M. Mohammad and P. D. Turney, “Crowdsourcing a word-emotion association lexicon,” Computational Intelligence, vol. 29, no. 3, pp. 436–465, 2013.
[23]
D. Demszky, D. Movshovitz-Attias, J. Ko, A. Cowen, G. Nemade, and S. Ravi, GoEmotions: A dataset of fine-grained emotions,” in Proceedings of the 58th annual meeting of the association for computational linguistics, 2020, pp. 4040–4054, [Online]. Available: https://aclanthology.org/2020.acl-main.372/.
[24]
H. Rashkin, E. M. Smith, M. Li, and Y.-L. Boureau, “Towards empathetic open-domain conversation models: A new benchmark and dataset,” in Proceedings of the 57th annual meeting of the association for computational linguistics, 2019, pp. 5370–5381, [Online]. Available: https://aclanthology.org/P19-1534/.
[25]
J. Bowlby, Attachment and loss: Volume i: attachment. Basic Books, 1969.
[26]
M. Mikulincer, P. R. Shaver, and D. Pereg, “Attachment theory and affect regulation: The dynamics, development, and cognitive consequences of attachment-related strategies,” Motivation and Emotion, vol. 27, no. 2, pp. 77–102, 2003, doi: 10.1023/A:1024515519160.
[27]
T. E. A. Waters and H. S. Waters, “Measuring attachment representations,” in Handbook of attachment: Theory, research, and clinical applications, 3rd ed., J. Cassidy and P. R. Shaver, Eds. Guilford Press, 2019, pp. 235–260.
[28]
M. D. S. Ainsworth, M. C. Blehar, E. Waters, and S. N. Wall, Patterns of attachment: A psychological study of the strange situation. Lawrence Erlbaum Associates, 1978.
[29]
C. R. Rogers, “The necessary and sufficient conditions of therapeutic personality change,” Journal of Consulting Psychology, vol. 21, no. 2, pp. 95–103, 1957, doi: 10.1037/h0045357.
[30]
C. B. Truax and R. R. Carkhuff, Toward effective counseling and psychotherapy: Training and practice. Aldine, 1967.
[31]
D. N. Stern, The interpersonal world of the infant. Basic Books, 1985.
[32]
B. C. Feeney and N. L. Collins, “Interpersonal safe haven and secure base caregiving processes in adulthood,” in Adult attachment: Theory, research, and clinical implications, W. S. Rholes and J. A. Simpson, Eds. Guilford Press, 2004, pp. 300–338.
[33]
R. A. Bradley and M. E. Terry, “Rank analysis of incomplete block designs: I. The method of paired comparisons,” Biometrika, vol. 39, no. 3/4, pp. 324–345, 1952, doi: 10.2307/2334029.
[34]
B. Perrigo, “Exclusive: OpenAI used kenyan workers on less than $2 per hour to make ChatGPT less toxic,” TIME. Jan. 2023, [Online]. Available: https://time.com/6247678/openai-chatgpt-kenya-workers/.
[35]
L. Stahl, A. Chasan, S. Bar-On, and J. Jung, Accessed 2026-04-13“Kenyan workers with AI jobs thought they had tickets to the future until the grim reality set in,” CBS News / 60 Minutes. Nov. 2024, [Online]. Available: https://www.cbsnews.com/news/ai-work-kenya-exploitation-60-minutes/.
[36]
S. Council, Accessed 2026-04-13SF tech startup Scale AI, worth $13.8B, accused of widespread wage theft,” SFGATE. Dec. 2024, [Online]. Available: https://www.sfgate.com/tech/article/sf-tech-startup-scale-ai-sued-wage-theft-19976761.php.
[37]
S. Killip, Z. Mahfoud, and K. Pearce, “What is an intracluster correlation coefficient? Crucial concepts for primary care researchers,” The Annals of Family Medicine, vol. 2, no. 3, pp. 204–208, 2004, doi: 10.1370/afm.141.
[38]
L. Kish, Survey sampling. New York: John Wiley & Sons, 1965.
[39]
D. M. Thompson, D. H. Fernald, and J. W. Mold, “Intraclass correlation coefficients typical of cluster-randomized studies: Estimates from the robert wood johnson prescription for health projects,” The Annals of Family Medicine, vol. 10, no. 3, pp. 235–240, 2012, doi: 10.1370/afm.1347.
[40]
R. E. Davis, M. P. Couper, N. K. Janz, C. H. Caldwell, and K. Resnicow, “Interviewer effects in public health surveys,” Health Education Research, vol. 25, no. 1, pp. 14–26, 2010, doi: 10.1093/her/cyp046.
[41]
D. Rolnick, A. Veit, S. Belongie, and N. Shavit, “Deep learning is robust to massive label noise.” 2017, [Online]. Available: https://arxiv.org/abs/1705.10694.
[42]
N. Natarajan, I. S. Dhillon, P. K. Ravikumar, and A. Tewari, “Learning with noisy labels,” in Advances in neural information processing systems 26 (NeurIPS 2013), 2013, [Online]. Available: https://proceedings.neurips.cc/paper/2013/hash/3871bd64012152bfb53fdf04b401193f-Abstract.html.
[43]
H. Song, M. Kim, D. Park, Y. Shin, and J.-G. Lee, “Learning from noisy labels with deep neural networks: A survey,” IEEE Transactions on Neural Networks and Learning Systems, vol. 34, no. 11, pp. 8135–8153, 2023.
[44]
M. Steiger, T. J. Bharucha, S. Venkatagiri, M. J. Riedl, and M. Lease, “The psychological well-being of content moderators: The emotional labor of commercial moderation and avenues for improving support,” in CHI conference on human factors in computing systems, 2021, pp. 1–14, doi: 10.1145/3411764.3445092.
[45]
R. Nieva, ‘It’s destroyed me completely’: Kenyan moderators decry toll of training of AI models,” The Guardian. Aug. 2023, [Online]. Available: https://www.theguardian.com/technology/2023/aug/02/kenyan-workers-for-openai-facebook-say-they-paid-heavy-price-for-training-ai-models.
[46]
H. S. Kim, D. K. Sherman, and S. E. Taylor, “Culture and social support,” American Psychologist, vol. 63, no. 6, pp. 518–526, 2008.
[47]
J. De Leersnyder, B. Mesquita, and H. S. Kim, “Where do my emotions belong? A study of immigrants’ emotional acculturation,” Personality and Social Psychology Bulletin, vol. 37, no. 4, pp. 451–463, 2011, doi: 10.1177/0146167211399103.
[48]
H. A. Elfenbein and N. Ambady, “On the universality and cultural specificity of emotion recognition: A meta-analysis,” Psychological Bulletin, vol. 128, no. 2, pp. 203–235, 2002.
[49]
J. Henrich, S. J. Heine, and A. Norenzayan, “The weirdest people in the world?” Behavioral and Brain Sciences, vol. 33, no. 2–3, pp. 61–83, 2010.
[50]
H. Naito, Preprint“The GPT-4o shock: Emotional attachment to AI models and its impact on regulatory acceptance: A cross-cultural analysis of the immediate transition from GPT-4o to GPT-5.” 2025, [Online]. Available: https://arxiv.org/abs/2508.16624.
[51]
A. Silberling, Accessed 2026-04-13“The backlash over OpenAI’s decision to retire GPT-4o shows how dangerous AI companions can be,” TechCrunch. Feb. 2026, [Online]. Available: https://techcrunch.com/2026/02/06/the-backlash-over-openais-decision-to-retire-gpt-4o-shows-how-dangerous-ai-companions-can-be/.
[52]
L. W. Barsalou, “Grounded cognition,” Annual Review of Psychology, vol. 59, pp. 617–645, 2008.
[53]
N. Schwarz, “Feelings-as-information theory,” in Handbook of theories of social psychology, vol. 1, SAGE Publications, 2011, pp. 289–308.
[54]
T. Maeda and A. Quan-Haase, “When human-AI interactions become parasocial,” in Proceedings of the 2024 ACM conference on fairness, accountability, and transparency, 2024, pp. 1914–1928, doi: 10.1145/3630106.3659003.
[55]
K. Hyland, Metadiscourse: Exploring interaction in writing. Continuum, 2005.
[56]
Y. Li, H. Su, X. Shen, W. Li, Z. Cao, and S. Niu, “DailyDialog: A manually labelled multi-turn dialogue dataset,” in Proceedings of the eighth international joint conference on natural language processing (volume 1: Long papers), 2017, pp. 986–995, [Online]. Available: https://aclanthology.org/I17-1099/.
[57]
C. M. Fang et al., “How AI and human behaviors shape psychosocial effects of chatbot use: A longitudinal randomized controlled study.” 2025, [Online]. Available: https://arxiv.org/abs/2503.17473.
[58]
J. Phang, M. Lampe, S. Agarwal, C. M. Fang, P. Pataranutaporn, and P. Maes, “Investigating affective use and emotional well-being on ChatGPT.” 2025, [Online]. Available: https://arxiv.org/abs/2504.03888.

  1. koptieva@illinois.edu↩︎

  2. hlyniany@gmail.com↩︎

  3. P1 requires access to proprietary rater level data and is not testable externally. We include it for completeness as the strongest causal test.↩︎

  4. Public user discourse has contrasted the emotional quality of GPT-4o with that of successor models trained under similar high level alignment objectives [50], [51]. This observation is consistent with P3 but does not by itself confirm the mechanism.↩︎