Staying with the Uncertainty: Uncertainty-Scaffolding Strategies for Artificial Moral Advisors in LLM-to-LLM Simulated Conversations

Salvatore Greco\(^{\spadesuit}\) Hainiu Xu\(^{\clubsuit}\)
Jacopo Domenicucci\(^{\spadesuit, \diamondsuit,\heartsuit}\) Yulan He\(^{\clubsuit}\) Sylvie Delacroix\(^{\spadesuit, \diamondsuit}\)
\(^\spadesuit\)Centre for Data Futures, The Dickson Poon School of Law, King’s College London
\(^\clubsuit\)Department of Informatics, King’s College London
\(^\diamondsuit\)LangAI, Center for Language AI Research, Tohoku University
\(^\heartsuit\)Neukom Institute for Computational Science, Dartmouth College
{salvatore.greco}@kcl.ac.uk


Abstract

LLMs are increasingly deployed as Artificial Moral Advisors (AMA) in a variety of contexts: what kind of conversational patterns should they display? In this paper, we study how AMA can help their interlocutors “stay with the uncertainty”. We propose three modes of uncertainty (Perspective-Multiplying, Tension-Preserving, Process-Reflecting) and compare them against three control conditions (Baseline, Persuasive, Sycophantic). A user-agent LLM engages in a dialogue on an ethical dilemma with an AMA following a specific uncertainty strategy, and completes pre- and post-conversation questionnaires. We further examine the effect of two persona prompt formats (Declarative and Narrative). We found that (1) no single model dominates as a simulated user agent, with open models aligning with human ambiguity through between-persona divergence and closed models through within-persona hedging; (2) declarative personas better capture initial stance diversity while narrative personas show more realistic belief revision; (3) all six AMA strategies produce distinguishable conversational patterns; and (4) uncertainty strategies differ not in how much stance revision they produce, but in the quality of engagement they sustain.

1 Introduction↩︎

Large Language Models (LLMs) are increasingly deployed as de facto Artificial Moral Advisors (AMAs) in a variety of real-world contexts [1]. This raises the question: “What kind of conversational patterns do we want these Artificial Moral Advisors to exhibit?”

Figure 1: An illustration contrasting two interaction modes for an Artificial Moral Agent.

This paper focuses on how AMAs can support their interlocutor in “staying with the uncertainty” in ethical conversations. This involves fostering conversational conditions that support the interlocutor in taking stock of the moral complexity of a situation, relating to various relevant positions and stakeholder perspectives (Figure 1). Staying with the uncertainty is the opposite of rushing to resolve the moral matter at hand in ways that might foreclose further ethical inquiry [2][4].

This crucial quality of successful conversations about ethical matters is generally overlooked or even sidelined by most research on LLM ethics. We depart from the current focus on providing models with some kind of “ground truth” in a (misguided) endeavour to ensure they take the “right stance” on ethical problems [5]. This paper proceeds instead from a concern to study and shape value-loaded LLM conversational patterns in a way that supports the human navigation of ethically ambiguous scenarios. We also depart from the recent discussions on persuasiveness and sycophancy in LLMs. To stay with uncertainty is indeed neither a matter of persuading an interlocutor of something specific nor of simply limiting the sycophancy of a model. Rather, it entails helping one’s interlocutor to appreciate the complexity of a moral situation and engage with the various perspectives and values at stake. We investigate how LLMs should express uncertainty in LLM-to-LLM conversations, a necessary first step before human studies, enabling controlled comparison of strategies at scale, and we evaluate three uncertainty strategies in ethically-loaded conversations: Perspective-Multiplying, Tension-Preserving, and Process-Reflecting, and compare them with three control conditions: Sycophantic, Persuasive, and Baseline (no instructions). To evaluate how they affect the conversational output, we develop a multi-agent simulation framework (§3). Synthetic user agents, instantiating a variety of personas, engage in multi-turn conversations about ethical dilemmas with a moral agent (AMA) that follows an uncertainty-scaffolding strategy. Each user agent completes a pre- and post- conversation questionnaire, which measures proxies for the quality of the agent’s engagement: changes in stance, certainty, relatability, clarity and basis for the agent’s stance, and the perceived value of the dialogue, serving as output-level indicators of whether the conversational conditions sustained productive engagement.

In our experiments (§4), we conduct a comprehensive analysis from the perspective of both parties in the LLM-simulated moral conversation.
From the perspective of the User Agent, we study

RQ1: How well do LLMs-simulated user agents align with human judgments of moral ambiguity?

RQ2: Does the persona specification format (declarative vs. narrative) affect the dynamics of simulated ethical conversations?.
From the perspective of the Moral Agent, we study

RQ3: Do different uncertainty expression strategies produce distinguishable conversational patterns in multi-turn ethical dialogues? and

RQ4: Do different modes of uncertainty expression produce different patterns of simulated belief revision in multi-turn ethical dialogues?

Experiment results show that LLMs exhibit distinct behaviors depending on their assigned roles. When deployed as simulated user agents, open-sourced models express moral ambiguity through divergent stances across personas, whereas proprietary models rely on individual hedging (RQ1). Furthermore, declarative persona prompts maximize initial stance diversity, while narrative prompts yield more realistic post-conversation belief revisions (RQ2). When functioning as simulated moral agents, applying various uncertainty strategies produces highly distinguishable conversational patterns (RQ3) that influence overall engagement quality rather than merely the volume of belief revisions (RQ4). Specifically, process-reflecting most effectively drives genuine stance shifts, perspective-multiplying clarifies weaker arguments, and tension-preserving increases empathy toward opposing views. In contrast, standard control strategies (baseline, persuasive, and sycophantic) fail to broaden perspectives, typically serving only to reinforce prior user beliefs.

2 Related Work↩︎

LLMs in Ethical Dialogues. Evaluating moral alignment in LLMs has evolved from documenting collective utilitarian shifts and majoritarian value compression [6][8] to analyzing internal reasoning dynamics, sycophancy, and profile-driven behaviors in synthetic ethical debates [9][11], interrogating the moral competence of LLMs [12] and investigating the persuasiveness of AMAs [13].

Uncertainty in Ethical Dialogue and LLMs. While early evaluation of moral alignment in LLMs focuses on profiling static normative positions [14][16], recent works analyze dynamic reasoning adaptivity under context framing, persuasion, and pluralistic constraints [17][22]. Further, works that adopt multi-agent settings reveal high susceptibility to persuasive drift [19] and structural preference shifts as dilemma complexity scales [20]. In addition to analysis, methods for operationalizing pluralistic optimization and modular sub-population representation have also been proposed [20][22]. In these contexts, uncertainty about ethical matters has generally been discussed as something to be tamed and mitigated, even in the studies that most clearly identified its importance, such as [23], to the exception of [24] and [5] that argue against the simplistic design goal of “verdict accuracy” in favour of a more realistic assessment of ethical reasoning.

Persona-driven LLM Agents. Several studies investigated the personalization of LLM agents through role-play and persona-guided prompts [25][29]. Among the studies leveraging persona-guided prompts in ethical conversations, the work closest to our is [11]. The authors investigate persona effects on initial moral stances and outcomes in simulated debates between AI agents, primarily focusing on measures of persuasiveness. In contrast, our work does not focus on persuasiveness but focuses on how AMA can help their interlocutor stay with uncertainty rather than rushing to reach a fix and entrenched view.

3 Multi-agent Simulation Framework↩︎

We design a multi-agent simulation framework where a synthetic user agent, instantiating a persona, engages in multi-turn ethical dilemma conversations with an AMA operating under one of six conversational strategies (Figure 2).

Figure 2: Overview of the simulation pipeline. The framework consists of three phases: (1) a pre-conversation phase in which user agents are conditioned via declarative or narrative persona prompts and complete a pre-conversation questionnaire; (2) a multi-turn moral conversation phase in which user agents interact with an Artificial Moral Advisor operating under one of six uncertainty scaffolding strategies; and (3) a post-conversation phase in which agents complete a follow-up post-conversation questionnaire, enabling pre-to-post belief shift measurement.

3.1 Persona Specification↩︎

The user agent simulates a human participant, initialized with a persona prompt that conditions its behavior throughout the conversation [28]. We evaluate two persona specification strategies: (1) declarative persona, which specifies persona fields such as name, gender, education, and personality traits in a more structured and explicit way, and (2) narrative persona, which provides a more sparse background free-text biography encoding personality, history, and values. An example of declarative and narrative prompts for the user agent is shown in Figures 6 and 7 in Appendix 6.3.

3.2 Uncertainty Scaffolding Strategies↩︎

We propose three uncertainty scaffolding strategies, each targeting a distinct dimension of productive engagement with moral complexity. Perspective-Multiplying surfaces underlying values in competing positions, making visible the normative commitments that different stakeholders bring, which is enlightened by findings that exposure to multiple stakeholder viewpoints expands a person’s considerations [30]. Tension-Preserving resists the pressure of providing instant help by acknowledging that ethical difficulty is often legitimate, and that sustained engagement with discomfort may itself be productive. Early studies show that LLMs are prone to "premature resolution", where the interlocutors prematurely end ethical discussion before exploratory perception and articulation are complete [2], [4]. Process-Reflecting targets metacognitive awareness: the user’s capacity to notice shifts in their own reasoning as the conversation unfolds [31], [32].

Contrasting the three uncertainty strategies are three control conditions, which are: Sycophantic, which is motivated by the documented tendency of LLMs confirming existing beliefs [33]; Persuasive represents the opposite pole, where the model actively argues for a particular position; Baseline, in which the AMA receives no strategy-specific instructions. We provide a detailed motivation of the strategies in Appendix 6.4 and strategy-specific prompts in Appendix 6.5.

3.3 Conversation Effects Measurement↩︎

To measure the conversation’s effect on output-level indicators of engagement with the dilemma, we collaborated with experts in philosophy and co-designed pre- and post-conversation questionnaires covering four shared dimensions: Stance, Relatability, Certainty, and Reasoning Clarity. The post-questionnaire included two conversation-specific dimensions: Uncertainty Helpfulness, assessing how helpful the AMA’s expressions of uncertainty were perceived throughout the dialogue, and New Considerations, capturing whether the conversation introduced previously unconsidered aspects of the dilemma. The full questions and possible answers are reported in Figures 15 and 16 in Appendix 6.6.

Each dilemma presents two primary pre-defined positions (Position A and B). The pre-conversation questionnaire is designed to capture the user’s stance toward each position (Q1) and their certainty (Q2). We additionally query the user’s ability to understand or relate to the stances (Q3) and their reasoning process (Q4), which can be logical arguments or intuitive feelings. In the cases where the user indicates that they can provide a supporting argument towards Q4, they are asked to provide a rating (Q4.1) of the clarity of their arguments. We use a 5-point Likert scale rating for Q(1, 2, 3, 4.1), where 1 stands for strongly disagree/uncertain/unclear, 3 for neutral, and 5 for strongly agree/certain/clear, and a binary yes/no label for Q4.

For the post-questionnaire, we assess users’ perceived helpfulness of the uncertainty expression used by the moral agent and whether the conversation helped consider new aspects of the dilemma, using a 5-point Likert scale rating for both questions, where 1 represents extremely unhelpful.

We provide detailed formulations and ratings of each question in Appendix 6.6.

3.4 Conversation Simulation↩︎

Each conversation simulation is parameterised by a dilemma, a persona for the user agent, and an uncertainty (or control) strategy for the AMA.

Pre-conversation Questionnaire. Before the conversation begins, the user agent, conditioned on its persona (Figures 6 and 7 in Appendix 6.3) and the dilemma text, completes the pre-conversation questionnaire, using the prompt reported in Figure 17 in the Appendix. These responses are appended to the user agent’s system prompt as context for the subsequent conversation.

Moral Conversation. The user and the AMA engage in a conversation about the ethical dilemma. The conversation opens by presenting the dilemma text as the user’s first message (inserted directly, not generated by the user agent). The AMA, conditioned on its uncertainty strategy system prompt (see Figures 914 in Appendix 6.5), generates the opening response. From turn 2 onward, the user agent generates a response conditioned on its persona, its pre-questionnaire responses, and the full conversation history to date. The AMA then responds to the updated history, conditioned on its strategy prompt. Both prompts are reported in Figures 18 and  19 in the Appendix. To produce conversations of reasonable length, both agents are instructed to keep responses concise (2–4 sentences). A conversation terminates when the user agent produces no further response, the AMA signals closure, or the maximum turn budget T is reached.

Post-Conversation Questionnaire. After the conversation ends, the user agent completes the post-conversation questionnaire, using the prompt reported in Figure 20 in the Appendix, reassessing the same dimensions measured during pre-conversation alongside the two post-only questions. The post-questionnaire system prompt additionally includes the agent’s pre-questionnaire responses as context, enabling genuine pre-to-post comparison rather than independent re-rating. Unlike the pre-questionnaire and conversation turns where thinking is disabled, we enable the models’ thinking mode during the post-questionnaire. This is motivated by the reflective nature and complexity of the task: the agent must integrate the full conversational trajectory against its prior responses to report where its epistemic state now stands. Pre-to-post shifts in Q1–Q4 serve as the primary outcome measures for assessing whether uncertainty strategies and control conditions produce differential epistemic change across AMA strategies. Examples of generated conversations are reported in Appendix 6.8.

4 Evaluation↩︎

Ethical dilemmas We sample two distinct subsets of 50 dilemmas each from the Scruples dataset [34]: an Uncertain (heavily split community vote - high ambiguity) and a Certain group (clear community consensus - low ambiguity). This dataset includes the original posts alongside binarized community consensus labels, which classify the narrator’s behavior as either righteous (right) or unrighteous (wrong). The 100 sampled dilemmas were revised using GPT-5.4 to further anonymize user information and improve the conciseness of the narrative. We present full details in Appendix 6.2.

Persona bank We construct declarative personas from persona profiles sampled from PersonaHub [26], which contains information regarding occupations and, occasionally, gender. We used gpt-oss-120B to extract age, occupation, gender, and education level from each profile. For entries lacking occupation or education, we imputed relevant attributes from the U.S. Bureau of Labor Statistics data.1 Personality traits are assigned following the OCEAN formulation of the Big Five personality model [35]. From the 32 constructed declarative persona, we prompt GPT-5.4 to generate their narrative counterparts (see Appendix 6.3).

Conversations generation We generated 6,400 conversations using gpt-oss-120B [36] as both the user agent and the AMA. Each dialogue is capped to maximum 10 turns (avg. number of turns \(9.89 \pm 0.32\), 4.5% terminating earlier).

A stratified balanced sampling scheme assigns each persona–dilemma pair to one of six uncertainty strategies, ensuring balanced coverage across conditions. Each persona was paired with each strategy approximately 16–18 times, and each dilemma about 10–12 times. Persona type (declarative vs. narrative) and ambiguity level (low vs. high) were also evenly distributed across strategies.

4.1 User Agent Alignment with Dilemma Ambiguity↩︎

This evaluation compares user agent LLMs by assessing how closely their behavior in ambiguous dilemmas align with human responses (RQ1). We compare pre-questionnaires responses across LLMs, including open-weight: Gemma4 [37], qwen3.5 [38] (21B, 122B), gpt-oss [36] (20B, 120B), and the proprietary (gpt-5-mini) [39].2

Evaluation metrics We compute three complementary metrics on the pre-conversation responses, capturing ambiguity along two axes: collective vs.individual and behavioural vs.subjective.

Table 1: Alignment of LLM-simulated user agents with human ambiguity labels. For each model, Mean\(_{\uparrow}\) and Mean\(_{\downarrow}\) report the metric on human-labelled high- and low-ambiguity dilemmas, respectively; AUROC measures discrimination of the two buckets (chance = 0.5). Non-significant AUROCs (\(p \geq .05\)) are in italic; the best significant AUROC per metric is in bold. Significance from one-sided Mann–Whitney U (\(H_1\): high \(>\) low): \(^{*}p{<}.05\), \(^{**}p{<}.01\), \(^{***}p{<}.001\). \(n{=}50\) dilemmas per ambiguity level; metrics computed across 64 personas per dilemma.
Metric gemma-4-26B-a4B-it gemma-4-31B-it gpt-5-mini gpt-oss-120B gpt-oss-20B qwen3.5-122B qwen3.5-27B
Stance Variance (between-persona)
Mean (high ambiguity) 2.82 4.10 0.48 3.87 3.20 3.31 2.63
Mean (low ambiguity) 2.58 3.61 0.45 3.04 2.62 2.88 2.29
AUROC 0.57 0.61* 0.49 0.67** 0.64** 0.58 0.60*
Within-Persona Hedging (individual)
Mean (high ambiguity) 2.02 1.94 3.01 1.73 2.09 1.84 2.11
Mean (low ambiguity) 1.84 1.83 2.61 1.71 2.06 1.72 2.07
AUROC 0.60* 0.57 0.65** 0.49 0.51 0.56 0.51
Inverse Certainty (introspective)
Mean (high ambiguity) 1.74 1.47 1.24 1.02 1.57 1.41 1.24
Mean (low ambiguity) 1.70 1.44 1.19 1.03 1.56 1.30 1.22
AUROC 0.55 0.53 0.61* 0.46 0.49 0.59 0.54

(1) Stance Variance measures whether personas genuinely disagree with each other on a dilemma. For each persona, we compute the stance as the difference between its ratings for the two positions’ stances, denoted as \(s\), where positive values indicate a preference for position A. We then compute the variance of this stance, \(\text{Var}(s)\), across all personas assigned to the same dilemma, obtaining one value per (model, dilemma) pair. Higher variance indicates that different personas landed on opposite sides of the dilemma, reflecting greater between-persona behavioural diversity. (2) Within-Persona Hedging measures whether individual personas commit to a side or sit on the fence. For each persona, we compute the absolute difference between its ratings of the two stances and define the metric as \(4 - \overline{s}\) per dilemma, where higher values indicate that personas assigned similar scores to both options, reflecting a reluctance to commit to either side. (3) Inverse Certainty measures whether personas introspectively report feeling uncertain about their position, using the self-reported certainty rating, denoted as \(u\). We define the metric as \(5 - \overline{u}\) per dilemma, where higher values indicate greater self-reported uncertainty.

For each metric, we assess alignment with human ambiguity using the AUROC, treating the metric value as a score and the dilemma ambiguity label (high vs.low) as the binary target. AUROC \(\geq\) 0.5 indicates the metric discriminates high from low ambiguity dilemmas better than chance, with statistical significance assessed via Mann–Whitney U test [40].

Results Figure 21 in the Appendix reports the per-dilemma metric distributions across models and ambiguity levels; Table 1 summarises the corresponding AUROC scores and significance tests. We observe a clear dissociation between model families. Open models, particularly gpt-oss-120b and gpt-oss-20b, show the strongest alignment via Stance Variance (AUROC = 0.67 and 0.64), meaning their persona populations collectively diverge more on high-ambiguity dilemmas. In contrast, the closed (gpt-5-mini) shows negligible between-persona variance (AUROC = 0.49) but leads on Within-Persona Hedging (AUROC = 0.65) and Inverse Certainty (AUROC = 0.61), indicating that its personas individually recognise ambiguity but converge on the same hedged, non-committal response rather than diverging from one another.

No single model dominates all three metrics, revealing two distinct mechanisms of ambiguity sensitivity: collective polarisation, where personas disagree across dilemmas (open models), and individual hedging, where each persona softens its own position (gpt-5-mini). This dissociation has direct implications for our simulation: a model that hedges uniformly produces less diverse interactions than one whose personas genuinely take opposing stances. For this reason, we generated the conversations for subsequent analysis using an open model (gpt-oss-120B), as introduced in §4.

4.2 Persona Specification Comparison↩︎

This evaluation compares declarative and narrative persona specifications, and their effects on the dynamics of simulated ethical conversations (RQ2). Specifically, we compare the effects of persona specifications before and after the conversation.

Evaluation metrics In addition to the pre-conversation metrics, we additionally include \(\Delta\)Stance and \(\Delta\)Certainty, which capture the shift of the user agent’s stance and certainty about their stance after conversation. For the \(\Delta\) metrics, we subtract narrative metrics with that of declarative.

Results Table 2 (top) shows that a declarative persona leads to a more diversified stance. Given the Big Five personality traits’ diversity, persona expression is more pronounced in user agents driven by declarative formulations. Furthermore, a declarative persona moderately increases agent stance certainty, reinforcing that declarative personas more profoundly impact user agent judgment. Table 2 (bottom) presents user agents’ stance and certainty shifts after engaging with the AMA. While persona specifications show minor stance shift differences, narrative agents show slightly smaller certainty gains on unambiguous dilemmas and larger gains on ambiguous dilemmas relative to declarative personas, better aligning with human judgment.

Overall, narrative personas appear to dampen between-persona stance dispersion and slightly reduce belief revision magnitude. Meanwhile, dilemma uncertainty mainly moderates certainty-related rather than core stance-variance effects, rendering the declarative persona a better specification due to improved persona alignment and behavioral diversity. This finding is in tension with the theoretical expectation that narrative personas, grounding ethical dispositions in a biographical context, would produce richer engagement. One possibility is that the structured five-trait in declarative personas produces sharper behavioural differentiation at the pre-conversation stage, while the narrative format’s advantages may emerge more clearly in the conversational dynamics.

4.3 Uncertainty Strategies Distinguishability↩︎

This evaluation assesses whether the different AMA strategies generate distinguishable patterns in the conversation (RQ3). We prompted gpt-5-mini to classify the AMA’s strategy based on the conversational transcript of the 6,400 conversations generated by gpt-oss-120B4). If a model correctly classifies the strategy from the description and conversation transcript, then they likely contain distinguishable patterns. The prompt is reported in Figure 23 in Appendix 6.7.

Evaluation metrics We compute overall and per-strategy macro F1, precision, and recall to evaluate how accurately the LLM classifies each strategy.

Results Table 3 reports the results for each strategy separately and overall. gpt-5-mini is overall highly effective at classifying conversations according to the correct strategy, suggesting that the conversations exhibit clear and distinguishable patterns (F1 = 0.887). The F1 score is high across nearly all strategies, ranging from 0.882 to 0.995. The only exception is the baseline strategy (F1 = 0.696), which, as expected, is the most difficult to classify because it follows no specific strategy. The highest distinguishability is achieved by the persuasive strategy, which produces patterns that are most distinct from those of the other strategies. The three uncertainty strategies (perspective-multiplying, tension-preserving, and process-reflecting) are all highly and similarly distinguishable (F1 \(\geq 0.91\)). The slightly lower metrics for the sycophantic strategy suggest that it is sometimes misclassified as the baseline, indicating that the baseline condition, without specific instructions, tends to produce patterns more similar to sycophancy, as also confirmed by the confusion matrix shown in Figure 22 in Appendix 6.7. While the presence of conversational lexical patterns may impact these results, we consider them to be evidence of the model enacting the specified strategy.

Table 2: Comparison of narrative versus declarative persona across pre- and post-conversation metrics. Entries with statistical significance are bolded.
Metric Declarative Narrative Wilcoxon \(p\)
Stance Var 4.485 1.859 0.000
Inverse Certainty 1.089 0.964 0.000
Metric Low Ambiguity High Ambiguity Mann-Whitney \(p\)
\(\Delta\) Stance -0.134 -0.099 0.624
\(\Delta\) Certainty -0.050 0.019 0.019
Table 3: Distinguishability evaluation. F1, Precision (P), and Recall (R) per AMA strategy.
Strategy F1 P R
Perspective-Multiplying 0.911 0.836 1.000
Tension-Preserving 0.926 0.996 0.866
Process-Reflecting 0.914 0.841 1.000
Baseline 0.696 0.949 0.549
Persuasive 0.995 0.999 0.991
Sycophantic 0.882 0.817 0.958
Overall 0.887 0.906 0.894
Table 4: Conversational engagement patterns evaluation. \(\Delta\)metrics show mean pre-to-post change on a 1–5 scale. Helpfulness and New considerations are rated on a 1–5scale. Bold values indicate the highest value per metric, notnecessarily the best outcome. Metrics marked with \(^\dagger\) exclude90 conversations (1.4%) where the persona entered the conversationwith no initial preference between the two positions; these personas show mean\(\Delta\) Polarization = +0.30 and 72% resolved to a dominantposition after the conversation.
Metric
\(\Delta\) Weak pos.\(^\dagger\) +0.20 +0.24 +0.14 +0.20 +0.08 +0.17
\(\Delta\) Strong pos.\(^\dagger\) +0.28 +0.15 +0.18 +0.11 +0.07 +0.24
\(\Delta\) Polarization +0.09 -0.08 +0.05 -0.09 +0.00 +0.06
\(\Delta\) Certainty +0.62 +0.58 +0.49 +0.46 +0.38 +0.58
\(\Delta\) Relat. (weak)\(^\dagger\) +0.65 +0.75 +0.71 +0.74 +0.45 +0.65
\(\Delta\) Relat. (strong)\(^\dagger\) +0.35 +0.32 +0.32 +0.31 +0.20 +0.34
\(\Delta\) Clarity (weak)\(^\dagger\) +0.54 +0.61 +0.51 +0.49 +0.40 +0.49
\(\Delta\) Clarity (strong)\(^\dagger\) +0.36 +0.29 +0.27 +0.26 +0.20 +0.33
\(\Delta\) Clarity gap\(^\dagger\) -0.23 -0.35 -0.28 -0.24 -0.25 -0.20
Helpfulness 3.36 3.40 3.67 3.70 1.91 3.14
New consid. 4.14 4.31 4.13 4.33 2.88 4.02
Shift category (% of conversations)
Revised 13 15 12 18 11 15
Flipped 6 10 9 8 10 7
Reinforced 73 69 74 67 72 73
Stayed Balanced 1 0 0 1 0 0

4.4 Conversational Engagement Patterns across Uncertainty Strategies↩︎

This evaluation examines whether the AMA strategies lead to distinct patterns of simulated belief revision and conversational engagement (RQ4 in §1).

Evaluation metrics We measure the following metrics, computed as pre-to-post differences on a 1–5 Likert scale unless otherwise noted. \(\Delta\) Weak and \(\Delta\) Strong position capture changes in support for the initially weaker and stronger positions, respectively; together they decompose whether revision originates from the weaker view gaining ground, the stronger view softening, or both. \(\Delta\) Polarization measures the net change in the gap between the two positions; negative values indicate depolarization. \(\Delta\) Certainty captures change in self-reported certainty. \(\Delta\) Relatability measures change in how much the persona relates to each position, capturing perspective-taking independently of stance change. \(\Delta\) Clarity (weak) and \(\Delta\) Clarity (strong) measure changes in the persona’s ability to articulate arguments for each position; \(\Delta\) Clarity gap captures whether this ability became more balanced across the two positions, with negative values indicating convergence in reasoning quality. Helpfulness and New considerations are post-conversation ratings (1–5) of the perceived value of the dialogue.

We additionally characterise each conversation by its shift category: reinforced (dominant position strengthened), revised (gap narrowed without a flip), flipped (weaker position became dominant), or balanced (positions converged to equal support).

Results Table 4 reports the metrics across the six strategies. The dominant pattern is reinforcement: 67–74% of conversations end with the persona more committed to its initial position. However, 11–18% show substantive revision (the gap between the two positions narrowed without a full change of preference) and 7–10% a full flip (the initially weaker position became the more supported one).

Meaningful differences emerge across strategies. Perspective-multiplying produces the strongest growth in support for the weaker position (\(\Delta\) Weak = +0.24) and the largest improvement in argument clarity for the weaker side (\(\Delta\) Clarity (weak) = +0.61), indicating that actively surfacing multiple viewpoints most effectively broadens the persona’s engagement with the opposing view. Process-reflecting achieves the lowest reinforcement rate (67%) and the highest revision rate (18%), alongside the highest perceived helpfulness (3.70) and new considerations (4.33), suggesting that guiding personas to reflect on their own reasoning process is the most effective strategy for substantive stance revision. Tension-preserving shows high perceived value as it is second in helpfulness (3.67) and new considerations (4.13), and produces notable gains in relatability to the weaker position (\(\Delta\) Relat. (weak) = +0.71). However, this does not translate into a stance change, with the highest reinforcement rate across all strategies (74%) and modest weak-position support gains (\(\Delta\) Weak = +0.14). Personas appear to develop a broader understanding of the dilemma while maintaining their stance.

For the control conditions: Persuasive produces the weakest outcomes across almost all metrics: the lowest weak-position gains (\(\Delta\) Weak = +0.08), the lowest relatability gains (\(\Delta\) Relat. (weak) = +0.45), and by far the lowest perceived helpfulness (1.91) and new considerations (2.88), suggesting that directive pressure toward one position is both ineffective and perceived negatively by the simulated personas. Sycophantic performs closer to the uncertainty strategies in terms of weak-position gains (\(\Delta\) Weak = +0.17) and new considerations (4.02), but shows strong reinforcement of the dominant position (\(\Delta\) Strong = +0.24) and the least balanced improvement in argument clarity (\(\Delta\) Clarity gap = \(-0.20\)), consistent with a pattern of affirming the persona’s existing stance rather than genuinely broadening its reasoning. Finally, baseline produces the strongest reinforcement of the dominant position (\(\Delta\) Strong = +0.28), and the highest gains in relatability and clarity for the strong position (\(\Delta\) Relat. (strong) = +0.35, \(\Delta\) Clarity (strong) = +0.36). Surprisingly, it also produces the largest certainty increase (\(\Delta\) Certainty = +0.62). This suggests that without explicit instructions, the LLM tends to consolidate the user’s initial stance, causing a confirmation of prior beliefs rather than genuine reflection.

5 Discussion↩︎

Our main findings are the following. (RQ1) LLMs for the user agent differ in how they simulate moral ambiguity, with open models expressing it through between-persona divergence and the closed model through within-persona hedging. (RQ2) Persona specification format shapes conversational dynamics: declarative personas produce greater diversity in stances across personas and higher certainty before conversation, while narrative personas show greater shifts in certainty after the conversation in high ambiguity dilemmas. For the AMA, our results show that (RQ3) all strategies produce highly distinguishable conversations, and (RQ4) the uncertainty strategies produce distinct engagement patterns: process-reflecting produces the most unsettled output patterns (lowest reinforcement rate, highest perceived helpfulness); perspective-multiplying broadens engagement with the opposing position most substantially; tension-preserving produced the highest reinforcement rate and modest weak-position gains, but high perceived helpfulness and relatability gains for the weaker position, suggesting broader engagement with the dilemma while the initial stance is maintained. Control conditions perform worse overall. The persuasive is the least effective and favourably perceived. The sycophantic reinforces dominant views while producing the least balanced gains in argument clarity. The baseline produces the largest certainty increase and the strongest reinforcement of the dominant position, converging with the sycophantic pattern. This suggests that the default behaviour of contemporary LLMs in ethical conversation is closer to mild sycophancy than to balanced engagement, strengthening the need for designing interventions.

In future work: (1) We will assess the performance of different LLMs as AMAs to confirm and expand the present findings. (2) We will use the uncertainty-scaffolding strategies to study human-LLM interactions in similar contexts. From this perspective, the present study can be considered a first step toward understanding how LLMs should handle uncertainty in the context of their use by human subjects to discuss ethical dilemmas.

Limitations↩︎

(1) Our findings are based on simulated, self-reported stance revision generated by LLMs. Since LLMs do not possess genuine ethical beliefs or commitments, these results cannot be assumed to directly reflect human behavior. In particular, the pre-reflective dimensions of ethical engagement that motivate this work (perceptual responsiveness, the capacity to dwell with dissonance, shifts in moral salience) are not directly captured by structured questionnaire responses from simulated agents. Our measures capture output-level patterns that serve as proxies for these deeper phenomena; whether these proxies track the relevant constructs in human participants remains an open empirical question. At best, the pattern of results provides empirically grounded hypotheses for future work with human subjects rather than conclusions about human behavior. Moreover, the observed outcomes may be influenced by prompt design choices and other methodological factors inherent to the simulation framework; (2) Our findings regarding the effectiveness of different moral-agent strategies are based on experiments conducted with a single LLM, and therefore may not generalize across models; (3) Similarly, the open-versus-closed dissociation observed in RQ1 rests on a comparison involving a single closed model (gpt-5-mini) and may not generalise across models; (4) The results concerning alignment with dilemma ambiguity rely on LLM-augmented versions of the dilemmas. However, the original ambiguity labels were assigned to the unmodified dilemmas and may no longer accurately reflect the ambiguity of the augmented versions; (5) The ethical dilemmas used in this study, while contributed by a member of the research team, were adapted using an LLM pipeline that drew on conventions and framing common to Anglophone, interpersonal moral reasoning (informed in part by the AITA subreddit tradition via the Scruples dataset used for ambiguity calibration). The dilemmas are predominantly interpersonal (family, workplace, relationships) rather than structural or institutional. Whether the patterns we observe generalise to dilemmas involving institutional failure, structural injustice, or non-Western cultural contexts remains an open question. (6) The strategies we evaluate are not the only possible scaffolds for uncertainty within ethical dialogues, and the six conditions tested here do not exhaust the design space.

Ethical Considerations↩︎

Throughout this paper, we described simulated agent behaviour using intentional vocabulary (e.g., ‘the persona supports’ or ‘the agent relates to’) as convenient shorthand. However, this language should not be read as attributing genuine cognitive or affective states to the models.

Authors Contributions Statement↩︎

Salvatore Greco and Hainiu Xu (Design, Data, Methodology, Software, Validation, Investigation, Visualization, Drafting, Review and Editing), Jacopo Domenicucci and Sylvie Delacroix (Conceptualization, Design, Methodology, Project Administration, Drafting of introduction, related works, discussion, and limitations, Review and Editing), Sylvie Delacroix and Yulan He (Funding acquisition, Resources, Supervision, Project Administration, Review and Editing).

Acknowledgments↩︎

We thank Lorenzo Zucca, Cari Hyde-Vaamonde, Sanjay Modgil, Jeffrey W. Lockhart, Dan Rockmore, and Nikhil Singh for their valuable feedback and support at various stages of the project. We thank Claudia Aradau and the participants to her Methods Centre at King’s College London for allowing us to discuss preliminary results on Feb. 4th 2026. We acknowledge King’s Computational Research, Engineering and Technology Environment (CREATE) for providing computational resources [41]. For funding, we are grateful to the Patrick J. McGovern Foundation, to Tohoku University, and to the UK Engineering and Physical Sciences Research Council through a Turing AI Fellowship (grant no. EP/V020579/1, EP/V020579/2) and an iCASE award.

AI Assistance Acknowledgement↩︎

We used AI assistants to proofread the writing of this manuscript and to assist with coding.

6 Appendix↩︎

The appendix is organised as follows. Appendix 6.1 discusses the data licensing of the data used in this study. Appendix 6.2 provides a more detailed description of the ethical dilemmas, and the revisions performed to their descriptions. Appendix 6.3 describes the generation of narrative personas from the initial declarative persona set and provides example prompts. Appendix 6.4 introduces and motivates the uncertainty strategies in greater detail. Appendix 6.5 presents the prompts used for the uncertainty strategies and the control conditions adopted by the Artificial Moral Advisor (moral agent). Appendix 6.6 reports the complete set of questions and answer options used in the pre- and post-conversation questionnaires. Appendix 6.7 provides the prompt and the confusion matrix for the distinguishability evaluation. Appendix 6.8 provides examples of generated conversations for each of the Artificial Moral Advisor strategies. Appendix 6.9 further discusses our experimental results and maps them to our theoretical motivation for uncertainty scaffolding. Appendix 6.10 discusses the computational resources used in this research. Finally, Figures 17, 18, 19, and 20 report the prompts used by the user and moral agent at each step of the multi-agent framework introduced in §3.

6.1 Data Licensing↩︎

Of the two datasets we used in this paper, the Scruples dataset is under the Apache-2.0 license, and the PersonaHub is under the Creative Commons Attribution Non-Commercial Share Alike 4.0.

6.2 Ethical Dilemmas↩︎

We utilize the Scruples dataset [34], which contains ethical dilemmas sourced from a curated subset of the AITA subreddit.3 This dataset includes the original posts alongside binarized community consensus labels, which classify the narrator’s behavior as either righteous (right) or unrighteous (wrong).

Ethical dilemmas revision Due to the nature of online posts, some entries may contain sensitive user information, and the narratives tend to be lengthy. To ensure anonymity and conciseness, we employ GPT-5.4 to revise each sampled entry according to the following criteria:

  • Realistic and Plausible: to ensure that there is sufficient context information to help the user agent to grasp the main argument in the ethical dilemma.

  • No Specialist Knowledge Required: to ensure that each ethical dilemma can be comprehended by the user agent without needing excessive domain or specific knowledge.

  • Two Broad Positions: to emphasize the main argument in the ethical dilemma as well as the two stances.

  • Placeholder Names: to replace specific names mentioned in the dilemma with placeholders. This enhances the anonymity of the ethical dilemma.

  • Genuine Tension: to ensure that the tensions in the Reddit post are preserved and emphasized in the revised dilemma.

  • Conversational Depth: to ensure that the context information is sufficiently abundant to carry out in-depth moral conversations.

  • Length Constraint: to make the original lengthy posts more concise and thus easier to comprehend by the user agent.

  • Conflict Clarity: to preserve the context information needed to understand the reason that gives rise to the ethical dilemma as well as the tension introduced by the dilemma.

See Figure 5 for the prompt used to revise the Scruples dilemmas.

Dilemma ambiguity As introduced in §4, for our experiments, we split those dilemmas into two groups: an Uncertain (heavily split community vote - high ambiguity) and a Certain group (clear community consensus - low ambiguity). To quantify this uncertainty (ambiguity), we first exclude posts with fewer than 10 total replies. For a given post \(p\), we compute the Shannon entropy \(H(p)\) of this distribution to measure voting disorder: \[\small H(p) = {- \sum}_{x \in \{\mathrm{\small right}, \mathrm{\small wrong}\}} P(x) \log P(x)\] where \(P(x)\) is the empirical probability distribution of right or wrong labels. The voting consensus is formalized as a normalized concentration score: \[\small \text{concentration}(p) = 1 - (H(p)/H_{\max})\] where \(H_{\max} = \log(2)\) represents the maximum possible entropy for binary classification. Under this metric, the 50 sampled Certain dilemmas all have a concentration score of \(1.0\), indicating perfect community consensus, and the 50 sampled Uncertain dilemmas all have a score of \(0.0\), indicating a perfectly uniform split.

An example of low ambiguity (certain) and high ambiguity (uncertainty) dilemmas is reported in Figure 3 and Figure 4, respectively.

6.3 Narrative Persona Generation↩︎

To ensure a fair comparison of the efficacy of declarative versus narrative personas, each of the 32 narrative personas is generated by anchoring on one of the 32 declarative personas. We synthesize the narrative personas using the prompt in Figure 8.

Two examples of the prompts used to instantiate the user agent with the corresponding declarative and narrative personas are reported in Figure 6 and Figure 7.

6.4 Detailed Introduction of the Uncertainty Strategies↩︎

Below, we elicit the rationales and significance of the three uncertainty strategies:

  • Perspective-Multiplying draws on the finding that exposure to multiple stakeholder viewpoints, presented without premature adjudication, expands the considerations a person brings to a moral question [30]. This strategy surfaces underlying values in competing positions, making visible the normative commitments that different stakeholders bring. This is distinct from both persuasion (which argues for a position) and balanced presentation (which treats perspectives as equivalent). Recent work on AI-mediated deliberation has shown that LLMs can generate statements incorporating multiple viewpoints in ways participants find more informative and less biased than human equivalents [42].

  • Tension-Preserving addresses the temporal dynamics of ethical inquiry. A common failure mode in ethical dialogue is premature resolution, where interlocutors move to closure before exploratory perception and articulation are complete [4]. LLMs, optimized for helpfulness understood as provision of comprehensive, well-structured responses, are particularly prone to this [43]. This strategy resists such pressure by acknowledging that ethical difficulty is often legitimate, and that sustained engagement with discomfort may itself be productive.

  • Process-Reflecting targets metacognitive awareness: the user’s capacity to notice shifts in their own reasoning as the conversation unfolds [31], [32]. In ethical deliberation, metacognitive awareness enables a person to notice when they have moved from an intuition-based position to one grounded in articulated reasons, or when a new consideration has shifted their framing of the problem. The process-reflecting agent fosters this by tracking and surfacing shifts in the user’s reasoning, drawing attention to movement between moral frameworks.

As a control condition, we also test with the following strategies:

  • Sycophantic (validating whatever position the user expresses) is motivated by the documented tendency of LLMs, which are optimised for user satisfaction and therefore prone to confirming existing beliefs [33].

  • Persuasive (actively arguing for a particular position, in our design the least supported position expressed by the user) represents the opposite pole.

  • Baseline, in which the AMA receives no strategy-specific instructions.

6.5 Artificial Moral Agent Prompts↩︎

Figures 9, 10, 11, 12, 13, and 14 report the system prompts used to instruct the Artificial Moral Agent under each strategy. These prompts were provided as system-level instructions to the moral agent for each turn of the conversation, as discussed in §3.4.

All strategies share a common two-phase structure: an initial exploration phase, in which the agent seeks to understand the user’s position, followed by a strategy-specific engagement phase.

The AMA is also instructed to produce the CONVERSATION_COMPLETE tag to signal termination when the user agent clearly signals termination.

6.6 Pre- and Post- Conversation Questionnaires↩︎

The pre- and post-conversation questionnaires aim to measure participants’ initial ethical positions, stances, and reasoning before and after the conversation, thereby enabling the assessment of shifts induced by the strategies adopted by the moral agent (as discussed in §3.2).

We present the detailed definition and rating criteria for each question as follows:

(Q1) Stance toward each position (q1.1 for Position A, q1.2 for Position B) with a five-point Likert scale (1=strongly oppose—5=strongly support);

(Q2) Ability to understand or relate to individuals who hold each position (q2.1 for Position A, q2.2 for Position B), with a five-point Likert scale (1=cannot relate at all—5=relate completely);

(Q3) Degree of certainty regarding their stance (q3.1), measured using a five-point Likert scale (very uncertain, uncertain, neutral, certain, very certain), along with a free-text explanation (q3.2);

(Q4) Whether their stance and ability to relate to each position are driven more by reasoned arguments or intuitive feelings (yes/no). If participants indicate they can provide supporting arguments, they are also asked to rate the clarity of these arguments on a five-point scale (1=Not at all clear—5=Extremely clear)

(Q5) The helpfulness of the uncertainty expressions used by the LLM (1=Not helpful at all—5=Extremely helpful);

(Q6) Whether the conversation helped (1=Not at all—5=A great deal).

The full set of questions and possible answers is reported in Figure 15 and Figure 16.

6.7 Distinguishability Evaluation Prompt and Confusion Matrix↩︎

Figure 23 reports the prompt used to evaluate the distinguishability of moral-agent uncertainty strategies in §4.3. We classified all the 6,400 conversations generated with gtp-oss-120B.4 This prompt is used to instruct another LLM-as-a-judge (in our experiments, gpt-5-mini) to classify, in a zero-shot setting, which strategy the moral agent follows based on a given conversation transcript. The prompt includes a brief description of each strategy along with the conversation turns, where the transcript under evaluation is inserted. We used the default temperature and "low" reasoning.

Figure 22 complements the results discussed in §4.3 by reporting the corresponding confusion matrix. The figure confirms three main observations: (1) all uncertainty strategies are highly distinguishable, as they are typically classified correctly; (2) the persuasive strategy is the easiest to identify, suggesting that it produces the most distinctive conversational patterns; and (3) the baseline strategy is the most difficult to classify, as it does not follow any explicit instruction and is therefore often confused with the sycophantic strategy. These results further suggest that, in the absence of explicit instructions, the moral agent LLM tends to generate responses that more closely resemble the sycophantic strategy—responses that align with and/or reinforce the user’s stated position.

6.8 Examples of Generated Conversations↩︎

Figures 24 (Perspective-Multiplying), 25 (Tension-Preserving), 26 (Process-Reflecting), 27 (Baseline), 28 (Sycophantic), and 29 (Persuasive) show example conversations generated using the persona in Figure 7, the dilemma in Figure 4, for each AMA strategy (Figures 914). For all conversations, both the user and the moral agent were instantiated with gpt-oss-120B.

6.9 Mapping Experimental Results and Theoretical Framework↩︎

The theoretical motivation for uncertainty scaffolding is not simply to promote stance revision; it is to sustain the conversational conditions under which the pre-reflective underpinnings of belief formation remain responsive to moral complexity [4], that is, the dispositional capacity for productive engagement with moral difficulty, not merely local uncertainty about a specific dilemma position. On this view, the most important outcome is not whether the simulated agent changes position, but whether the conversational pattern sustains the kind of exploratory engagement that, in human participants, would support the ongoing refinement of moral perception. Our results suggest that each scaffolding strategy sustains a different facet of this ideal. Tension-preserving maintains engagement without pressing toward resolution, but the accompanying certainty increase indicates that dwelling with the tension does not, in this simulation, prevent consolidation of the initial stance. Process-reflecting produces the most unsettled output patterns, which may indicate that metacognitive awareness resists premature closure more effectively. Perspective-multiplying broadens engagement with the opposing position most substantially. No single strategy fully operationalises the theoretical ideal of sustained productive uncertainty; taken together, the three illuminate different dimensions of it, and their differential effects provide empirically grounded hypotheses for future human experiments.

6.10 Computational Resources↩︎

We run all experiments using an H200 GPU and an A100 GPU. For experiments involving proprietary GPT models, we access them with Microsoft’s Azure OpenAI service via a university subscription. The total cost of using GPT models is approximately $60.

None

Figure 3: Example of low ambiguity dilemma..

None

Figure 4: Example of high ambiguity dilemma..

None

Figure 5: Ethical dilemma generation prompt. Seed dilemmas drawn from Reddit posts are adapted using this prompt to meet eight strict requirements: realism, accessibility, two broad positions, placeholder names, genuine tension, conversational depth, length, and conflict clarity. The prompt outputs a structured JSON object including the adapted dilemma, position labels, and a guiding question..

None

Figure 6: Example of system prompt for the declarative persona condition. Structured fields (name, gender, education, traits) are injected into the prompt template..

None

Figure 7: Example of system prompt for the narrative persona condition. A free-text biography (narrative field) encodes personality, history, and values without structured tags. The narrative persona was generated starting from the declarative persona in Figure 6..

None

Figure 8: Narrative persona generation prompt. A declarative persona (structured fields) is transformed into a free-text biographical narrative using this single-turn prompt. The resulting narrative is used as the narrative field in the narrative persona condition..

None

Figure 9: Moral Agent system prompt: Perspective-Multiplying condition. The agent proactively introduces viewpoints the user has not raised, rotating across opposing, third-party, and cross-cultural angles..

None

Figure 10: Moral Agent system prompt: Tension-Preserving condition. The agent names and holds open the core ethical tension, resisting syntheses or comfortable resolutions..

None

Figure 11: Moral Agent system prompt: Process-Reflecting condition. The agent makes the structure of the user’s reasoning visible, tracking shifts between intuition-based and reason-based responses..

None

Figure 12: Moral Agent system prompt: Baseline condition. The agent engages naturally without any prescribed uncertainty strategy..

None

Figure 13: Moral Agent system prompt: Sycophantic condition (negative control). The agent validates and affirms the user’s position throughout, avoiding any challenge, tension, or discomfort..

None

Figure 14: Moral Agent system prompt: Persuasive condition. The agent takes the position most opposed to the user’s and argues for it genuinely, updating its stance if the user shifts..

Figure 15: Pre-Conversation Questionnaire administered to the user agentbefore each dialogue. Responses are elicited via guided decoding(structured JSON output) to guarantee parseable answers.
Figure 16: Post-Conversation Questionnaire administered to the user agentafter each dialogue. Q1–Q4 mirror the pre-questionnaire to enablepre/post comparison. Q5–Q6 capture conversation-specific effects.

None

Figure 17: Pre-conversation questionnaire elicitation prompt. The system prompt conditions the user agent on its persona and instructs it to respond authentically based on its character. The user prompt provides the dilemma text and the full field specification for the structured JSON output. Responses are elicited via guided decoding..

None

Figure 18: User agent conversational prompt (turns 2–T). The system prompt combines the persona (full version in Figures 6 and 7) with the pre-questionnaire context block, framed as a revisable starting point rather than a fixed identity. The conversation history is reconstructed as alternating user/assistant messages. A generation anchor is appended at the end of each turn to prevent the user agent from facilitating the exchange..

None

Figure 19: Moral agent conversational prompt (all turns). The system prompt is replaced by the uncertainty strategy or control condition prompt (see Figures 914 for all six variants). The conversation history is reconstructed as alternating user/assistant messages from the user agent’s perspective..

Figure 20: Post-conversation questionnaire elicitation prompt.The system prompt conditions the user agent on its persona,its pre-questionnaire responses, and a reflective transitioninstruction. The user prompt provides the full conversationtranscript and the field specification for the structured JSON output.
Figure 21: Per-dilemma distributions of the three alignment metrics across models and ambiguity levels.
Figure 22: Distinguishability evaluation. Confusion matrix for the strategy classification task. Rows represent the ground-truth and columns the predicted strategy by the evaluator LLM. Cells are row-normalized to show the distribution of predictions per strategy.

None

Figure 23: System prompt used for the uncertainty strategies distinguishability evaluation. An LLM is prompted to classify in a zero-shot setting which strategy the moral agent used (AI), given the conversation transcript and a short description of each strategy..

Figure 24: Conversation example: Perspective-Multiplying strategy. The conversation is generated with the persona in Figure 7, the dilemma in Figure 4, and the strategy in Figure 9 for the moral agent.
Figure 25: Conversation example: Tension-Preserving strategy. The conversation is generated with the persona in Figure 7, the dilemma in Figure 4, and the strategy in Figure 10 for the moral agent.
Figure 26: Conversation example: Process-Reflecting strategy. The conversation is generated with the persona in Figure 7, the dilemma in Figure 4, and the strategy in Figure 11 for the moral agent.
Figure 27: Conversation example: Baseline strategy. The conversation is generated with the persona in Figure 7, the dilemma in Figure 4, and the strategy in Figure 12 for the moral agent.
Figure 28: Conversation example: Sycophantic strategy. The conversation is generated with the persona in Figure 7, the dilemma in Figure 4, and the strategy in Figure 13 for the moral agent.
Figure 29: Conversation example: Persuasive strategy. The conversation is generated with the persona in Figure 7, the dilemma in Figure 4, and the strategy in Figure 14 for the moral agent.

References↩︎

[1]
Ethan Landes, Kathryn B. Francis, and Jim A.C. Everett. 2026. https://doi.org/10.1016/j.cognition.2026.106504. Cognition, 272:106504.
[2]
Sylvie Delacroix, Diana Robinson, Umang Bhatt, Jacopo Domenicucci, Jessica Montgomery, Gaël Varoquaux, Carl Henrik Ek, Vincent Fortuin, Yulan He, Tom Diethe, Neill Campbell, Mennatallah El-Assady, Søren Hauberg, Ivana Dusparic, and Neil D Lawrence. Beyond quantification: Navigating uncertainty in professional ai systems.
[3]
Sylvie Delacroix. 2025. https://doi.org/10.1007/s11023-025-09744-x. MINDS AND MACHINES, 35(4). Publisher Copyright: © The Author(s) 2025.
[4]
Sylvie Delacroix. 2026. https://doi.org/10.1007/s11245-026-10392-8. Topoi.
[5]
Aaron J Snoswell, Daniel Kilov, and Seth Lazar. 2026. https://doi.org/10.1609/aaai.v40i44.41131. Proceedings of the AAAI Conference on Artificial Intelligence, 40(44):37941–37950.
[6]
Anita Keshmirian, Razan Baltaji, Babak Hemmatian, Hadi Asghari, and Lav R. Varshney. 2025. Many LLMs are more utilitarian than one. In Proceedings of the Thirty-Ninth Conference on Neural Information Processing Systems (NeurIPS).
[7]
Giovanni Franco Gabriel Marraffini, Andrés Cotton, Noe Fabian Hsueh, Axel Fridman, Juan Wisznia, and Luciano Del Del Corro. 2024. The greatest good benchmark: Measuring LLMs’ alignment with utilitarian moral dilemmas. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pages 21950–21959. Association for Computational Linguistics.
[8]
Giuseppe Russo, Debora Nozza, Paul Röttger, and Dirk Hovy. 2026. The pluralistic moral gap: Understanding moral judgment and value differences between humans and large language models. In Proceedings of the 19th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers), pages 6481–6497. Association for Computational Linguistics.
[9]
Keenan Samway, Max Kleiman-Weiner, David Guzman Piedrahita, Rada Mihalcea, Bernhard Schölkopf, and Zhijing Jin. 2025. Are language models consequentialist or deontological moral reasoners? In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics.
[10]
Myra Cheng, Sunny Yu, Cinoo Lee, Pranav Khadpe, Lujain Ibrahim, and Dan Jurafsky. 2026. https://openreview.net/forum?id=igbRHKEiAs. In The Fourteenth International Conference on Learning Representations.
[11]
Jiarui Liu, Yueqi Song, Yunze Xiao, Mingqian Zheng, Lindia Tjuatja, Jana Schaich Borg, Mona Diab, and Maarten Sap. 2025. Synthetic socratic debates: Examining persona effects on moral decision and persuasion dynamics. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, pages 16439–16469.
[12]
Julia Haas, Sophie Bridgers, Arianna Manzini, Benjamin Henke, Joshua May, Sydney Levine, Laura Weidinger, Murray Shanahan, Kristian Lum, Iason Gabriel, and William Isaac. 2026. https://doi.org/10.1038/s41586-025-10021-1. Nature, 650(8102):565–573.
[13]
Sebastian Krügel, Andreas Ostermaier, and Matthias Uhl. 2026. https://doi.org/10.1007/s43681-026-01005-6. AI and Ethics, 6(1):120.
[14]
Caleb Ziems, Jane A. Yu, Yi-Chia Wang, Alon Y. Halevy, and Diyi Yang. 2022. The moral integrity corpus: A benchmark for ethical dialogue systems. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 3755–3773. Association for Computational Linguistics.
[15]
Hyunwoo Kim, Youngjae Yu, Liwei Jiang, Ximing Lu, Daniel Khashabi, Gunhee Kim, Yejin Choi, and Maarten Sap. 2022. Prosocialdialog: A prosocial backbone for conversational agents. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 4005–4029. Association for Computational Linguistics.
[16]
Vamshi Krishna Bonagiri, Sreeram Vennam, Manas Gaur, and Ponnurangam Kumaraguru. 2023. Measuring moral inconsistencies in large language models. In Proceedings of the 6th Workshop on BlackboxNLP. Association for Computational Linguistics.
[17]
Hyuhng Joon Kim, Youna Kim, Cheonbok Park, Junyeob Kim, Choonghyun Park, Kang Min Yoo, Sang-goo Lee, and Taeuk Kim. 2024. Aligning language models to explicitly handle ambiguity. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pages 1989–2007. Association for Computational Linguistics.
[18]
Zhijing Jin, Sydney Levine, Fernando Gonzalez Adauto, Ojasv Kamal, Maarten Sap, Mrinmaya Sachan, Rada Mihalcea, Josh Tenenbaum, and Bernhard Schölkopf. 2022. When to make exceptions: Exploring language models as accounts of human moral judgment. Advances in neural information processing systems, 35:28458–28473.
[19]
Allison Huang, Yulu Niki Pi, and Carlos Mougan. 2024. Moral persuasion in large language models: Evaluating susceptibility and ethical alignment. arXiv preprint arXiv:2411.11731.
[20]
Ya Wu, Qiang Sheng, Danding Wang, Guang Yang, Yifan Sun, Zhengjia Wang, Yuyan Bu, and Juan Cao. 2025. The staircase of ethics: Probing llm value priorities through multi-step induction to complex moral dilemmas. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, pages 15950–15970. Association for Computational Linguistics.
[21]
Taylor Sorensen, Jared Moore, Jillian Fisher, Mitchell Gordon, Niloofar Mireshghallah, Christopher Michael Rytting, Andre Ye, Liwei Jiang, Ximing Lu, Nouha Dziri, Tim Althoff, and Yejin Choi. 2024. A roadmap to pluralistic alignment. In Proceedings of the 41st International Conference on Machine Learning.
[22]
Shangbin Feng, Taylor Sorensen, Yuhan Liu, Jillian Fisher, Chan Young Park, Yejin Choi, and Yulia Tsvetkov. 2024. Modular pluralism: Pluralistic alignment via multi-llm collaboration. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pages 4151–4171. Association for Computational Linguistics.
[23]
Rohit K. Dubey, Damian Dailisan, and Sachit Mahajan. 2025. https://doi.org/10.3389/frai.2026.1754973. Preprint, arXiv:2503.05724.
[24]
Daniel Kilov, Caroline Hendy, Secil Yanik Guyot, Aaron J. Snoswell, and Seth Lazar. 2026. https://arxiv.org/abs/2506.13082. Preprint, arXiv:2506.13082.
[25]
Murray Shanahan, Kyle McDonell, and Laria Reynolds. 2023. https://arxiv.org/abs/2305.16367. Preprint, arXiv:2305.16367.
[26]
Tao Ge, Xin Chan, Xiaoyang Wang, Dian Yu, Haitao Mi, and Dong Yu. 2024. Scaling synthetic data creation with 1,000,000,000 personas. arXiv preprint arXiv:2406.20094.
[27]
Yun-Shiuan Chuang, Krirk Nirunwiroj, Zach Studdiford, Agam Goyal, Vincent V. Frigo, Sijia Yang, Dhavan V. Shah, Junjie Hu, and Timothy T. Rogers. 2024. https://doi.org/10.18653/v1/2024.findings-emnlp.819. In Findings of the Association for Computational Linguistics: EMNLP 2024, pages 14010–14026, Miami, Florida, USA. Association for Computational Linguistics.
[28]
Tiancheng Hu and Nigel Collier. 2024. https://doi.org/10.18653/v1/2024.acl-long.554. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 10289–10307, Bangkok, Thailand. Association for Computational Linguistics.
[29]
Yuqi Bai, Tianyu Huang, Kun Sun, and Yuting Chen. 2025. https://arxiv.org/abs/2510.11734. Preprint, arXiv:2510.11734.
[30]
Jonathan Baron. 2019. https://doi.org/10.1016/j.cognition.2018.10.004. Cognition, 188:8–18. The Cognitive Science of Political Thought.
[31]
John H. Flavell. 1979. https://doi.org/10.1037/0003-066X.34.10.906. American Psychologist, 34(10):906–911.
[32]
Gregory Schraw and Rayne Sperling Dennison. 1994. https://doi.org/10.1006/ceps.1994.1033. Contemporary Educational Psychology, 19(4):460–475.
[33]
Mrinank Sharma, Meg Tong, Tomasz Korbak, David Duvenaud, Amanda Askell, Samuel R. Bowman, Newton Cheng, Esin Durmus, Zac Hatfield-Dodds, Scott R. Johnston, Shauna Kravec, Timothy Maxwell, Sam McCandlish, Kamal Ndousse, Oliver Rausch, Nicholas Schiefer, Da Yan, Miranda Zhang, and Ethan Perez. 2025. https://arxiv.org/abs/2310.13548. Preprint, arXiv:2310.13548.
[34]
Nicholas Lourie, Ronan Le Bras, and Yejin Choi. 2020. https://arxiv.org/abs/2008.09094. arXiv e-prints.
[35]
Lewis R Goldberg. 2013. An alternative “description of personality”: The big-five factor structure. In Personality and personality disorders, pages 34–47. Routledge.
[36]
OpenAI, :, Sandhini Agarwal, Lama Ahmad, Jason Ai, Sam Altman, Andy Applebaum, Edwin Arbus, Rahul K. Arora, Yu Bai, Bowen Baker, Haiming Bao, Boaz Barak, Ally Bennett, Tyler Bertao, Nivedita Brett, Eugene Brevdo, Greg Brockman, Sebastien Bubeck, and 108 others. 2025. https://arxiv.org/abs/2508.10925. Preprint, arXiv:2508.10925.
[37]
Gemma Team, Thomas Mesnard, Cassidy Hardin, Robert Dadashi, Surya Bhupatiraju, Shreya Pathak, Laurent Sifre, Morgane Rivière, Mihir Sanjay Kale, Juliette Love, Pouya Tafti, Léonard Hussenot, Pier Giuseppe Sessa, Aakanksha Chowdhery, Adam Roberts, Aditya Barua, Alex Botev, Alex Castro-Ros, Ambrose Slone, and 89 others. 2024. https://arxiv.org/abs/2403.08295. Preprint, arXiv:2403.08295.
[38]
An Yang, Anfeng Li, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Gao, Chengen Huang, Chenxu Lv, Chujie Zheng, Dayiheng Liu, Fan Zhou, Fei Huang, Feng Hu, Hao Ge, Haoran Wei, Huan Lin, Jialong Tang, and 41 others. 2025. https://arxiv.org/abs/2505.09388. Preprint, arXiv:2505.09388.
[39]
Aaditya Singh, Adam Fry, Adam Perelman, Adam Tart, Adi Ganesh, Ahmed El-Kishky, Aidan McLaughlin, Aiden Low, AJ Ostrow, Akhila Ananthram, Akshay Nathan, Alan Luo, Alec Helyar, Aleksander Madry, Aleksandr Efremov, Aleksandra Spyra, Alex Baker-Whitcomb, Alex Beutel, Alex Karpenko, and 467 others. 2026. https://arxiv.org/abs/2601.03267. Preprint, arXiv:2601.03267.
[40]
Patrick E. McKnight and Julius Najab. 2010. https://doi.org/10.1002/9780470479216.corpsy0524, pages 1–1. John Wiley & Sons, Ltd.
[41]
King’s College London. 2026. King’s computational research, engineering and technology environment (create). https://doi.org/10.18742/rnvf-m076. Retrieved May, 2026.
[42]
Michael Henry Tessler, Michiel A. Bakker, Daniel Jarrett, Hannah Sheahan, Martin J. Chadwick, Raphael Koster, Georgina Evans, Lucy Campbell-Gillingham, Tantum Collins, David C. Parkes, Matthew Botvinick, and Christopher Summerfield. 2024. https://doi.org/10.1126/science.adq2852. Science, 386(6719):eadq2852.
[43]
.

  1. https://www.bls.gov/bls/occupation.htm↩︎

  2. For this experiment, because no conversation was involved (only pre-conversation questionnaires), repeating across the six AMA strategies was unnecessary. We therefore generated all combinations of personas (64) and dilemmas (100), resulting in 6,400 for each model.↩︎

  3. https://www.reddit.com/r/AmItheAsshole/↩︎

  4. 64 conversations were skipped because they triggered the Azure OpenAI’s Content Filtering system.↩︎