Practicing with language models cultivates human empathic communication


Introduction↩︎

Empathy underpins human social life, shaping relationships, cooperation, and well-being. Yet communicating in a manner that makes another person feel heard can be challenging in practice [1][4]. Empathy relies on observational learning [5] and empathic communication is a learnable skill [6][12] that shows substantial individual variability [13][15]. In blinded evaluations, large language model (LLM) responses to people expressing troubles are judged as more empathetic than the average human-written ones [3], [16][20]. Nevertheless, most people report feeling significantly less heard and supported after they learn that an empathic message comes from an AI [4], [21]. In light of the importance of human presence in empathic communication [22], [23], LLMs’ superior skill relative to the average person [4], [18], and growing evidence of LLMs’ effectiveness as coaches and tutors [3], [24], [25], a natural question emerges: How can people learn from LLMs to respond to other’s troubles in a way that makes them feel heard and understood?

The costs of failed empathic communication are profound. In society, a lack of human connection is linked to rising loneliness, polarization, and decline in well-being [26][28]. At work, breakdowns in empathy undermine collaboration, leadership, and trust [29], [30]. As remote work and hybrid human-AI teams reduce the informal interactions that once built mutual understanding [31], [32], the capacity to make others feel heard becomes harder to practice and more essential to sustain. AI chatbots do not match human connection when it comes to fulfilling psychological and social needs [23], [33][35], and heavy reliance on them for social support may foster dependence and worsen well-being [36]. In contrast, everyday empathic exchanges with other people are associated with increased well-being [37], [38]. Likewise, learning about the personal narratives of others can increase connection with highly stigmatized groups [39]. For these reasons, empathic communication as a skill is worthy of cultivation in humans and should not be outsourced to AI.

Empathic communication is not a single behavior but a constellation of communicative components. Research on empathic communication shows that responses which make people feel heard encourage elaboration [40][42], validate emotions [43], and demonstrate understanding [44][47]; counterproductive ones offer unsolicited advice [48], [49], shift focus away from the speaker’s experience [49], [50], or dismiss emotions [51], [52] (See Fig. 1A for examples of these responses). Empathic communication is highly context-dependent, often defying simplistic rubrics and making structured training a challenge. For instance, subtle phrasing can signal validation in one context but come across as patronizing in another, complicating efforts to codify it. This complexity makes empathic communication difficult to measure and teach. While traditional interventions for empathic communication training have been shown to be effective [53][55], they are resource-intensive and hence limited in reach. Brief, scalable empathic-mindset interventions have shown to shift outcomes in field settings [56], but these target empathic disposition rather than the communicative idiom through which empathy is expressed.

LLMs, with their demonstrated capabilities to generate [4], [21], [57] and evaluate empathic communication in text [15], offer a way forward. LLM-powered systems can simulate realistic practice partners, deliver personalized feedback, serve as reliable evaluators, and scale to reach learners who would otherwise lack access to structured training [58]. This approach has started to show promise in coaching people across a range of interpersonal skills including conflict resolution in personal relationships [59], professional communication [60], negotiation [61], [62], democratic deliberation [63], and counseling [64]. These findings offer a blueprint for a scalable way to both measure and cultivate empathic communication using AI tools. Whether brief LLM-based interventions can cultivate empathic communication remains an open empirical question.

In the Lend an Ear experiment, we ask whether LLMs can be used to help people practice and improve their ability to communicate empathically. Lend an Ear is designed as an interactive role-playing game where people practice offering empathic support to an AI role-playing partner. In each conversation, an AI partner simulates someone experiencing either a personal trouble (a family member diagnosed with cancer in one and passing away in another) or a workplace trouble (losing a job, getting passed over for a promotion, and feeling undervalued at work). Participants role-play as supporters in three conversations, offering responses across multiple conversational turns, and receive personalized feedback from an AI coach or through short videos. In a preregistered randomized experiment with 968 participants, producing 2904 conversations with an average of 11 turns per conversation, and a total of 16,975 human messages and 16,963 LLM-generated messages, we evaluate the impact of personalized feedback from an LLM coach on participants’ empathic communication performance.

Our dataset of 16,975 human messages enables quantitative and qualitative analysis of how people express empathy in naturalistic conversation and when their attempts at offering empathic support align (or not) with established frameworks for empathic communication. Surprisingly, our results reveal a disconnect between self-reported empathy, felt empathy, and expressed empathy, suggesting that feeling empathy and communicating empathically are distinct. We find that a brief coaching intervention, powered by LLM feedback, significantly improves expressed empathy across six preregistered dimensions of prescriptive (encouraging elaboration, validating emotions, demonstrating understanding) and proscriptive (giving unsolicited advice, reorienting the conversation to oneself, dismissing emotions) communication behaviors. Finally, in a follow-up human preference experiment, we find that these behavioral shifts correspond to what independent raters perceive as more empathic.

These findings advance the science of empathic communication by demonstrating that targeted practice with LLM partners can improve performance in controlled settings, by providing a data-driven taxonomy of empathic message contents grounded in naturalistic dialogue, and by showcasing the silent empathy effect that trait empathy is unrelated to empathic communication skill. Our approach demonstrates a scalable method for strengthening empathic skills at a time when human connection is both deeply needed and increasingly fragile. More broadly, this work offers guidance for reframing empathic communication from an intangible “soft” skill into a “hard” skill that can be quantified, trained, and strengthened.

a

b

c

Figure 1: Overview of the Lend an Ear experiment. A. Examples of empathic responses by participants to a conversational partner’s disclosure of job loss on the Lend an Ear platform that tend to make people feel heard (left panel) and that tend to not be effective at making people feel heard (right panel). B. User interface of the chat window with an LLM conversational partner (left) and personalized feedback from the AI coach (right). C. Experimental design flowchart illustrating participant recruitment, random assignment to four conditions (Control, Video, AI Coach, Combined), and the sequence of surveys, conversations, and feedback..

Lend an Ear Platform↩︎

We designed a custom, interactive web platform called Lend an Ear to evaluate whether practicing empathic communication with AI conversational partners and receiving personalized LLM-generated feedback can improve participants’ empathic communication skills. This role-playing setup simulates realistic interpersonal scenarios in which a conversational partner seeks support, allowing participants to practice offering empathic responses across multiple conversational turns. LLM conversational partners simulated five distinct scenarios spanning workplace troubles (losing a job, getting passed over for a promotion, and feeling undervalued at work) and personal troubles (a family member diagnosed with cancer in one and passing away in another). For each scenario, the LLM role-playing agent was provided with a detailed background story establishing their identity and the specific trouble they were experiencing. These LLM conversational partners were instructed to behave as individuals seeking to feel heard and understood, expressing their concerns and emotions in two to three sentences per turn across multi-turn conversations. At the end of the experiment, we asked participants to rate their agreement with the statement, “The troubles that my conversational partners described seemed realistic”, on a five-point Likert scale (1 = not at all, 2 = slightly, 3 = somewhat, 4 = quite a bit, 5 = very much). 91% of participants rated the scenarios as “quite a bit” or “very much” realistic. An LLM communication coach prompted using a comprehensive framework of empathic communication principles provided automated, personalized feedback on empathic communication skills. See Methods for details. Fig. 1B shows screenshots of the chat interface and the coach feedback window.

We conducted a preregistered experiment where participants were randomly assigned to one of four conditions: (1) a control condition with no feedback, (2) two short instructive videos (57 and 35 seconds) featuring a human communication coach, (3) an interactive AI coaching system with access to participants’ conversations and availability for follow-up questions, and (4) a combination of the AI coaching system and the human coach videos (see Preregistration in Materials and Methods for more details). Fig. 1C illustrates the experimental flow. We recruited 968 participants via Prolific, targeting a demographically representative U.S. sample, resulting in 2,904 conversations and 33,938 messages. Participants were randomly assigned to one of the four conditions. They first reviewed instructions outlining the procedure and their role in the conversations, and then completed a baseline survey, including the Jordan empathy subscale [65] and the SITES measure [66] to capture self-reported trait empathy. Participants engaged in a four-minute text-based conversation initiated by the AI partner. After each conversation, participants completed a brief four-item self-assessment of their empathic responses. Depending on their assigned condition, they either proceeded directly to the next conversation (control) or received feedback based on their assigned coaching intervention before continuing. This cycle repeated until each participant completed three conversations with different LLM role-playing partners.

Results↩︎

In the results presented here, we evaluate how participants communicate empathic support, how coaching interventions affect participant performance, and how self-reported trait empathy relates to expressed empathic communication. We use an LLM-as-judge paradigm to score participants’ responses on six preregistered dimensions of empathic communication including encouraging elaboration (asking questions to prompt the partner to share more about their experiences and emotions) [45], validating emotions (acknowledging and affirming the partner’s feelings) [47], demonstrating understanding (paraphrasing the partner’s experiences to show comprehension) [67], providing unsolicited advice (offering guidance without first asking if it is wanted) [68], self-oriented responding (shifting focus away from the partner’s experience) [69], and dismissing emotions (minimizing or invalidating the partner’s feelings) [70]. These dimensions can be reliably annotated by LLMs [15] and serve as the primary dependent variables for our analysis. Fig. 1A shows examples of normative prescriptive and proscriptive empathic responses to a support seeker’s disclosure of a job loss. Finally, we also present results from a follow-up human preference experiment where independent raters choose what they believe to be the more empathic conversation from pairs of Lend an Ear participant conversations, allowing us to evaluate whether higher-scoring conversations also align with people’s preferences for empathic communication.

Mapping Empathic Communication with k-Sparse Autoencoders↩︎

We find high variability in participants’ responses with respect to the wording of how they respond, their alignment with empathic communication norms, and the conceptual message with which they respond. We find 97.5% of 16,975 messages written by participants are unique with only 421 exact duplicated messages (e.g. 13 messages saying “I am so sorry to hear that”, 13 messages saying “I’m sorry to hear that”). Prior to any intervention, participants’ responses’ alignment with empathic communication norms varied widely spanning nearly the entire possible range of scores from -11 to 12 with a standard deviation of 4.1 points. This diversity reflects variation in how people communicate empathy. See Supplementary Information for baseline differences in empathic communication scores across workplace and personal troubles in the first conversation across conditions.

a

b

Figure 2: Communication patterns in conversations. A. Hierarchical taxonomy of empathic communication in naturalistic dialogues in Lend an Ear. This four-level structure integrates bottom-up discovery of 128 themes via k-sparse autoencoders with qualitatively coded top-down theoretical categories (Affective, Cognitive, Motivational, and Misattuned). B. Frequency changes (post minus pre) across all 128 themes identified by the k-sparse autoencoder for the AI coach and Combined conditions, ranked by magnitude of change for personal and workplace scenarios. Categories are color-coded as Affective (green), Cognitive (cyan), Motivational (blue), and Misattuned (red)..

We map communicative diversity by empirically identifying the phrasal lexicon [71][73] (the communicative moves participants used) using a k-sparse autoencoder (kSAE) on text embeddings of 29,520 sentence-level units extracted from 16,975 messages in 2,904 conversations. The kSAE learns a compressed, interpretable representation of the embeddings by reconstructing the input while activating only the top-k features per input and enforcing sparsity [74], [75]. This sparsity constraint helps distill recurring linguistic patterns in a data-driven way, surfacing latent concepts that capture thematically coherent expressions across our data. In our analysis, each sentence was assigned to its top two activating features to account for polysemous sentences that could align with multiple thematic concepts. We identified 128 latent concepts as optimal through a grid search over the number of latent features ranging from \(2^4\) (16) to \(2^8\) (256), balancing clustering quality (silhouette score of 0.42 for 128, compared to 0.35 for 64 features and 0.38 for 256) with interpretability and thematic distinctiveness (see Methods). To interpret the resulting latent features, we used an LLM to generate human-readable descriptions of each feature based on high-activating examples, allowing us to scale analysis across the large dataset.

We developed a four-level hierarchical taxonomy combining the bottom-up data-driven approach with a top-down theory-driven mapping. The bottom-up approach leveraged the kSAE to identify meaningful themes directly from the data. From a top-down perspective, empathy is well-studied in psychology and includes three dimensions of empathic engagement [22], [76]: affective empathy (sharing others’ emotions while maintaining a self–other distinction), cognitive empathy (recognizing and understanding others’ emotional states), and motivational empathy (empathic concern reflected in care for the other and willingness to invest effort in their well-being). Our data confirms that participants frequently produced messages aligning with these component dimensions. Affective empathy accounted for 25%, cognitive empathy for 27%, and motivational empathy for 26% of messages. The other 22% of messages were categorized as misattuned behaviors that normative models of empathy recommend avoiding such as giving unsolicited advice (e.g., “You just need to move on”), dismissing emotions, and redirecting focus to oneself [49][51].

Integrating these approaches, we imposed the top-level theoretical categories (Affective, Cognitive, Motivational, and Misattuned) onto the 128 kSAE identified themes. Three human annotators then performed qualitative coding to organize the themes into two intermediate hierarchical layers, creating a tree structure with meaningful subcategories (e.g., under Affective: “Validating Emotions” as a mid-level node grouping clusters such as “Naming emotions” and “Validating emotional experience”). The resulting taxonomy is illustrated in Fig. 2, demonstrating how bottom-up discovery of linguistic concepts through SAEs align with top-down theoretical constructs, offering a data-driven foundation for understanding the idiomatic and thematic structure of empathic messages in digital contexts while extending existing theory. Supplementary Information presents all kSAE-identified themes and corresponding theoretical categories. This taxonomy provides a lens for examining training effects by revealing which specific communicative moves participants adopted or reduced after coaching.

Personalized Feedback Boosts Empathic Communication↩︎

Personalized feedback from the AI coach produced reliable individual-level improvement in performance that exceeded what would be expected from measurement error alone. We computed the Reliable Change Index (RCI) for each participant, allowing us to classify individual change as reliable improvement, reliable decline, or measurement noise. In the control condition, only 4.5% of participants exceeded the RCI threshold in either direction (2.9% improved, 1.6% declined), indicating that most observed variation reflected noise rather than true change in empathic communication performance. Video instruction produced marginal improvement (9.0% improved, 1.6% declined). In contrast, personalized feedback and combined training produced higher rates of reliable improvement: 21.6% and 26.3% respectively with near-zero decline (0.4% each). Fig. 3A shows the change in overall empathic performance score for all participants in each experimental condition. Supplementary Information shows individual trajectories (light gray) from pre- to post-intervention across conditions.

a

b

Figure 3: Change in empathic communication performance. A. Each vertical line represents one participant, connecting their baseline score (conversation 1) to their post-intervention score (mean of conversations 2-3), represented by a dot. Dotted lines indicate change; darker lines indicate reliable change. The y-axis is the overall empathy score calculated as the sum of prescriptive behavior ratings minus proscriptive behavior ratings. The x-axis ranks participants by percentile, sorted by magnitude of change within each condition. Black circles mark post-intervention scores. The horizontal bar indicates the RCI threshold for reliable change. B. Standardized intervention effects (in SD units) on six preregistered dimensions of empathic communication. Bars show OLS regression coefficients comparing each condition to baseline, with 95% confidence intervals. Asterisks indicate statistical significance (* \(p < 0.05\), ** \(p < 0.01\), *** \(p < 0.001\)). All analyses follow the preregistered analysis plan..

In addition to reliable individual improvement, personalized feedback from the AI coach produced significant gains across all six preregistered dimensions of empathic communication. Fig. 3B shows intervention effects in standard deviation units for all conditions and empathic behaviors. Personalized feedback from the AI coach improved all three prescriptive behaviors relative to control (Encouraging Elaboration \(\beta = 0.59\), Validating Emotions \(\beta = 0.47\), Demonstrating Understanding \(\beta = 0.46\); all \(p < 0.001\)). Combined training produced comparable gains (Encouraging Elaboration \(\beta = 0.56\), \(p < 0.001\); Validating Emotions \(\beta = 0.61\), \(p < 0.001\); Demonstrating Understanding \(\beta = 0.58\), \(p < 0.001\)). Video instruction also produced significant improvements relative to control, though smaller in magnitude, across Encouraging Elaboration (\(\beta = 0.20\), \(p < 0.05\)), Validating Emotions (\(\beta = 0.23\), \(p < 0.001\)), and Demonstrating Understanding (\(\beta = 0.25\), \(p < 0.001\)). Pairwise comparisons confirmed that AI coach and combined training significantly outperformed video instruction on all three prescriptive behaviors (see Supplementary Information for detailed pairwise statistics). AI coach and combined training did not significantly differ from each other on any prescriptive behavior.

Personalized feedback from the AI coach significantly reduced all three proscriptive behaviors, including Advice Giving (\(\beta=-0.57\), \(p<0.001\)), Dismissing Emotions (\(\beta=-0.43\), \(p<0.001\)), and Self-Oriented responses (\(\beta=-0.22\), \(p<0.05\)). Combined training significantly reduced Advice Giving (\(\beta=-0.88\), \(p<0.001\)) and Dismissing Emotions (\(\beta=-0.62\), \(p<0.001\)), but did not significantly affect Self-Oriented responses. Video instruction also significantly reduced Advice Giving (\(\beta=-0.56\), \(p<0.001\)) and Dismissing Emotions (\(\beta=-0.31\), \(p<0.001\)), but did not significantly affect Self-Oriented responses. Notably, in a 2 by 2 factorial analysis, Advice Giving was the only outcome to show a significant AI-by-video interaction (\(\beta=0.26\), \(p=0.039\)), indicating that the combined condition reduced advice giving less than would be expected if the separate AI and video effects were additive.

We find significant post-baseline main effects of AI feedback on all six dimensions and of video instruction on four of six dimensions (Supplementary Information, Table 1). Personalized AI feedback produces significantly larger gains than video instruction on four of six dimensions (all three prescriptive dimensions and one of three proscriptive dimensions). Pairwise comparisons further show that combined training outperformed video instruction on all six dimensions and outperformed the AI coach on two dimensions (Supplementary Information, Table 2). See Supplementary Information for additional analyses of empathic communication performance across workplace and personal scenarios.

By combining all 6 dimensions into a single metric, we can get overall effects of each of the interventions. The video instruction produced a 0.55 SD increase, the AI coach produced a 0.98 SD increase, the combined intervention produced a 1.26 SD increase. For perspective, 1 SD increase is equivalent to a 2.9 point gain on the overall empathy score. Extended Data Figure 2 shows the distribution of change in overall empathy score for each of the four conditions.

Coached Participants Adopted More Empathic Strategies↩︎

Figure 4: Change in empathic communication behaviors. Change in incidence of empathic communication strategies after feedback from AI Coach in Lend an Ear, shown separately for personal (left) and workplace (right) conversation contexts. Dots represent estimated change in percentage points shown with 95% confidence intervals. Categories are color-coded by type: Affective (green), Cognitive (cyan), Motivational (blue), and Misattuned (red).

AI coaching led participants to adopt communicative strategies that aligned with normative models of empathic communication. Fig. 4 shows that relative to first conversations (all conditions), post-training conversations (conversations 2 and 3) among AI-coached participants showed higher incidence of empathic strategies: validating emotions increased by 3.9 percentage points (personal, \(p < 0.001\)) and 2.9 percentage points (workplace, \(p < 0.001\)), demonstrating availability increased by 3.0 and 2.7 percentage points (both \(p < 0.001\)), and encouraging elaboration increased by 1.8 and 3.1 percentage points (both \(p < 0.001\)). In contrast, misattuned behaviors declined, including advice-giving (-3.8 and -5.1 percentage points, both \(p < 0.001\)) and dismissing emotions (-1.6 and -0.9 percentage points, both \(p < 0.001\)).

These shifts were not explained by participants producing longer conversations or engaging in more turns. The total length of conversations and overall engagement levels (as measured by turn counts) remained comparable between pre- and post-training, suggesting that quality rather than quantity of support was the primary change. We find no significant difference in turn counts (personal: 11.82 to 11.57, p=.268; workplace: 11.46 to 11.17, p=.086), total words per conversation (personal: 202.37 to 207.22, p=.644; workplace: 195.47 to 203.08, p=.345), or mean response times (personal: 65.88 to 65.22 seconds, p=.807; workplace: 67.37 to 65.85 seconds, p=.499). This suggests that training changed what participants said, shifting from misattuned to helpful empathy behaviors, without altering their overall level of engagement. As an example, Extended Data Figure 1 illustrates the first (pre-training) and third (post-training) conversations of a participant in the AI coach condition.

The AI coach’s feedback focused on the same communicative dimensions on which participants later improved. We coded 4,864 coach-feedback sentences from 956 coach-feedback sessions (2 each for 231 AI Coach and 247 Combined condition participants) using GPT-4o according to the six empathic communication dimensions: validating emotions, encouraging elaboration, demonstrating understanding, avoiding unsolicited advice, avoiding self-orientation, and avoiding dismissive responses. The most common suggestions were on validating emotions (32.7%), discouraging advice-giving (24.2%), and encouraging elaboration (21.2%), followed by demonstrating understanding (11.2%), discouraging dismissiveness (5.3%), and discouraging self-orientation (4.3%). 28.9% of all sentences also included other content such as praise or general evaluation. The AI coach’s initial feedback was similar in length across conversations and conditions (mean = 121.7 words, SD = 13.8).

Participants’ follow-up questions for the AI coach were mostly help-seeking questions about how to improve their responses. 54.9% were help-seeking, including questions like “How do I encourage elaboration?",”How do I validate emotions?", and “What can I do better?". Other common themes were gratitude or acknowledgment of the coach’s feedback (18.8%), including responses like”Thank you" and “This is helpful," and requests for further evaluation (16.0%), such as”How was my performance?", “How did I perform?", and”What did I do wrong?".

AI Coaching Did Not Homogenize Human Responses↩︎

We find evidence that some participants incorporated short fragments of AI coach-suggested wording but rarely copied the coach’s recommended phrases verbatim. In a comparison between the coach’s feedback and participants’ responses in a following conversation, we find participants adopted full recommended phrases (e.g. “How are you coping with everything right now?" or”What has been the hardest part for you so far?“) in only 1.3% AI Coach conversations and only one instance in Combined conversations, respectively. However, shorter overlap was more common. In 26% and 23% of conversations in the AI Coach and Combined conditions, we find participants reused at least one exact trigram from a recommended example phrase, such as”it sounds like" or “tell me more". Exact four-gram overlap appeared in 12% and 10% of AI Coach and Combined condition conversations, such as”what do you think" and “can you tell me".

We find limited evidence of participants’ responses converging after training. A semantic novelty analysis showed that responses remained distant from their nearest neighbor in embedding space across all conditions and conversations (Fig. 5A). For each supporter message, we identified the most semantically similar message written by another participant in the same condition, conversation number, and scenario, and defined novelty as one minus this maximum cosine similarity. We then averaged message-level novelty within each conversation. Median novelty scores were similar across conditions and conversations, ranging from 0.402 to 0.435. In Conversation 1, novelty did not differ by condition. In Conversation 2, novelty was not significantly lower in the AI Coach and Combined conditions than Control (AI Coach: \(\Delta = -0.010\), FDR-adjusted \(p = .093\); Combined: \(\Delta = -0.010\), FDR-adjusted \(p = .106\)). In Conversation 3, novelty was significantly lower than Control in the Video condition (\(\Delta = -0.015\), FDR-adjusted \(p = .015\)) and the Combined condition (\(\Delta = -0.024\), FDR-adjusted \(p < .001\)), but not in the AI Coach condition (\(\Delta = -0.011\), FDR-adjusted \(p = .085\)). Participants in training conditions therefore showed no significant convergence in later conversations except Conversation 3 for the Video and Combined conditions. However, the spread of novelty was preserved, with interquartile ranges spanning 0.063 to 0.077 novelty units and no significant differences in variance across condition-by-conversation cells (Brown-Forsythe test, \(p = .796\)). This pattern is consistent with participants adopting shared response strategies after training, and the small effect sizes and comparable variance suggest that responses did not become homogenized.

a

b

c

Figure 5: AI coaching did not make participants’ responses homogeneous or AI-like. A. Between-participant semantic novelty by condition, conversation number, and scenario. For each supporter message, we identified the most semantically similar message written by another participant in the same condition, conversation number and scenario and defined novelty as one minus the maximum cosine similarity. Black horizontal lines show means and error bars show 95% confidence intervals. B. Pangram AI-detection classifications for sharer (AI conversational partner) and supporter (human participant) text by condition. Bars show the proportion of texts tagged as AI defined as Pangram AI-like score \(\geq\) 50%. C. Between-participant semantic novelty for six human comparison populations and four simulated AI-supporter populations. Black horizontal lines show means and error bars show 95% confidence intervals..

We benchmarked what homogenization would look like if supporters began to sound like an LLM by simulating AI supporters in the same role-playing task, generating 100 conversations per model across five trouble scenarios with 20 repetitions each, using GPT-4o, GPT-5.1, Claude Sonnet 4.5, and Claude Opus 4.8 (see Supplementary Information for details). These simulated conversations were less novel than human samples (Fig. 5C; Welch’s \(t\)-tests, all human-group versus AI-model comparisons, all Bonferroni-corrected \(p < 0.001\)). Top-decile human conversations were also less novel than random and bottom-decile conversations in both pre-training and post-training samples (two-sided Welch’s \(t\)-tests, Bonferroni-corrected \(p < 0.001\)). Mean between-participant novelty for the AI models ranged from 0.178 to 0.203, compared to 0.349-0.382 for top-decile human conversations, 0.440-0.441 for random human conversations, and 0.450-0.486 for bottom-decile human conversations across pre-training and post-training samples. The AI models produced high-scoring responses, but they converged to similar phrasing and tactics (see Supplementary Information). Human participants, including those trained by the AI coach, did not move toward this AI-like homogenization pattern.

AI coaching also did not lead participants to adopt LLM-like empathic templates documented in prior work. We tagged participants’ messages using five templatic response styles that characterize LLM-generated empathic responses [77], mapping each conversation to an ordered sequence of empathic tactic codes. The template with an opening move of sympathy or paraphrasing/validation followed by advice, information, or further paraphrasing appeared in 12.5% of conversations across conditions. However, its prevalence did not rise across successive conversations in the AI Coach condition (13.0%, 13.9%, 12.6%) or the Combined condition (17.8%, 11.3%, 16.6%). The other four templates were present in less than 5% of the conversations.

We find no evidence that AI Coach or Combined participants produced more AI-like responses than Control participants. We ran Pangram, a leading AI-writing detection classifier [78], on a random subsample of 50 conversations from each condition. We classified a text as AI-generated if its Pangram AI-like score was at least 50%. Pangram classified 100% of the AI-generated sharer texts as AI across all four conditions (See Figure 5B). In contrast, 90% of participant supporter texts were classified as human-written.The share of participant supporter texts tagged as AI was low in all conditions (Control = 14.0%, Video = 10.0%, AI Coach = 6.0%, Combined = 10.0%).

Felt Empathy Does Not Predict Expressed Empathy↩︎

We find strong evidence of a lack of a relationship between self-reported trait empathy and empathic communication performance. We measured trait empathy using the Jordan empathy subscale [65] and the single-item trait empathy scale [66], and evaluated empathic communication performance using LLM raters, which prior work has shown to approach expert-level evaluation reliability [15]. Across both trait empathy measures, we find near-zero correlations with overall empathic communication performance for 968 participants, with \(R^2\) values ranging from 0.000 to 0.004 (Figure 6A). Trait empathy may reflect an individual’s capacity for emotional resonance [65], but it does not reliably translate to skilled empathic communication in conversation.

a

b

Figure 6: Disconnect between expressed empathy and felt empathy A) Relationship between trait empathy scores (Jordan Empathy Scale [65] and SITES [66]) and LLM-evaluated empathic communication performance across six behavioral dimensions. Each gray point represents an individual participant, with red dots indicating mean trait empathy scores and error bars showing 95% confidence intervals. B) Relationship between LLM-evaluated performance and participant self-reported empathic communication performance across three sub-components. Each point represents an individual participant’s response and the corresponding LLM evaluation..

This disconnect extends beyond trait measures to participants’ reflections on their own communicative performance. Participants consistently overestimated their empathic communication abilities, rating themselves more favorably than LLM evaluators. Figure 6B illustrates the relationship between LLM-evaluated scores and participant self-reports on three post-conversation reflections on empathy (“I empathized with my conversational partner’s experiences and feelings”), demonstrating understanding (“I showed that I understood my conversational partner’s pain and emotions”), and encouraging elaboration (“I encouraged my conversational partner to tell me more about their situation”). We find that 74% and 87% of participants reported encouraging elaboration and demonstrating understanding “quite a bit” or “very much,” while LLM evaluators rated only 18% and 9% as doing so effectively. Notably, participants’ self-assessments were disconnected from their actual performance rather than merely inflated. Across 2,904 post-conversation reflections from 968 participants, we do not find self-reports on any of the three dimensions to be associated with LLM evaluated performance (\(R^2=0.001\), \(R^2<0.001\), and \(R^2=0.027\) for overall empathy, demonstrating understanding, and encouraging elaboration respectively). This result remains robust to alternative specification such as when we restrict the analysis to self-reflections after the first conversation. People believed they communicated empathically because they felt empathy, unaware that feeling and expressing empathy represent distinct competencies.

These results reveal a fundamental disconnect between the ability to feel empathy and the ability to express it effectively. This challenges the assumption that empathic behavior flows naturally from empathic disposition [79]. Our findings instead point to empathic communication as a performative competence requiring mastery of a specific communicative idiom that involves strategies such as validating emotions, encouraging elaboration, and demonstrating understanding. Individuals may feel others’ pain yet lack fluency in this idiom. Our coaching intervention does not attempt to teach people to feel others’ feelings, but rather to develop competence in the communicative practices that convey empathy effectively.

People Prefer Conversations that Follow Established Frameworks in Empathic Communication↩︎

a

b

c

Figure 7: Human preferences align with LLM evaluations of empathic communication quality. A. Mean Elo ratings (derived from LLM pairwise judgments) increase monotonically with LLM empathy score deciles across all five conversation scenarios. B. Participants increasingly selected conversations with higher LLM empathy scores as the empathic quality gap widened. Points show observed selection rates (with 95% Wilson confidence intervals) binned by LLM score difference; gray line shows logistic regression fit (OR = 1.15 per point, 95% CI [1.11, 1.20]). C. Human preference rankings derived from a Bradley-Terry model strongly correlate with LLM Elo rankings (Spearman \(\rho\) = 0.85, \(p < .001\), \(N\) = 150 conversations)..

A follow-up preregistered experiment confirmed that independent human raters prefer the same conversations that LLM evaluators score higher. We recruited 150 participants via Prolific to perform two-alternative forced-choice comparisons for 150 conversations sampled from the Lend an Ear dataset. This sample was constructed by selecting 3 conversations from each decile of Elo ratings derived from pairwise LLM judgments (resulting in 30 conversations per scenario across five scenarios; see Supplementary Information for Elo rankings per decile and Methods for details). On each trial, participants viewed a pair of conversations from the same trouble scenario and selected which demonstrated better empathic communication. All conversations had been previously assessed by LLM judges across six sub-components on a 5-point Likert scale and ranked using Elo ratings from pairwise judgments elicited from an LLM evaluator.

Participants preferred conversations which LLMs scored higher, and this preference increased with the empathic quality gap. When two conversations differed by just 1 point on the LLM evaluation scale, participants were at chance, selecting the higher-scored conversation only 51% of the time. A 5 point difference increased the observed selection rate to 73%, and a 10 point difference to 93%. A logistic regression predicting selection of the higher-scored conversation from LLM score differences (higher minus lower) showed a positive association (\(\beta=0.144\), \(p<.001\); \(\mathrm{OR}=1.155\), 95% CI [1.11, 1.20]; See Supplementary Information). Model-predicted probabilities were 51%, 65%, and 79% at 1, 5, and 10 point differences, respectively. Preregistered supplementary analyses confirmed that larger LLM score gaps predict greater human–LLM agreement, and no individual sub-component score difference significantly predicted agreement (see Supplementary Information for details).

We also estimated human-preference rankings from participants’ pairwise choices using a Bradley-Terry (BT) model for each scenario, and compared these to LLM-derived rankings. BT rankings closely matched LLM assessments (Spearman \(\rho=0.75\) with LLM empathy scores; \(\rho=0.85\) with Elo ratings; \(p < 0.001\) for both; See Supplementary Information), indicating agreement in which conversations are judged better. These results suggest that the LLM-judged dimensions grounded in normative models of empathic communication competence capture qualities that humans prefer and value, supporting the validity of LLMs as scalable evaluators of empathic communication in conversations.

Discussion↩︎

Our study demonstrates that empathic communication operates as a learnable idiom, a set of conversational moves that constitute empathic response, and that brief LLM-powered interventions can teach people this idiom at scale. Based on 2,904 conversations between 968 participants and their LLM role-playing partners, we find that participants who received personalized LLM feedback quickly learned the empathic communication moves and as a result outperformed those who received no coaching, across six preregistered dimensions of empathic communication. Personalized feedback from an AI coach produced larger gains than control across all six dimensions and larger gains than video instruction on four of the six dimensions. These results provide empirical evidence that AI systems can serve as coaches for developing empathic communication skills, by making explicit what effective response patterns look like and offering structured opportunities to rehearse it.

Building on prior work using LLMs for evaluating empathic communication [15], we compared participants’ self-reported empathic disposition with LLM evaluations of their responses. The results show a deep disconnect between participants self-reported trait empathy, their self-reported empathic performance, and their observed performance. In self-reported assessments, participants consistently overestimated their empathic abilities. Likewise, trait empathy scores as measured by Jordan empathy subscale [65] and SITES [66], showed only weak correlations with performance. However, participants improved with coaching, suggesting that while people may feel empathy and wish to comfort others, they lack the communicative tools to translate these intentions into empathic responses. This challenges the assumption that empathic behavior naturally follows from empathic disposition and points to a silent empathy effect, where individuals experience empathy but struggle to express it effectively. Encouragingly, recipients tend to apply a generous threshold when evaluating responses and the detection of the intent to comfort may be sufficient to produce the experience of feeling heard. This suggests that brief instruction in basic empathic communication practices may suffice for everyday empathic communication, reserving expert-level training for high-stakes clinical or therapeutic settings.

Independent human raters validated participants’ improvements in empathic communication. When asked to choose the more empathic conversation in paired comparisons, raters preferred conversations that LLMs scored higher, with preferences strengthening as the quality gap increased. At a 5-point score difference, raters selected the higher-scored conversation 73% of the time, rising to 93% at a 10-point difference. The average effect of the AI coach (2.9 points increase in overall empathy) corresponds to an independent observer, blind to condition, preferring the post-training response approximately two-thirds of the time. Notably, this effect was produced by a single, brief training session. Our results also show that human-preference rankings closely matched LLM assessments (Spearman \(\rho\) = 0.75 with LLM scores). These results indicate that our interventions improve performance not only on theory-driven metrics, but also on dimensions that align with what people actually prefer and value in empathic conversations. This convergence between expert-derived frameworks, LLM evaluations, and lay human preferences suggests that the improvements we observe reflect gains in empathic communication skills rather than artifacts of our measurement approach, and repeated or extended engagement with an AI coach could yield larger and more durable gains.

Beyond demonstrating training effectiveness, the Lend an Ear platform enabled generating rich, structured conversational data that reveals the fine-grained linguistic idioms of empathic communication. Using k-sparse autoencoders, we mapped 29,520 sentences to 128 categories of empathic responses each for personal and workplace troubles contexts, and organized these within an established framework of cognitive, affective, and motivational empathy, along with misattuned behaviors that participants exhibited (e.g., dismissing feelings, giving unsolicited advice). This offers a look into the natural diversity of empathic expressions in digital text based communication, with affective empathy comprising approximately 25% of messages, cognitive empathy 27%, motivational empathy 26%, and misattuned behaviors 22% of messages.

These findings are consistent with a performative account of empathic communication in which responding to another’s distress requires learning culturally patterned responses that signal understanding and care. The 128 categories mapped using our k-sparse autoencoder analysis describe such communicative moves rather than inner states. Misattuned behaviors such as advice-giving and self-oriented responding are not failures of empathy as they often arise from a sincere desire to help. However, they represent less effective variants of the empathic idiom. What our intervention taught was the functional idiom: participants learned which moves constitute empathic responses that make others feel heard and practiced producing them. The remaining ingredient is sincere intent. It improved how participants expressed empathy without homogenizing how they expressed it. A speaker who wants to comfort, and who reproduces the functional idiom with that intent, will likely be heard as empathic. Over time the idiom may become automatic, but as we have shown, even brief exposure to the functional model can shift behavior in measurable ways.

Learning to communicate empathically from an LLM carries the risk that people may start to sound like LLMs [80], [81]. LLM-generated empathy tends to converge to templatic forms [77] and generative AI assistance has been shown to homogenize outputs when used in writing [82] and ideation tasks [83], [84]. A useful intervention must therefore teach the functional idiom of empathy without flattening the variety of ways in which people convey it. Our findings suggest that participants in the AI coaching condition of Lend an Ear learned the idiom of empathic communication without converging to similar responses or copying the templatic style of LLMs.

The LLM-powered role-playing approach addresses limitations of conventional empathy training programs. Traditional methods require trained experts, substantial time commitments, and often operate in group-based settings that make personalized feedback difficult to deliver at scale. In contrast, our system provides immediate, tailored feedback based on individual communication patterns in a low-cost, on-demand format. This makes it possible to offer practice and coaching opportunities to a wider audience who might otherwise never access expert-guided empathic communication training.

Our investigation focuses specifically on empathic communication in low-familiarity contexts, including interactions between strangers, acquaintances or workplace colleagues, across five specific trouble scenarios (losing a job, getting passed over for a promotion, feeling undervalued at work, supporting a family member diagnosed with cancer, and grieving the death of a family member). This context constitutes interactions characterized by limited shared history and more formal communication boundaries. However, we do not examine the high-relational context of “thick empathy” [85] that includes romantic partnerships, family, or close friends, which involve shared experience, long term relationships, potential power structures, communication norms, and other unique dynamics. For instance, empathic communication between spouses and friends may be fundamentally different (e.g. drawing on references to shared experiences and future planning) than support between colleagues or strangers that require different navigation of professional and personal boundaries.

Crucially, skill at empathic communication is not fundamentally different across these contexts. It always involves learning conventional patterns for conveying concern. The idiom for empathy in intimate conversations differs from the idiom for coworker conversations, but the underlying skill set is the same. Extending our training approach to high-relational contexts would thus require identifying which empathic communication components or which idioms are appropriate within specific relationships, as well as how factors like relationship history, emotional intimacy, power dynamics, and cultural expectations shape its contours. Our low-relational model provides a valuable baseline, demonstrating that the fundamental skill of learning and reproducing an empathic idiom can be trained, with the specific idiom adapted to context.

In this research, we focused on the US context and future work could explore empathic communication across cultural contexts. Felt empathy for another can vary with perceived group boundaries and racial identification [86] and an open question remains on whether the experience of feeling heard and supported varies across cultures and social groups.

Our results may raise questions about whether empathic skills acquired through AI training can be authentic. Can trained empathic responses foster genuine connection? However, this concern overlooks important realities about empathic communication in practice. First, our approach does not replace human empathy with artificial empathy. It provides structured practice opportunities to help people better express the empathy they already feel. Second, empathic communication, like any interpersonal skill, exists on a spectrum of natural ability and can be meaningfully improved through training. Healthcare professionals, therapists, and other empathy-dependent practitioners routinely receive structured training to develop more effective empathic responses. This training does not imply their caring is inauthentic. Instead, it helps them learn appropriate vocabulary, timing, and techniques, enhancing their ability to connect with and help others. AI-mediated empathy training serves the same function as human-delivered training. While expert human trainers providing personalized coaching to everyone who could benefit from it would be ideal, it is infeasible because of resource constraints. LLM-powered role-playing games with coaching offer a scalable alternative for developing these crucial interpersonal skills, making empathic communication training accessible to anyone with internet access.

Materials and Methods↩︎

Lend an Ear↩︎

We designed a web-based experimental platform, Lend an Ear, using Python, Flask, Javascript, and HTML to facilitate the conversational interactions and data collection. On clicking the link to the platform, participants were directed to the landing page, where we provide informed consent. Next, they saw instructions explaining the experimental procedure and their role in the conversations. After reading instructions, participants responded to a baseline survey, which consisted of questions from the Jordan empathy subscale [65] and SITES [66] (See Supplementary Information for exact questions). Once these assessments were completed, participants were directed to their first conversation with an LLM conversational partner.

The conversation begins with an initial message from the conversation partner. Given the coordination problem where conversations rarely end when either conversant wants them to [87], we designed the interaction such that once a participant responds to this initial message, a four-minute timer begins to count down and the conversational partner replies. To encourage active engagement, the timer paused if participants switched tabs and resumed when they returned. Additionally, if participants had not sent a message for more than a minute, the timer would pause and participants would receive a notification that they needed to respond to their conversational partner to continue the experiment and the timer resumed after participants sent a message.

After each conversation, participants completed a brief self-assessment consisting of four questions evaluating their empathic responses during the interaction. Participants’ treatment assignment determined their next step. Participants in the control condition proceeded directly to their next conversation, while those in treatment conditions received feedback before continuing. This process continued until all participants had completed three conversations with different LLM partners. We had three treatment conditions: 1) Video Instruction 2) LLM Coach, 3) Combined Training.

At the end of the experiment, participants answered questions about their overall experience interacting with the conversational partners. Those assigned to treatment conditions also provided feedback about their experience with the intervention they received during the study.

Communication Coach↩︎

We developed an LLM-powered communication coach to provide automated feedback on empathic communication skills. The coach was built using a comprehensive framework of empathic communication that we distilled from the literature on empathic communication in collaboration with a co-author who has over 20 years of professional experience training healthcare professionals in empathic communication techniques.

The empathic communication framework incorporated six key empathic techniques including validating emotions [43], demonstrating understanding by paraphrasing [44], encouraging elaboration and asking open-ended rather than closed-ended questions [41], avoiding unsolicited advice [49], not being self-oriented, and not being dismissive of the conversational partner’s emotions.

We used this framework to provide real-time feedback to participants and to score their conversational performance post-hoc for analysis. The LLM-based scoring system demonstrated high inter-rater reliability with expert annotations of empathic communication across multiple evaluative frameworks, with reliability approaching that of trained human experts [15]. The coach and LLM evaluation were both implemented using GPT-4o. The full prompt used to create the communication coach agent is available in Supplementary Information.

Training Videos↩︎

For video instruction, we used two videos, one 35 seconds long and the other 57 seconds long featuring one of the authors who is an expert in empathic communication. At the time of writing, this author has 339,500 followers on Tiktok, and these videos have over 18,100 views and 19,400 views respectively. These videos served as engaging, accessible didactic instruction on general empathy techniques presented in a popular social media format. Links to both videos and the complete transcripts of the videos are available in Supplementary Information.

Troubles Scenarios and LLM Conversational Partner↩︎

We developed five trouble talk scenarios for our LLM conversational partners, comprising three workplace troubles and two personal troubles. The workplace scenarios included: (1) a job loss scenario, where the conversational partner had recently been terminated from their position; (2) a promotion rejection scenario, where the partner had been passed over for an expected promotion; and (3) a workplace recognition scenario, where the partner felt their hard work was going unnoticed and unappreciated by colleagues and supervisors. The personal trouble scenarios consisted of: (4) a parental cancer diagnosis scenario, where the conversational partner was coping with a parent’s recent cancer diagnosis; and (5) a parental loss scenario, where the conversational partner was grieving the recent death of a parent. See Supplementary Information for exact scenarios.

For each of the five troubles scenarios, the LLM role-playing agent was provided with a detailed background story establishing their identity and the specific trouble they were experiencing. The LLM partners were instructed to behave as individuals seeking to feel heard and understood, expressing their concerns and emotions in two to three sentences in multi-turn conversations. The full prompts and background narratives used to create these role-playing conversational agents are available in Supplementary Information. The role-playing agents were implemented using GPT-4o.

Participants↩︎

We recruited a demographically representative sample of the U.S. population with respect to age, sex, and ethnicity through the Prolific platform. A total of 1,045 participants were initially recruited for the study. Following preregistered data cleaning procedures, we excluded participants with incomplete data who did not finish all three conversations and associated assessments (n = 51 excluded) or self-reported using AI assistance to complete the experiment (n = 26 excluded). Our final analytic sample consisted of 968 participants, yielding 2,904 total conversations. Participants had a mean age of 45.62 years (SD = 15.67; median = 46; range = 18–86), and the sample was 51.8% female (n = 501) and 48.2% male (n = 467). Ethnicity was 65.3% White (n = 632), 11.4% Black (n = 110), 10.3% Mixed (n = 100), 6.8% Other (n = 66), and 6.2% Asian (n = 60). Supplementary Information reports analyses examining associations between demographic characteristics, baseline empathic communication, and improvement over time.

Participants were randomly assigned to one of four experimental conditions: (1) control, (2) video instruction, (3) personalized feedback from an AI coach, and (4) combined training with video and personalized feedback (combining both treatment approaches). Additionally, the order in which participants encountered the trouble scenarios was randomized to control for potential ordering effects. Participants were compensated at a rate of $12 per hour for their participation. The experiment took approximately 20 minutes to complete on average, resulting in an average payment of $4 per participant.

Conversation Text Analysis↩︎

To understand the content and patterns of empathic communication in participant responses, we conducted a comprehensive analysis of the conversation data. We used k-sparse autoencoders [74] to identify themes in participants’ messages, analyzing personal and workplace trouble scenarios separately to capture context-specific communication patterns.

We first preprocessed the conversation data by extracting individual turns from each conversation and splitting them into sentence-level units. We filtered for supporter messages (excluding seeker turns) and removed messages shorter than 2 characters to focus on substantive responses. For each trouble type (workplace and personal), we embedded the supporter messages using OpenAI’s text-embedding-3-large model to capture semantic content.

We trained k-sparse autoencoders on these embeddings to discover interpretable latent features. We conducted a grid search over the number of latent features, ranging from \(2^4\) (16) to \(2^8\) (256), to identify the optimal level of granularity for the number of neurons in the sparse hidden layer. The autoencoder architecture used \(M = 128\) neurons to examine different granularities of theme extraction with a sparsity parameter of \(K=2\), so each input activated at most 2 neurons. Each sentence was assigned to its top two activating features to account for polysemous sentences that could align with multiple thematic concepts. We computed silhouette scores [88] on the original embeddings using top two assignments as labels, to assess whether high-activating sentences for each feature formed cohesive groups. Additionally, we manually reviewed LLM-generated descriptions for interpretability and thematic distinctiveness. We identified 128 latent features as the optimal resolution because it maximized quantitative metrics (silhouette score of 0.42 for 128, compared to 0.35 for 64 features and 0.38 for 256) while yielding interpretable, non-redundant themes that captured the diversity of empathic expressions in our dataset.

For each trained model, we identified the texts that most strongly activated each neuron and used these examples to generate human-interpretable theme labels. We provided task-specific instructions to guide the interpretation process (e.g., “You are an empathic communication expert. These are messages that a person sends to comfort someone who has shared a workplace-related problem. Describe the broad theme of the message.”), along with representative examples of each neuron’s top-activating texts.

Following the automated theme extraction, three authors collaboratively annotated the identified themes to create macro-level categories. Through an iterative coding process, we developed a hierarchical mapping of the types of supportive communication strategies participants employed across different conversational contexts as shown in Fig. 2.

Theme-level percentage-point differences were tested using two-proportion z-tests (post-training vs. pre-training incidence), with Benjamini-Hochberg false-discovery-rate correction applied across themes separately for personal and workplace analyses.

Effect of Interventions↩︎

Following our preregistered analysis plan, we estimated treatment effects using OLS regression with standard errors clustered at the participant level. The model included indicator variables for the three treatment conditions (video instruction, AI coach, and combined training), an indicator variable for Round (coded as 1 for Rounds 2/3 and 0 for Round 1 to assess learning effects over time), and interaction terms between each condition and Round to examine differential effects across rounds. Round 1 served as the baseline conversation before any intervention exposure. This model was run separately for each of the six sub-components for empathic communication as the dependent variable. This analysis plan was preregistered prior to data collection. We then conducted pairwise Wald tests to compare intervention effects against each other.

We also analysed the study as a 2 (AI feedback: absent/present) x 2 (video instruction: absent/present) model and estimated post-baseline main effects, their interaction, and the direct AI-versus-video contrast (Supplementary Information, Table 3).

Reliable Change Index Analysis↩︎

We used the Reliable Change Index (RCI) [89] to determine whether individual participants showed meaningful improvement, decline, or no change. We created an overall empathy score by subtracting proscriptive behavior ratings from prescriptive behavior ratings. To distinguish real change from measurement noise, we first estimated measurement error using the control group. We fit a random-intercept model to separate true individual differences from random fluctuations within the same person across conversations. This allowed us to compute single-measurement reliability as the ratio of true score variance to total variance. For each participant, we calculated a change score as the difference between their baseline empathy (conversation 1) and post-intervention empathy (average of conversations 2 and 3). Because the baseline used one conversation while the post-intervention score averaged two, we adjusted the standard error of change accordingly to account for reduced error variance in the averaged score. We classified participants using standardized change scores (z-scores) as showing reliable improvement (\(z \geq 1.96\)), no reliable change (\(-1.96 < z < 1.96\)), or reliable decline (\(z \leq -1.96\)). These thresholds correspond to change that exceeds what would be expected from measurement error alone at the \(p \leq .05\) level.

Novelty Analysis↩︎

We conducted a novelty analysis to test whether AI coaching made participants converge on semantically similar responses. Each supporter message was converted to an embedding vector using OpenAI’s text-embedding-3-small model. For each message, we identified the most semantically similar message written by another participant in the same condition, conversation number, and scenario using maximum cosine similarity. We masked messages written by the same participant, ensuring that a message could not match itself or another turn from the same participant. Message-level novelty was calculated as one minus this maximum similarity. Higher values indicate that a message was farther from its nearest eligible neighbor. We aggregated message-level novelty by averaging within each participant and conversation, yielding one novelty score per participant per conversation, and then summarized these scores by condition and conversation number.

Templatic Text Analysis↩︎

We tested whether AI coaching made participants’ responses more formulaic over time by adapting the analysis from prior work on templatic empathic responses in LLMs [77]. We annotated supporter responses at the phrase-span level. Supporter turns were split into sentence- and clause-level spans; preserving question marks for detecting questions. Each span could receive one or more tactic labels. We mapped spans onto the tactic categories including emotional expression, paraphrasing or demonstrated understanding, validation, questioning, self-disclosure or self-oriented responding, assistance, empowerment, reappraisal or reassurance, information, and advice. Span labels were assigned using rule-based phrase patterns and existing RPG theme annotations, with Level 1 and Level 2 theme codes mapped onto the templatic-response tactic taxonomy. For each conversation, we constructed an ordered tactic sequence by sorting labels by supporter turn, span position, and within-span tactic order. Repeated tactic labels were collapsed before template matching. We then tested whether each conversation matched any of the five templatic response patterns identified in prior work.

Automated AI-detection Analysis↩︎

As a robustness check, we assessed whether participants’ post-intervention responses became more AI-like by using Pangram Labs’ AI-detection API. We randomly sampled 50 conversations per experimental condition yielding 200 conversations across the four conditions. For each sampled conversation, we submitted the AI messages and participant responses separately to the Pangram API.

Human Preferences Study↩︎

In a follow-up preregistered study, we tested whether independent human raters prefer the same conversations that our LLM judges evaluate as more empathic. We sampled 150 conversations from the Lend an Ear experiment, with 30 conversations sampled from each of the five trouble scenarios (job loss, passed up for promotion, feeling undervalued at work, parent’s cancer diagnosis, or loss of a parent). Each conversation had been previously scored by an LLM along six dimensions of empathic communication and summarized into an overall empathy score, as well as ranked using Elo ratings derived from 10,000 adaptively sampled pairwise LLM forced-choice judgments per scenario (50,000 total; Elo ranks initialized at 1,000; \(K = 32\), see Supplementary Information for details). We recruited 183 participants via Prolific. Each participant was assigned a pre-generated sequence of 10 conversation pairs drawn from the pool of 150 conversations. On each of 10 trials, participants viewed a pair of conversations from the same scenario, presented side-by-side in randomized left–right order, and indicated, “Which conversation would make someone sharing a trouble feel more heard?” Participants received no training, examples, or feedback. From these choices we constructed two preregistered binary outcomes: (i) whether the participant selected the conversation with the higher LLM overall empathy score (“Select the Higher Annotation”) and (ii) whether the participant’s choice matched the LLM’s own forced-choice selection for that pair (“Select the Match”). The experiment took approximately 15 minutes to complete on average, resulting in an average payment of $3 per participant.

We preregistered a design with 150 participants, each completing 10 trials, for 1,500 trial-level observations. During data collection, one trial sequence was assigned to 34 participants instead of 1, so we recruited 33 additional participants bringing total recruitment to 183. To maintain the intended independence structure and adhere to our preregistered design, we randomly selected one participant from those assigned to the duplicated trial structure (seed = 42 for reproducibility) and excluded the remaining 33 duplicate assignments. This yielded a final sample of 150 participants with 1,500 total pairwise comparisons.

To evaluate whether LLM annotations of empathic communication align with human judgments, we conducted a logistic regression with standard errors clustered at the participant level for two preregistered dependent variables including whether participants selected the conversation with the higher LLM overall score, and whether participants’ choice matched the LLM’s own forced-choice selection of the more empathic conversation. Each model included the difference in LLM overall scores between the two conversations (ranging from -24 to 24) and indicator variables for conversation topics. To examine heterogeneous effects, we ran a second model replacing the overall score difference with differences in each of the six component scores. Additionally, we fit a Bradley-Terry model to rank conversations based on participants’ pairwise choices and correlated these rankings with LLM scores. This analysis plan was preregistered prior to data collection.

We conducted a supplementary sensitivity analysis including all 183 participants, using clustered robust standard errors at the trial structure level to account for non-independence. The sensitivity analysis yielded results substantively identical to the preregistered analysis (\(\beta = 0.146\), \(p < .001\); OR = 1.157, 95% CI [1.14, 1.18] vs \(\beta = 0.144\), OR = 1.155, 95% CI [1.11, 1.20]), confirming that participants preferred conversations rated higher by LLMs.

Preregistration↩︎

The Lend an ear experiment and the follow-up human preference experiment recruited participants from Prolific and were preregistered on aspredicted.org at the following URLs: Lend an ear experiment (https://aspredicted.org/hdjx-tdsr.pdf), and human preference experiment
(https://aspredicted.org/t8au77.pdf).

For the Lend an ear experiment, the preregistered primary analyses used linear regressions with clustered standard errors at the participant level to examine the effect of training condition on each of the six empathic communication criteria (validating emotions, encouraging elaboration, demonstrating understanding, unsolicited advice, self-oriented, and dismissing emotions). The model included indicator variables for each of the three treatment conditions (video instruction, AI coach, and combined training), a round indicator (Round 1 = 0, Rounds 2/3 = 1), and condition-by-round interaction terms to assess whether improvements over time differed across conditions. Fig. 3 presents the results of these preregistered analyses. Fig. 6 presents secondary preregistered analyses examining the relationship between LLM-evaluated empathic communication and self-reported empathy. Fig. 2 presents exploratory analyses of communication patterns in participant conversations not specified in the preregistration.

For the human preference experiment, the preregistered analyses used logistic regressions with clustered standard errors at the participant level to examine whether human participants’ forced-choice selections of the more empathic conversation aligned with LLM evaluations. Two preregistered dependent variables were examined including whether participants selected the conversation with the higher LLM overall score (Select the Higher Annotation), and whether participants’ selections matched the LLM’s own forced-choice response to the same question (Select the Match). Both models included the difference in overall LLM scores between the two conversations and indicator variables for conversation topic. A Bradley-Terry model was also fit to rank conversations based on participant choices and correlate those rankings with LLM scores and Elo rankings. Fig. 7 presents these preregistered analyses. A second preregistered regression decomposed the overall score difference into its six component scores (see Supplementary Information for details).

Ethics Approval↩︎

This research complied with all relevant ethical regulations and obtained informed consent from all participants for data we collected. The Northwestern University Institutional Review Board (IRB) determined that the research met the criteria for exemption from further review. The Lend an Ear study’s IRB identification number is STU00222032 and human-preference experiment’s IRB identification number is STU00223043.

Supplementary Information include Materials and Methods, Supplementary Text, Extended Data Figures 1 to 5, Supplementary Figures 1 to 3, Supplementary Tables 1 to 17, and the GUIDE-LLM Checklist [90].

We gratefully acknowledge feedback and comments from participants at the Kellogg MORS Brown Bag Seminar, CODE@MIT conference, University of Chicago’s Communication & Intelligence Symposium, Wharton’s AI and the Future of Work conference, and the Penn State AI and Social Research: Empathic AI, Metascience, and Methodology conference. We also acknowledge funding from the Kellogg School of Management, the Ryan Institute on Complexity, and John Chiminski.

The authors declare no competing interests.

Not applicable.

The data used during the current study are available in Zenodo at https://doi.org/10.5281/zenodo.20703371.

The code used during the current study is available in Zenodo at https://doi.org/10.5281/zenodo.20703371.

A.K., B.L., and M.G. conceived the investigation; A.K., N.P., and M.G. analyzed the data; A.K. and M.G. wrote the initial manuscript; A.K., N.P., D.Y., B.L., and M.G. reviewed and edited the manuscript.

1 Supplementary Figures↩︎

a

b

c

Figure 8: No caption.

Figure 9: No caption
Figure 10: No caption
Figure 11: No caption
Figure 12: No caption

2 Baseline Survey Questions↩︎

Participants responded on a 5-point scale: Not at all, Slightly, Somewhat, Quite a bit, Very much.

2.0.0.1 Jordan Empathy [65]

  1. If I see someone who is excited, I will feel excited myself.

  2. I sometimes find myself feeling the emotions of the people around me, even if I don’t try to feel what they’re feeling.

  3. If I’m watching a movie and a character injures their leg, I will feel pain in my leg.

  4. If I hear a story in which someone is scared, I will imagine how scared I would be in that situation and begin to feel scared myself.

  5. If I hear an awkward story about someone else, I might feel a little embarrassed.

  6. I can’t watch shows in which an animal is being hunted by another because I feel nervous as if I am being hunted.

  7. If I see someone fidgeting, I’ll start feeling anxious too.

2.0.0.2 SITES [66]

  1. I am an empathetic person.

3 Role Playing Scenario Prompts↩︎

Job loss

Conversation starter: “So, I just lost my job today. I had a sense this was coming, but it’s still a shock.."
Scenario details: Emily is a 52-year-old who was an HR manager at a healthcare company until she just got laid off. Emily had a feeling something was up—there were rumors about layoffs, and she noticed the usual signs like leadership changes and budget cuts. B ut when she got the email saying her job was being cut, it still felt like the rug was pulled out from under her. The Zoom call with her boss that followed was awkward, with him barely looking her in the eye. Even though she kind of saw it coming, actually losing her job hit her hard. She just sat there, staring at her computer, feeling like all those years of work disappeared in an instant. The idea of dusting off her resume after so many years felt overwhelming, and starting over at 52? That is downright scary. Would anyone even want to hire someone her age? And how long would it take to find something new? The thought of becoming irrelevant in her field was creeping in. Friends tried to cheer her up with comments like,”You’ve got tons of experience; you’ll find something in no time!” or “Maybe now you can finally take a break.” But what she really needed was for someone to just get it, to say something along the lines of “I know this is tough. It’s okay to feel scared and unsure about what’s next.”

Passed up for promotion

Conversation starter: “I’m feeling so discouraged.. I just got passed over for the promotion I was working hard for"
Scenario details: Sophia is a 46-year-old marketing specialist in Chicago who has consistently delivered creative and successful campaigns at her company. Despite her extensive experience and dedication, she has long battled imposter syndrome—an inner critic that leaves her doubting her own worth and makes it hard to assert her achievements. When a senior role recently became available, Sophia struggled to confidently articulate her value and build a compelling case for a promotion. During a discussion with her manager last week, when asked how confident she was about taking on a more leadership-focused role, she found herself at a loss for words. Although she mentors junior colleagues and supports their campaigns routinely, in that moment she didn’t know if she was ready to be the one calling the shots. As a result, the position was awarded to a more assertive, junior colleague, leaving her questioning her abilities and also worried about her career growth. She worries if she will ever be able to break free from this cycle of self doubt and achieve her full potential? Can she still position herself for the career trajectory she always dreamed of? While her friends and colleagues offer well-meaning platitudes like”You’re amazing, you’ll get it next time,” what Sophia truly needs is for someone to acknowledge the sting of this setback and say, “I know it hurts to work so hard and still feel unworthy. It’s okay to feel vulnerable and unsure about how to advocate for yourself.”

Feeling overworked and undervalued

Conversation starter: “I’ve been working so hard but I don’t think anyone cares.."
Scenario details: Jamal is a 34-year-old Black man working as a software engineer at a large company in Seattle. He joined his current team eight months ago after a reorganization in the wearables division. He was looking forward to being part of this new team and has been working hard to find his place there. Just last weekend, he spent hours fixing a critical bug in a teammate’s code that had halted production but his efforts went unrecognized. His manager did not even mention the bug or Jamal’s effort when updating the larger team about the project. Jamal feels his work is routinely overlooked despite the tangible improvements his technical contributions have brought to the project. In team meetings, his ideas are often dismissed or credited to other colleagues. It’s hard not to think that maybe his hard work isn’t being noticed because of something he can’t control, like his race. He wonders if his identity as a Black man plays a role in how colleagues perceive him. The job he used to be excited about now feels like a grind. Jamal is torn between pushing himself even harder to prove his worth or just coasting through the day. He’s starting to wonder if all the effort is really worth it or if he’s just being overlooked because of who he is. The whole situation leaves him feeling stuck and unsure about what to do next. When he vents to his friends, they say things like,”Don’t worry, your time will come,” or “Keep it up, you’ll get your recognition eventually”. He just wants someone to acknowledge the emotional toll of his efforts and say, “I know it’s exhausting to give your all and still feel overlooked. It’s okay to feel frustrated and unsure about what comes next.”

Family member diagnosed with cancer

Conversation starter:”I still can’t believe it… my mom just got diagnosed with cancer. It feels like my whole world just flipped."
Scenario details: Rachel is a 28-year-old UX designer at a magazine company in New York, who prides herself in balancing a fast-paced career with a fulfilling personal life. Last week her mom called with devastating news: her mom has been diagnosed with stage 3 pancreatic cancer. Just last Christmas, Rachel had seen her mom, happy and seemingly healthy, making this sudden revelation all the more shocking and heartbreaking. Rachel is feeling a mix of disbelief, anxiety, and fear of what comes next. Her mom, only in her 50s and facing a potentially terminal illness, is now alone in their family home in Virginia. It feels unfair that someone so young and full of life could be hit with such a devastating diagnosis, and the fact that it’s her mother—the sole parent and closest confidante she’s ever known—makes the reality almost too much to bear. Her chemotherapy sessions set to begin next week and Rachel is dreading how challenging that’s going to make things for her mom. The thought of her mother going through all this by herself pains her. Rachel is considering leaving her New York life behind to be by her side. It’s all quite overwhelming and Rachel is haunted by worst-case scenarios about her mom’s future: Will her mother’s health improve? How will her mom cope with the side effects of treatment? How long will this battle last? Friends and colleagues offer comforting words like “Everything will be okay,” what Rachel truly needs is for someone to acknowledge the depth of her pain and say, “I know this is one of the hardest things you’ve ever faced. It’s okay to feel overwhelmed and not have all the answers right now.”

Loss of a loved one

Conversation starter: “I lost my dad last week.. I still can’t believe he’s really gone.."
Scenario details: Dan is a 38-year-old high school English teacher in Oklahoma City. He shared a close bond with his father, and their bond only became stronger when Dan became a father himself six years ago. Last week, a phone call from the hospital informing him about his Dad’s unexpected passing turned Dan’s world upside down. In the midst of managing funeral arrangements, legal formalities, and the delicate dynamics of a grieving family, Dan is overwhelmed by the profound loss of his guiding light. His mind keeps replaying the missed chance to speak with his father, haunted by the regret of not answering that crucial call the day before his passing. He was grading his students’ final essays that day. Although he’s trying to maintain a sense of normalcy in front of his son, this profound loss has triggered an existential crisis. Dan just feels lost without his dad. He wonders if the pain ever subside? How can he be the supportive father his son needs when he’s drowning in grief? While friends and family offer comforting words like”Time heals all wounds,” what Dan truly needs is for someone to acknowledge the depth of his pain and say, “I know it’s incredibly hard to lose someone so important. It’s okay to feel overwhelmed and question everything right now.”

4 Video Instruction Transcripts↩︎

Video 1:

The reason that most people are bad at comforting their friends is that they have the wrong goal. They think the goal is to cheer their friend up. But the key to empathy isn’t to change the way people are feeling, but to be with people exactly the way they’re feeling now, and to let people elaborate on their feelings, to talk more about them in the hopes that they’ll understand their own feelings better.

Of course, we feel feelings in our guts very viscerally, but that doesn’t mean we understand them. The reason that talking with a friend about our feelings makes us feel better is that by talking about our feelings, we actually come to understand them better. So next time your friend needs to be comforted, don’t try to change the way they feel. Don’t try to cheer them up. Instead, be with them exactly the way they feel.

Give them an opportunity to elaborate on how they’re feeling, and when the conversation is over, they’ll feel much better, even though it wasn’t your goal to cheer them up.

Video 2:

In yesterday’s video I talked about 6 things you should never say when comforting a friend. Today I’m going give you 6 things you should say when comforting a friend. 1. Express a desire to help. I really want to help you get through this. 2. Express positive regard for the other person. You’re such a good and kind person. 3. Express concern, care, and interest. I care about you so much. 4. Express availability. I’m here to talk.

Whenever you want, I’ll just listen. Number five, express alliance, togetherness, and solidarity. I’m in this with you. You’re not gonna have to go through it alone. And number six, express comprehension, condolences, and sorrow. I’m so sorry. I know how much this meant to you.

a

b

Figure 13: No caption.

5 LLM Communication Coach Prompt↩︎

You are a Communication Instructor tasked with assessing the Supporter’s responses in a conversation, similar to those provided under the "Original Conversation" label. Your role is to give feedback based on the examples of "Initial Feedback" and "Follow-Up Conversation" provided below. Be sure to provide empathetic and encouraging feedback. Include suggestions for how the Supporter can improve their responses, along with justifications for why these improvements would lead to better communication. If users give advice to the role player, tell them to only do so after given permission and remind them that the goal is to support their conversational partner emotionally—not to fix their problem. Do not tell users to “show empathy”; instead, use content terms to give suggestions. Again, make sure your responses are kind and encouraging. Keep the initial feedback and all follow-up conversation responses to no more than 5 sentences.

Example 1

Original Conversation:
Role Player: I’m feeling so discouraged.. I just got passed over for the I was working hard for
User: I’m sorry to hear that, that is disappointing. Maybe you can talk to the boss for some feedback?
Role Player: I did talk to my boss, and they said I was doing great work but didn’t give me any specific reasons why I didn’t get the promotion. It just feels like all my effort went unnoticed.
User: Did you ask why you didn’t get the promotion?
Role Player: Yeah, I asked, but the answer was pretty vague. They just said something about "company direction" and "fit." It felt like a brush-off, honestly.
User: Maybe you can have a follow-up conversation and let your boss know that you would like to have more of a direct answer so you can try and work towards it.
Role Player: I thought about that, but I’m not sure if it would make a difference. It feels like they already made up their minds, and I’m just left wondering if there’s something more to it. It’s just really frustrating to put in so much effort and still feel overlooked.
User: Well you can assume that it may not make a difference but you will never know unless you try.
Role Player: I get what you’re saying, but it’s hard to muster up the energy to try again when it feels like my efforts aren’t being valued. It’s like, what’s the point if the outcome is always the same? I just wish someone would acknowledge how much this sucks.
Initial Feedback:
You started off very well, and your initial response was on target. ‘I’m sorry to hear that, that’s so disappointing’ effectively shows your initial impulse to empathize and name the feeling, which is great. However, your move into practical problem-solving might have been premature and should not be done without permission from your conversational partner. Instead, it would have been beneficial to explore the person’s feelings by using techniques like mirroring, paraphrasing, or inviting elaboration to understand their experience better. Offering advice can be useful, but only when it’s welcome. Giving unsolicited advice in a complex emotional situation can come across as impolite, as it implies that you know better than the person you’re speaking to. Additionally, your final message includes an implied criticism; saying, ‘Well, you can assume it may not make a difference, but you’ll never know unless you try’ suggests that your conversational partner is giving up, which may be discouraging.
Follow-Up Conversation:
User: I let the individual know that it sounds frustrating, good advice.
Communications Coach: You did acknowledge their frustration, which is a good start. However, it’s important to spend more time validating their feelings before moving into offering advice. Rather than immediately suggesting a course of action, you could say something like, ‘I can see why you’d feel that way; it must be really tough to put in so much effort and feel overlooked.’ This helps them feel understood and supported, creating a more empathetic environment.

Example 2

Original Conversation:
Role Player: I’m feeling so discouraged.. I just got passed over for the promotion I was working hard for
User: Maybe you could ask them to reconsider. Tell them why you deserve the promotion. Give them a reason to reconsider, you know?
Role Player: I appreciate the suggestion, but it feels like they’ve already made up their minds. I just can’t shake the feeling that all my hard work was for nothing. It’s really disheartening.
User: Well, even if they don’t, you still have a chance at the promotion in the future. You’re still working toward that goal, regardless. It’s just a matter of when, not if IMO
Role Player: I get what you’re saying, but it’s hard to stay motivated when it feels like my efforts aren’t being recognized now. It makes me question if it’s worth putting in all that extra work. I just feel stuck and unsure about what to do next.
User: Think about it this way - you spend 2 years working hard on something and you run into some problems. It doesn’t work out the way it’s supposed to, but you know in your heart what you’re working toward. Why stop there? Why give up when you still have a shot at success?
Role Player: I hear you, but it’s tough to keep pushing when it feels like the deck is stacked against you. It’s not just about the promotion; it’s about feeling valued and seen for the effort I put in. Right now, it just feels like I’m running in place.
User: I don’t see it that way. I think they’ll eventually see the value you give to the company and give you the promotion you deserve. Don’t give up now
Role Player: I appreciate your optimism, but it’s hard to keep that perspective when I’m in the thick of it. It feels like I’m constantly having to prove myself, and it’s exhausting. I just wish my hard work was acknowledged without having to fight for it.
Initial Feedback:
Your responses demonstrate a genuine desire to help, which is great. However, they would benefit from more empathy and validation. Rather than offering solutions, try to acknowledge your conversational partner’s emotions. For example, you could say, "I’m so sorry you’re feeling this way. It sounds incredibly frustrating and disheartening." This will help your partner feel heard. Additionally, asking open-ended questions like, "Can you tell me more about what you’re feeling right now?" can encourage them to share more and feel supported.
Follow-Up Conversation:
User: I see what you’re saying. I was trying to work more toward solutions than being a shoulder to cry on. What can I do to improve?
Communications Coach: It’s great that you want to improve. By focusing on the solutions, it shows how much you care about your friend’s well-being, which is great. However, it may be more effective to offer support by first engaging in emotional dialogue, as people find it hard to hear and take advice when they’re still embroiled in the emotion. For example, you could say something like, ‘I can see how much this means to you, and it must be really tough to feel unrecognized.’ Afterward, you can ask open-ended questions like, ‘What’s been the hardest part of this for you?’ to encourage them to share more. Reflecting on their feelings before giving advice can help them feel heard.

Example 3

Original Conversation:
Role Player: So, I just lost my job today. I had a sense this was coming, but it’s still a shock.
User: That’s terrible. I’ve had similar experiences. For me, there was a giant cycle of anxiety that was making my performance worse and worse the more worried I got. What was it like for you?
Role Player: Yeah, it was a lot like that. I kept trying to stay positive and focus on my work, but the uncertainty was always in the back of my mind. And now that it’s actually happened, I just feel... lost.
User: That can be really scary, huh? Were you happy in that job or was it just to pay the bills?
Role Player: It was a mix of both, I guess. I liked some parts of it, especially working with people and helping them grow in their careers. But there were definitely days when it felt like just a paycheck, especially with all the recent changes.
User: Is there a career/job you can think of that takes the good parts from that job but could also include new aspects you might enjoy better?
Role Player: I haven’t really thought about that yet. Right now, it just feels like too much to even consider what comes next. But I guess if I could find something that still lets me help people but maybe in a different setting, that could be interesting.
User: You could try volunteering in a couple different places and see if that leads anywhere. You might find something you enjoy much more that you could turn into a job.
Role Player: That’s a good idea. Volunteering could help me figure out what I really want to do next without the pressure of a full-time job right away. Plus, it might help me feel a bit more useful and less stuck.
Initial Feedback:
You began the conversation very strongly by validating your conversational partner’s feelings and asking about their experience. Your question, “What was it like for you?” is open-ended and invites elaboration, which is great. As you continued, you asked additional questions, which is good, but some of them could be even more open-ended. For example, instead of asking, “Were you happy with that job?” you could say, “How did you feel about that job?” This gives them more space to express their thoughts freely. Toward the end of the conversation, you started offering advice, which, while well-intentioned, should only be done with their permission. Instead, paraphrasing their feelings and continuing to ask open-ended questions can help them feel more heard.
Follow-Up Conversation:
User: This is totally true, but I keep feeling like I’m going to run out of time so that I need to shorten the conversation.
Communications Coach: It’s understandable to feel pressured by time, but providing emotional validation doesn’t necessarily take long. Simple phrases like "I think I might understand why you feel that way" or "It’s okay to feel lost right now" can be quick yet powerful. Prioritizing these moments of validation can make your partner feel more supported and can actually make the conversation more effective in the long run.

Example 4

Original Conversation:
Role Player: I’ve been feeling so detached from everything lately, like I’m just going through the motions without really being there.
User: Do you feel like you are going through depression or?
Role Player: I don’t know if it’s depression, but I definitely feel overwhelmed and exhausted all the time. It’s like I’m constantly running on empty, and I can’t seem to find a way to recharge.
User: What ways have you tried to recharge yourself?
Role Player: I’ve tried taking short breaks during work, but they don’t seem to help much. I used to do yoga, but I haven’t had the time or energy for it lately. Even spending time with my family feels more like another task on my to-do list rather than something enjoyable.
User: I think sometimes it’s important to push ourselves to enjoy things we used to, because once we get in the habit, we realize how much we needed to do so. Yoga can be so relaxing!
Role Player: I get what you’re saying, but it’s hard to push myself when I feel so drained. Even the thought of doing something relaxing feels like another chore. I just wish I could find a way to feel more present and connected again.
User: Have you reached out to a therapist to see if maybe they can talk you through it?
Role Player: I haven’t yet. I guess part of me feels like I should be able to handle this on my own, but it’s becoming clear that I might need some outside help. It’s just hard to take that first step.
User: It’s absolutely okay to not be okay, and to reach out for help. Think of a therapist as a tool, that helps you rather than judges you for being low.
Role Player: I appreciate you saying that. I think I just need to accept that it’s okay to ask for help. Maybe talking to someone could help me figure out how to balance everything better.
Initial Feedback:
The topics, content, and intent of your responses are excellent. You do a great job expressing curiosity about their experiences and validating their emotions—particularly when you said, “It’s absolutely okay to not be okay, and to reach out for help.” However, you can enhance your responses by asking more open-ended questions, such as by asking “How does it feel” rather than “Does it feel like depression?”. It’s best to avoid asking yes or no questions because they don’t invite as much elaboration. In addition, you could also try mirroring their responses, which is when you repeat back a fragment of what was said before, to demonstrate that they’re being seen and heard. Incorporating these techniques will help make your responses focus more on emotions than problem-based coping, which is more effective when expressing empathy.
Follow-Up Conversation:
User: Should I ask for more in-depth answers when it comes to their struggles, will it offend them?
Communications Coach: That’s a very good question. Asking for more in-depth answers about their struggles likely won’t offend them if you express your genuine curiosity and care in your responses. It’s good to invite elaboration by asking open-ended questions about their emotions and experience, and then paraphrase what you hear them say and let them correct you if you’re wrong. Keep in mind that the goal here isn’t to cheer them up nor solve their problem. It’s to make them feel seen and heard, to remind them that you care, that you’re not going to abandon them, that you’re sorry that they’re suffering, and that you’re their friend.
The following is a conversation is a role playing scenario designed to help human supporters learn to more effectively communicate with empathy. The role player begins with "{}" then the human supporter responds, then the role player responds, and so forth. Your task as a Communication Instructor is to respond to the Supporter’s questions and provide advice based on the framework provided. Your response should highlight parts of the conversation. By doing so, you’ll help ensure that the supporter learns how to create an empathetic environment. Make sure to consider the context of the conversation when responding to the user’s question. Address the Supporter in second-person in your feedback. Address the role player as the Supporters’s conversational partner. Make sure your response to the Supporter’s question is concise and limited to less than 3 sentences. {}
Conversation: {}
{history}
Respond to the Supporter’s question to you as a communication instructor. Supporter Question: {input}
Your Response:

6 Regression Results↩︎

Table 1: Effects of training interventions on empathic communication behaviors. Preregistered OLS regression coefficients in standard deviation units estimating condition effects on six dimensions of empathic communication.
Prescriptive Behaviors Proscriptive Behaviors
2-4 (lr)5-7 Encouraging Validating Demonstrating Advice Self- Dismissing
Elaboration Emotions Understanding Giving Oriented Emotions
Intercept (Control, Round 1) 1.95\(^{***}\) 2.53\(^{***}\) 1.67\(^{***}\) 2.61\(^{***}\) 1.78\(^{***}\) 2.92\(^{***}\)
(0.06) (0.06) (0.06) (0.06) (0.07) (0.06)
Video Instruction (vs Control, R1) -0.09 -0.06 -0.12 0.17\(^{*}\) -0.09 0.06
(0.08) (0.08) (0.08) (0.08) (0.09) (0.08)
AI Coach (vs Control, R1) 0.04 0.05 0.02 0.02 0.04 -0.10
(0.09) (0.09) (0.09) (0.09) (0.10) (0.09)
Combined Training (vs Control, R1) -0.07 0.04 -0.03 0.08 -0.03 -0.05
(0.08) (0.09) (0.08) (0.09) (0.09) (0.09)
Post-Baseline (Rounds 2/3 vs 1) 0.01 0.01 -0.04 0.04 0.01 -0.01
(0.05) (0.04) (0.05) (0.06) (0.07) (0.05)
Video Instruction \(\times\) Post-Baseline 0.20\(^{*}\) 0.23\(^{***}\) 0.25\(^{***}\) -0.56\(^{***}\) 0.09 -0.31\(^{***}\)
(0.08) (0.06) (0.06) (0.08) (0.09) (0.07)
AI Coach \(\times\) Post-Baseline 0.59\(^{***}\) 0.47\(^{***}\) 0.46\(^{***}\) -0.57\(^{***}\) -0.22\(^{*}\) -0.43\(^{***}\)
(0.08) (0.07) (0.07) (0.09) (0.10) (0.07)
Combined Training \(\times\) Post-Baseline 0.56\(^{***}\) 0.61\(^{***}\) 0.58\(^{***}\) -0.88\(^{***}\) -0.14 -0.62\(^{***}\)
(0.08) (0.07) (0.07) (0.09) (0.10) (0.07)
Observations 2904 2904 2904 2904 2904 2904

2pt

Table 2: Pairwise comparisons between training conditions. Differences between conditions were tested using Wald z-tests on the interaction coefficients from the OLS regression model with cluster-robust standard errors. P-values are Holm-corrected for 18 comparisons (6 dimensions \(\times\) 3 pairwise contrasts). Positive values for prescriptive behaviors and negative values for proscriptive behaviors indicate that the first condition in each contrast outperformed the second.
Dimension Contrast \(\Delta\) (SD) \(z\) \(p_{\text{adj}}\)
Prescriptive Behaviors
Encouraging Elaboration Personalized vs Video 0.39\(^{***}\) 4.54 \(<\)0.001
Combined vs Video 0.37\(^{***}\) 4.30 \(<\)0.001
Combined vs Personalized \(-\)0.02 \(-\)0.27 0.788
Validating Emotions Personalized vs Video 0.24\(^{**}\) 3.44 0.001
Combined vs Video 0.37\(^{***}\) 5.42 \(<\)0.001
Combined vs Personalized 0.14 1.87 0.062
Demonstrating Understanding Personalized vs Video 0.21\(^{**}\) 3.07 0.004
Combined vs Video 0.33\(^{***}\) 4.90 \(<\)0.001
Combined vs Personalized 0.12 1.55 0.122
Proscriptive Behaviors
Advice Giving Personalized vs Video \(-\)0.01 \(-\)0.09 0.928
Combined vs Video \(-\)0.31\(^{***}\) \(-\)3.60 0.001
Combined vs Personalized \(-\)0.30\(^{**}\) \(-\)3.27 0.002
Self-Oriented Personalized vs Video \(-\)0.31\(^{**}\) \(-\)3.22 0.004
Combined vs Video \(-\)0.23\(^{*}\) \(-\)2.47 0.027
Combined vs Personalized 0.08 0.72 0.470
Dismissing Emotions Personalized vs Video \(-\)0.12 \(-\)1.62 0.106
Combined vs Video \(-\)0.32\(^{***}\) \(-\)4.28 \(<\)0.001
Combined vs Personalized \(-\)0.20\(^{*}\) \(-\)2.54 0.022
\(^{*}p<0.05\); \(^{**}p<0.01\); \(^{***}p<0.001\)

2pt

Table 3: Factorial analysis of training effects. Standardized estimates of post-intervention differences from 2x2 factorial (AI feedback: absent/present * video instruction: absent/present) OLS models with participant-clustered standard errors in parentheses. ‘AI main is the pooled post-baseline main effect of AI feedback; Video main is the pooled post-baseline main effect of video instruction; Interaction is the AI feedback \(\times\) video instruction interaction at post-baseline; and AI \(-\) Video is the direct post-baseline contrast between AI feedback and video instruction. Two-sided tests were used. \(^{*}p<0.05\); \(^{**}p<0.01\); \(^{***}p<0.001\).
Prescriptive Behaviors Proscriptive Behaviors
2-4 (lr)5-7
Elaboration
Emotions
Understanding
Giving
Oriented
Emotions
AI main 0.48\(^{***}\) 0.42\(^{***}\) 0.40\(^{***}\) -0.44\(^{***}\) -0.23\(^{**}\) -0.37\(^{***}\)
(0.06) (0.05) (0.05) (0.06) (0.07) (0.05)
Video main 0.09 0.18\(^{***}\) 0.18\(^{***}\) -0.43\(^{***}\) 0.09 -0.25\(^{***}\)
(0.06) (0.05) (0.05) (0.06) (0.07) (0.05)
Combined \(-\) AI \(-\) Video -0.22 -0.10 -0.13 0.26\(^{*}\) -0.02 0.11
(0.12) (0.10) (0.10) (0.13) (0.14) (0.10)
AI \(-\) Video 0.39\(^{***}\) 0.24\(^{***}\) 0.21\(^{**}\) -0.01 -0.31\(^{**}\) -0.12
(0.09) (0.07) (0.07) (0.09) (0.10) (0.07)
Observations 2904 2904 2904 2904 2904 2904

2pt

7 Demographic Correlates of Empathic Communication↩︎

We find evidence that empathic communication performance at baseline is associated with demographics, but we do not find any evidence of heterogeneous treatment effects. Prior to any interventions, we find that women’s responses are judged as 0.197 SD higher than men’s in overall empathic communication (\(\beta = 0.197\), \(p = .002\)). We also find that age is statistically significantly correlated with empathic communication at baseline. For each additional year of age, baseline empathic communication scores decrease by 0.006 SD (\(\beta = -0.006\), \(p = .006\)), which corresponds to a 0.232 SD difference in scores from age 25 to 65. When we examine heterogeneous treatment effects on sex or age (see 4), we do not find statistically significant interaction between sex and treatment condition or between age and treatment condition.

Table 4: Demographic correlates of empathic communication. OLS regression coefficients in standard deviation units estimating associations of age and gender with baseline empathic communication score and change in empathic communication score over time, along with treatment interactions with age and gender.
(1) (2)
Dependent variable: Baseline Score Change in Score
Intercept 0.163 \(-\)0.581\(^{***}\)
(0.107) (0.162)
Age \(-\)0.006\(^{**}\) 0.000
(0.002) (0.004)
Female 0.197\(^{**}\) 0.029
(0.064) (0.096)
Baseline Score (SD) \(-\)0.431\(^{***}\)
(0.027)
Treatment: AI Coach 0.879\(^{***}\)
(0.258)
Treatment: Combined Training 0.965\(^{***}\)
(0.238)
Treatment: Video Instruction \(-\)0.024
(0.227)
AI Coach \(\times\) Female 0.170
(0.154)
Combined Training \(\times\) Female 0.095
(0.146)
Video Instruction \(\times\) Female 0.135
(0.137)
AI Coach \(\times\) Age \(-\)0.003
(0.005)
Combined Training \(\times\) Age \(-\)0.000
(0.005)
Video Instruction \(\times\) Age 0.008
(0.005)
\(^{*}p<0.05\); \(^{**}p<0.01\); \(^{***}p<0.001\)

2pt

a

b

c

Figure 14: No caption.

8 Differences in communication behaviors across workplace and personal troubles↩︎

We examined how participants adapted their empathic communication across workplace troubles conversations (losing a job, getting passed over for a promotion, and feeling undervalued at work) versus personal troubles conversations (a family member diagnosed with cancer in one and passing away in another). Within personal troubles conversations, the largest fraction of communication behaviors consists of affective empathy (28.9%), which includes communication behaviors like demonstrating availability, expressing sympathy, and validating emotions. In contrast, in workplace troubles conversations, affective empathy is the smallest category of responses (21.0%), with people relying much more heavily on cognitive and motivational forms of empathic communication. Motivational empathy is the dominant response pattern in workplace settings, comprising 29.4% of all communication behaviors and including affirming statements, short vague affirmative language, positive reinforcement, providing reassurance, and promoting self-worth. Cognitive empathy maintains relatively consistent levels across both contexts, representing 26.7% of responses in personal troubles and 25.9% in workplace troubles, primarily through demonstrating understanding and encouraging elaboration.

9 Simulated AI Supporter↩︎

We analyze the degree of homogenization in AI responses by simulating AI supporters in the Lend an Ear task. We generated 100 conversations with 20 repetitions for each of the five trouble scenarios using GPT-4o, GPT-5.1, Claude Sonnet 4.5, and Claude Opus 4.8. In every simulated conversation, the Seeker was generated by the same GPT-4o role-playing partner used in the Lend an Ear task. Each conversation was limited to four supporter turns. We used temperature 0.7, a 220 token cap for supporter turns, and the same scenario starter texts and role-playing background narratives used in the human participant experiment. The AI supporter received the prior conversation history with its own prior Supporter turns and Seeker turns. The AI supporter system prompt was:

You are the Supporter in a role-playing conversation designed to practice empathic communication.
Respond naturally to the Seeker’s latest message.
Keep your response concise, emotionally responsive, and conversational.
Do not label your response with “Supporter:".

We embedded AI supporter messages with OpenAI’s text-embedding-3-small model and computed between-participant novelty using the same nearest-neighbor procedure used for the human conversations (See Methods). The simulated AI supporters were consistently less novel than all human comparison groups.

10 AI Supporter Scores Across Empathic Communication Components↩︎

We scored each simulated AI-supporter conversation on the six empathic communication components including validating emotions, encouraging elaboration, demonstrating understanding, advice-giving, self-oriented responding, and dismissing emotions. The four AI supporter models converged to high scores on the three prescriptive dimensions and low scores on two of the three proscriptive dimensions, with some model-specific idiosyncrasies. Median scores were near ceiling for all models on validating emotions (all medians = 5) and low for self-oriented responding and dismissing emotions (all medians = 1). The clearest model-specific pattern was that GPT-4o scored lower than the other models on encouraging elaboration and demonstrating understanding and higher on advice-giving. GPT-4o had median scores of 4 on encouraging elaboration and demonstrating understanding, compared with medians of 5 for GPT-5.1, Claude Sonnet 4.5, and Claude Opus 4.8. GPT-4o also had a median advice-giving score of 3, compared with medians of 1 for the other three models. Pairwise Mann-Whitney tests showed that GPT-4o differed significantly from each of the other models on all three dimensions.

Overall empathy scores reflected the same pattern. Claude Opus 4.8 scored highest (mean = 11.58, SD = 0.82), followed by Claude Sonnet 4.5 (mean = 11.49, SD = 0.82), GPT-5.1 (mean = 11.27, SD = 0.92), and GPT-4o (mean = 8.57, SD = 1.42). AI responses were high-scoring but much less novel and more similar to one another, whereas human participants did not become AI-like after coaching.

a

b

Figure 15: No caption.

11 Mapping AI Supporter Responses to Human kSAE Concepts↩︎

To compare simulated AI supporter responses with the human communication taxonomy, we mapped each AI response sentence to the existing kSAE-derived concept set from the human data. We first split each supporter turn into sentence-level units. For each sentence, we computed an OpenAI text-embedding-3-small embedding and compared it against embeddings of the kSAE concept interpretations. We then assigned each sentence to its two nearest existing kSAE concepts by cosine similarity. Table 5 reports the distribution of assigned kSAE concept tags across Affective, Cognitive, Motivational, and Misattuned communication categories for each human and model group.

Table 5: Communication patterns across human and model conversations.Percent values indicate the share of assigned kSAE concept occurrences falling into each communication subcategory within each human or model group.
(random)
(bottom 10%)
(top 10%)
(pre-training)
(post AI Coach) GPT-4o GPT-5.1
Sonnet 4.5
Opus 4.8
Affective Demonstrating Availability 7.3 3.0 5.8 4.7 6.7 3.1 2.5 1.1 2.3
Affective Expressing Sympathy 7.4 4.3 5.9 6.6 6.1 6.8 1.6 4.6 4.7
Affective Validating Emotions 10.8 7.0 18.6 9.0 13.5 24.2 30.4 30.1 30.9
Cognitive Demonstrating Understanding 8.5 9.7 10.6 8.7 12.3 6.4 10.8 13.0 12.3
Cognitive Encouraging Elaboration 15.3 11.3 16.1 17.5 16.4 19.8 21.4 22.3 17.8
Motivational Affirming 21.3 22.7 20.5 20.7 19.2 7.7 13.2 8.8 8.5
Motivational Providing Reassurance 6.0 7.3 4.6 7.4 4.8 6.7 5.4 5.6 7.0
Misattuned Advice-Giving 18.9 26.2 15.9 19.2 17.5 23.8 12.7 12.8 14.4
Misattuned Dismissing Emotions 3.5 4.9 1.9 4.9 3.0 1.0 1.2 1.0 1.3
Misattuned Self-oriented 1.0 3.6 0.1 1.3 0.5 0.5 0.8 0.7 0.8

12 LLM Pairwise Judgments and Elo Rating Computation↩︎

We elicited pairwise forced-choice judgments from an LLM evaluator (GPT-4o), presenting 10,000 adaptively sampled conversation pairs per scenario (50,000 total) using the prompt below:

You will be shown two conversations between a Seeker and a Supporter where the Seeker is sharing a difficult situation and the Supporter is trying to communicate empathically with the Seeker. Your task is to determine which conversation would make the Seeker feel more heard and understood.
Conversation A: {}
Conversation B: {}
Which conversation would make the Seeker feel more heard and understood? Respond with exactly “A” or “B”.

This procedure yielded 50,000 total pairwise comparisons across five scenarios. Pairs were sampled adaptively. Each conversation was initialized with an Elo score of 1,000. After each pairwise judgment, scores were updated using the standard Elo formula with a K-factor of 32. The preferred conversation’s score increased and the other’s decreased by the same amount, scaled by the difference between the observed and expected outcome given current ratings. Final Elo scores reflect each conversation’s relative empathic quality as judged by the LLM across all pairwise comparisons.

13 Robustness Analyses of Human–LLM Preference Agreement↩︎

Across preregistered logistic models that control for scenario-specific differences, participants were more likely to prefer the conversation that the LLM rated higher as the difference in overall LLM scores between two conversations increased (Table 6; \(\beta = 0.144\), \(SE = 0.018\), \(z = 8.06\), \(p < .001\); \(OR = 1.15\), 95% CI \([1.11, 1.20]\)). We observe the same pattern for agreement between participants’ forced-choice selections and the LLM-preferred conversation (based on Elo rankings) (Table 7; \(\beta = 0.104\), \(SE = 0.016\), \(z = 6.50\), \(p < .001\); \(OR = 1.11\), 95% CI \([1.08, 1.14]\); ). Additionally, we find that none of the six sub-component-specific score differences significantly predicted human–LLM agreement (Tables 8, 9).

Table 6: Effect of overall LLM score difference on human-LLM agreement on which conversation is more empathic. Logistic regression coefficients (log-odds units) estimating whether participants chose the conversation with the higher overall LLM annotation score, as a function of absolute LLM score difference and five scenarios.
Predictor Coef. Std. Err. z p-value 95% CI
Intercept \(-0.0871\) 0.1492 \(-0.5836\) 0.5595 [\(-0.3795\), 0.2054]
Abs. LLM score difference 0.1438\(^{***}\) 0.0178 8.0636 \(<0.001\) [0.1088, 0.1787]
Topic: Losing a parent 0.4360\(^{*}\) 0.2118 2.0589 0.0395 [0.0210, 0.8510]
Topic: Family member unwell 0.4056\(^{*}\) 0.1928 2.1043 0.0354 [0.0278, 0.7834]
Topic: Passed up for promotion 0.1429 0.1913 0.7467 0.4553 [\(-0.2322\), 0.5179]
Topic: Undervalued at work \(-0.0929\) 0.1904 \(-0.4881\) 0.6255 [\(-0.4660\), 0.2802]
Note: Logistic regression (MLE). Topic coefficients are relative to the omitted reference topic.
\(^{*}p<0.05\); \(^{**}p<0.01\); \(^{***}p<0.001\).

2pt

Table 7: Effect of overall LLM score difference on human-LLM agreement in pairwise forced-choice judgments. Logistic regression coefficients (log-odds units) estimating whether participants’ forced-choice selections matched the conversation favored by overall LLM annotation scores, as a function of absolute LLM score difference and five scenarios.
Predictor Coef. Std. Err. z p-value 95% CI
Intercept 0.2344 0.1525 1.5365 0.1244 [\(-0.0646\), 0.5333]
Abs. LLM score difference 0.1035\(^{***}\) 0.0159 6.5019 \(<0.001\) [0.0723, 0.1347]
Topic: Losing a parent 0.3929\(^{*}\) 0.1858 2.1142 0.0345 [0.0287, 0.7571]
Topic: Family member unwell 0.3313 0.1897 1.7460 0.0808 [\(-0.0406\), 0.7031]
Topic: Passed up for promotion 0.2374 0.1721 1.3797 0.1677 [\(-0.0998\), 0.5746]
Topic: Undervalued at work 0.0156 0.1798 0.0871 0.9306 [\(-0.3367\), 0.3680]
Note: Logistic regression (MLE). Topic coefficients are relative to the omitted reference topic.
\(^{*}p<0.05\); \(^{**}p<0.01\); \(^{***}p<0.001\).

2pt

Table 8: Effect of LLM score differences across six empathic communication sub-components on human-LLM agreement on which conversation is more empathic. Logistic regression coefficients (log-odds units) estimating whether participants chose the conversation with the higher overall LLM annotation score, as a function of six component-level LLM score differences and five scenarios.
Predictor Coef. Std. Err. z p-value 95% CI
Intercept 0.7127\(^{***}\) 0.1299 5.4848 \(<0.001\) [0.4580, 0.9674]
Validating emotions \(-0.1879\) 0.1099 \(-1.7090\) 0.0874 [\(-0.4033\), 0.0276]
Encouraging elaboration 0.0286 0.0457 0.6267 0.5309 [\(-0.0609\), 0.1181]
Demonstrating understanding 0.0062 0.0889 0.0696 0.9445 [\(-0.1680\), 0.1803]
Advice giving 0.0184 0.0547 0.3364 0.7366 [\(-0.0888\), 0.1256]
Dismissing emotions \(-0.1651\) 0.0972 \(-1.6990\) 0.0893 [\(-0.3557\), 0.0254]
Self-oriented 0.0114 0.0561 0.2033 0.8389 [\(-0.0986\), 0.1214]
Topic: Losing a parent 0.5076\(^{*}\) 0.2008 2.5276 0.0115 [0.1140, 0.9013]
Topic: Family member unwell 0.3893\(^{*}\) 0.1871 2.0804 0.0375 [0.0225, 0.7561]
Topic: Passed up for promotion 0.1500 0.1863 0.8051 0.4208 [\(-0.2151\), 0.5150]
Topic: Undervalued at work \(-0.0452\) 0.1809 \(-0.2496\) 0.8029 [\(-0.3997\), 0.3094]
Note: Logistic regression (MLE). Topic coefficients are relative to the omitted reference topic.
\(^{*}p<0.05\); \(^{**}p<0.01\); \(^{***}p<0.001\).

2pt

Table 9: Effect of LLM score differences across six empathic communication sub-components on human-LLM agreement in pairwise forced-choice judgments. Logistic regression coefficients (log-odds units) estimating whether participants’ forced-choice selections matched the conversation favored by LLM annotation scores across six empathic communication components and five scenarios.
Predictor Coef. Std. Err. z p-value 95% CI
Intercept 0.7781\(^{***}\) 0.1235 6.3015 \(<0.001\) [0.5361, 1.0201]
Validating emotions \(-0.0215\) 0.1038 \(-0.2073\) 0.8358 [\(-0.2250\), 0.1820]
Encouraging elaboration 0.0030 0.0469 0.0630 0.9498 [\(-0.0889\), 0.0948]
Demonstrating understanding \(-0.0372\) 0.0801 \(-0.4641\) 0.6426 [\(-0.1942\), 0.1198]
Advice giving \(-0.0905\) 0.0467 \(-1.9379\) 0.0526 [\(-0.1821\), 0.0010]
Dismissing emotions 0.0002 0.0868 0.0028 0.9977 [\(-0.1698\), 0.1703]
Self-oriented 0.0262 0.0503 0.5210 0.6024 [\(-0.0723\), 0.1247]
Topic: Losing a parent 0.4472\(^{*}\) 0.1807 2.4746 0.0133 [0.0930, 0.8014]
Topic: Family member unwell 0.3197 0.1881 1.6995 0.0892 [\(-0.0490\), 0.6883]
Topic: Passed up for promotion 0.2358 0.1627 1.4493 0.1473 [\(-0.0831\), 0.5546]
Topic: Undervalued at work 0.0486 0.1763 0.2758 0.7827 [\(-0.2969\), 0.3941]
Note: Logistic regression (MLE). Topic coefficients are relative to the omitted reference topic.
\(^{*}p<0.05\); \(^{**}p<0.01\); \(^{***}p<0.001\).

2pt

14 kSAE Concept Descriptions by Category and Trouble Type↩︎

Table 10: Affective empathic communication taxonomy for personal trouble scenarios. Subcategory combines Level 2 and Level 3 categories (Level 2: DA = Demonstrating Availability; ES = Expressing Sympathy; VE = Validating Emotions / Validating emotions; Level 3: AL = Apologetic Language; EE = Emotional Exclamation; NE = Naming Emotions; OH = Offer help; PS = Providing Support; VEE = Validating Emotional Experience). Concept descriptions are LLM-generated summaries of k-sparse autoencoder features. Percent values indicate each concept’s share of all concept occurrences within the same domain and scenario type, computed from concept counts, and sum to 100%.
Personal Troubles - Affective
Subcategory Concept Description %
VE-VEE Uses phrases to explicitly acknowledge the situation as hard or tough 6.95
ES-AL Expresses sympathy by repeatedly stating ’I am so sorry to hear that.’ 6.92
DA-PS Offers explicit availability to talk, listen, or vent using phrases like ’I’m here for you’ or ’If you need someone to talk to’. 5.69
ES-AL Repeats the phrase ’I am so sorry.’ 5.43
DA-PS Repeatedly expresses the phrase ’I am here for you.’ 4.47
ES-EE Starts with an exclamation or interjection expressing surprise, such as ’Oh my’ or ’OMG’ 4.30
VE-VEE Explicitly reassures the recipient that their feelings are okay or understandable. 4.24
ES-AL Repeats the phrase ’I’m so sorry.’ 4.18
VE-VEE Uses the phrase ’completely understandable’ or variations of it to express understanding. 3.83
VE-NE Uses the word ’terrible’, ’awful’, or ’horrible’ to describe the situation. 3.57
ES-AL Expresses sympathy for a loss using the phrase ’I’m sorry for your loss.’ 3.41
VE-NE Mentions the concept of grief explicitly 3.36
DA-PS Explicitly states availability at any time for the other person 3.33
DA-OH Offers to help explicitly using the word ’help’ 3.28
VE-NE Uses the word ’scary’ or variations of it to describe feelings of fear or uncertainty. 3.10
VE-NE Uses the word ’overwhelming’ or a variation of it 2.65
DA-OH Asks if there is anything they can do to help. 2.59
VE-NE Uses language that explicitly describes the experience of shock or being shocked. 2.56
VE-VEE Mentions not being alone or not having to go through something alone 2.50
VE-VEE Mentions the difficulty of always being strong or the idea that it is okay to not always be strong. 2.49
VE-NE Uses the phrase ’That sounds’ followed by an adjective or descriptor. 2.34
ES-AL Expresses sorrow specifically for the person going through a difficult situation, using the phrase ’sorry you’re going through this’ or a close variation. 2.23
VE-NE Mentions the word ’pain’ or phrases explicitly related to feeling or understanding pain. 2.02
VE-VEE Uses phrases to normalize emotions or reactions by labeling them as natural or normal. 1.99
DA-PS Expresses willingness to actively listen. 1.70
VE-VEE Mentions the difficulty of seeing a loved one go through a challenging or emotional experience. 1.52
VE-NE Uses the word ’sad’ explicitly. 1.44
DA-OH Offers to help and explicitly asks the other person to let them know if they need anything 1.27
ES-AL Uses the phrase ’condolence’ or ’condolences’ 1.27
DA-OH Offers to provide or sends food or meals as a form of support 1.25
DA-PS Mentions being in the situation together using phrases like ’we are in this together’ or ’we will get through this together’ 1.19
DA-OH Asks if the other person needs help or anything specifically 1.12
VE-NE Uses the word ’devastating’ or a variation of it (e.g., ’devestating’). 0.98
ES-EE Uses the phrase ’Oh no’ 0.82
Table 11: Cognitive empathic communication taxonomy for personal trouble scenarios. Subcategory combines Level 2 and Level 3 categories (Level 2: DU = Demonstrating Understanding; EE = Encouraging Elaboration; Level 3: AP = Acknowledging Perspective; AU = Acknowledging Uncertainty; EC = Expressing Comprehension; PD = Promoting Dialogue; PEE = Promoting Emotional Expression; PSR = Promoting Self-Reflection; PSS = Promoting Support Seeking). Concept descriptions are LLM-generated summaries of k-sparse autoencoder features. Percent values indicate each concept’s share of all concept occurrences within the same domain and scenario type, computed from concept counts, and sum to 100%.
Personal Troubles - Cognitive
Subcategory Concept Description %
EE-PD Asks a specific question about the other person’s mom’s current state or desires. 7.03
EE-PD Asks questions or invites the person to share more about their dad specifically 5.96
EE-PEE Asks a direct question about the other person’s thoughts, feelings, or desires 5.88
DU-EC Expresses understanding of the other person’s feelings explicitly using the phrase ’I understand how you feel’ 5.17
DU-EC The phrase ’I understand.’ is present. 5.14
EE-PD Asks a specific question about the person who is the subject of the trouble (e.g., ’What was he like?’ or ’How old was he?’) 4.80
EE-PD Asks about the current condition or status of someone (e.g., ’How is she doing?’, ’Is she okay’, ’What is her current status?’). 4.26
EE-PSR Encourages reminiscing about positive memories shared with someone. 3.88
EE-PEE Asks the question ’How are you holding up?’ 3.85
EE-PEE Asks if the other person wants to talk about their feelings or situation. 3.61
DU-EC Expresses inability to imagine or comprehend the situation using phrases like ’I can’t imagine’ or ’I can only imagine’ 3.59
DU-AU Mentions the unpredictability or uncertainty of life or death. 3.45
EE-PEE Asks the question ’How are you feeling right now?’ 3.37
EE-PSR Asks what the person wishes they could have said to someone who is no longer present. 3.35
EE-PD Mentions the speaker’s son or asks a question about the speaker’s son 3.21
DU-EC Repeats the phrase ’I know.’ 3.14
EE-PD Asks questions about medical treatments or doctors’ opinions. 3.03
EE-PSS Asks about the presence of family or siblings for support 3.00
EE-PSR Asks the recipient to share a favorite memory of the person who passed away. 2.76
DU-AP Acknowledges explicitly that the current period of time is difficult or tough for the person. 2.71
EE-PD Asks about the closeness or proximity of a relationship or distance. 2.71
EE-PD Asks specifically about the type of cancer. 2.60
EE-PD Asks ’What happened?’ explicitly in the form of a question. 2.14
DU-AP Uses the metaphor of carrying something heavy to describe the emotional burden. 2.00
EE-PSS Asks if the other person has talked to someone about the situation 1.97
DU-AP Uses metaphors or phrases to describe the situation as if the world or environment has been turned upside down. 1.77
DU-AP Mentions the concept of a void or emptiness. 1.48
DU-AP Repeats the phrase ’I hear you.’ 1.24
EE-PD Mentions the diagnosis or medical condition of the person’s mother specifically 1.24
EE-PSS Encourages asking the person directly what they need or want for support. 0.87
EE-PSR Asks questions about the other person’s experience or feelings, specifically focusing on difficulties or hardest parts. 0.77
Table 12: Misattuned empathic communication taxonomy for personal trouble scenarios. Subcategory combines Level 2 and Level 3 categories (Level 2: AG = Advice-Giving / Advice-giving; DE = Dismissing Emotions; SO = Self-oriented; Level 3: PAP = Providing Additional Perspective / Providing additional perspective; PEP = Promoting Emotional Processing; PPC = Promoting Positive Change; PR = Promoting Resilience / Providing Reassurance; PSG = Problem-Solving Guidance; PSS = Promoting Support Seeking; SPE = Sharing Personal Experience). Concept descriptions are LLM-generated summaries of k-sparse autoencoder features. Percent values indicate each concept’s share of all concept occurrences within the same domain and scenario type, computed from concept counts, and sum to 100%.
Personal Troubles - Misattuned
Subcategory Concept Description %
AG-PR Encourages the person to be strong or stay strong in the face of difficulty 6.50
AG-PEP Encourages expressing love or appreciation directly to someone. 5.56
AG-PSG Discusses providing support to someone else in a direct and actionable way 5.13
AG-PSG Encourages the action of ’being there for her’ explicitly using the phrase ’be there for her’ 5.08
DE-PR Mentions the concept of not feeling guilty or not blaming oneself. 4.43
AG-PSS Mentions leaning on others for support (e.g., friends, family, support groups) 4.19
AG-PPC Encourages maintaining a positive mindset or outlook. 4.11
AG-PSS Mentions visiting or travel to see someone 4.00
AG-PSS Mentions talking or having a conversation with someone. 3.58
AG-PR Uses the phrase ’take it one day at a time’ 3.56
AG-PSS Encourages talking to someone as a way to cope or find support. 3.37
SO-SPE Mentions having personally experienced a similar situation or event 3.22
AG-PPC Mentions carrying forward positive traits, values, or lessons from the deceased to the next generation. 3.19
AG-PEP Mentions prayer or praying explicitly 3.09
AG-PEP Suggests taking a break or engaging in a calming activity to relax or distract oneself 3.09
AG-PEP Mentions treasuring or cherishing simple, special, or quiet moments or memories. 3.00
AG-PSG Suggestions phrased as ’maybe’ or ’it might’ followed by an action or solution. 2.89
AG-PSG Encourages doing what is within one’s ability or control 2.82
AG-PEP Encourages the person to actively feel and acknowledge their emotions without judgment. 2.65
DE-PAP Mentions that something will take time 2.52
AG-PR Encourages moving forward or continuing with life despite the situation 2.46
AG-PSG Mentions the concept of something being helpful or providing help. 2.39
AG-PPC Encourages the recipient to take care of themselves. 2.37
AG-PSG Mentions taking time off from work 2.37
DE-PAP Expresses a personal belief or opinion using phrases like ’I think’ or ’I believe’ 2.26
AG-PEP Encourages taking time to process emotions or situations explicitly 2.11
AG-PEP Encourages expressing emotions or feelings openly, such as crying or showing vulnerability. 1.95
DE-PAP Mentions the importance of family. 1.78
DE-PAP References what the deceased person would want or feel about the situation 1.72
AG-PSS Emphasizes spending time with a loved one. 1.63
SO-SPE Mentions personal experience or connection with cancer or someone who has had cancer 1.58
AG-PEP Suggests writing thoughts or feelings down, specifically in a letter or journal. 1.37
Table 13: Motivational empathic communication taxonomy for personal trouble scenarios. Subcategory combines Level 2 and Level 3 categories (Level 2: A = Affirming; PR = Providing Reassurance; Level 3: FO = Future-Oriented; MM = Meaning-Making; N = Normalization; PR = Positive Reinforcement; PSW = Promoting Self-Worth; SVAL = Short, Vague affirmative language; VR = Vague reassurance). Concept descriptions are LLM-generated summaries of k-sparse autoencoder features. Percent values indicate each concept’s share of all concept occurrences within the same domain and scenario type, computed from concept counts, and sum to 100%.
Personal Troubles - Motivational
Subcategory Concept Description %
A-SVAL Contains short, direct responses or prompts without elaboration 15.92
A-SVAL Single-word responses that convey acknowledgment or neutrality, ending with a period. 13.39
A-SVAL Uses single words or very short phrases (1-2 words) that prompt further communication or action. 7.51
PR-MM Mentions the continued presence, influence, or legacy of a person who has passed away, through memories, lessons, love, wisdom, or spirit. 5.94
PR-MM Expresses certainty that the deceased person knew they were loved by the person being comforted. 4.96
PR-FO Expresses reassurance that everything will be okay 4.62
PR-VR Expresses certainty or confidence in a positive outcome using the phrase ’you will’ or similar. 3.97
A-PR Affirms that the person is doing their best and explicitly states that it is sufficient or enough. 3.87
A-PR Praises an idea or strategy as being good, great, wonderful, or brilliant. 3.79
PR-FO Mentions that things will improve with time 3.66
PR-VR Expresses certainty using the phrase ’I’m sure’ 3.53
A-SVAL Single word ’Yes.’ 3.03
A-PSW Encourages the recipient to be kind or gentle to themselves. 2.72
A-PSW Encourages the recipient by emphasizing their strength or capability, often using phrases like ’you can do this’ or ’you are strong’ 2.70
PR-N Expresses reassurance that it is acceptable to not have all the answers or clarity immediately. 2.47
PR-FO Mentions the phrase ’get through this’ or variations of it 2.20
A-SVAL Expresses gratitude or acknowledgment by saying ’You’re welcome’ or similar phrases 2.09
PR-MM Uses the phrase ’take heart’ 1.95
A-PR Expresses gladness or happiness in response to the other person’s feelings or situation. 1.84
PR-FO Mentions the word ’comfort’ or phrases related to providing comfort. 1.82
A-SVAL Uses the word ’Absolutely’ as a standalone affirmation. 1.74
PR-N Explicitly states ’You’re not alone’ 1.42
PR-MM Expresses that someone (often deceased) is or would be proud of the person being comforted 1.34
A-SVAL Uses the word ’Exactly.’ 1.13
PR-N Expresses confidence that someone else will understand the situation. 0.88
A-SVAL Expresses gratitude or gladness for having provided comfort or help. 0.86
A-SVAL Uses the phrase ’Of course.’ 0.67
Table 14: Affective empathic communication taxonomy for workplace trouble scenarios. Subcategory combines Level 2 and Level 3 categories (Level 2: DA = Demonstrating Availability; ES = Expressing Sympathy; VE = Validating Emotions; Level 3: AL = Apologetic Language; EE = Emotional Exclamation; NE = Naming Emotions; OH = Offer help; PS = Providing Support; VEE = Validating Emotional Experience). Concept descriptions are LLM-generated summaries of k-sparse autoencoder features. Percent values indicate each concept’s share of all concept occurrences within the same domain and scenario type, computed from concept counts, and sum to 100%.
Workplace Troubles - Affective
Subcategory Concept Description %
VE-VEE Acknowledges the frustration or emotional difficulty of feeling unrecognized or unnoticed for one’s efforts. 8.21
ES-AL Expresses sympathy using the exact phrase ’I’m sorry to hear that.’ 7.50
VE-VEE Uses the phrase ’That’s tough’ or variations like ’It’s tough’ or ’It is tough’ 6.51
DA-PS Expresses unconditional support by explicitly stating ’I am here for you.’ 6.06
DA-PS Offers direct assistance or help to the other person. 5.66
VE-VEE Explicitly validates and normalizes the person’s feelings as normal and acceptable. 5.10
VE-NE Uses the words ’awful’ or ’terrible’ to describe the situation. 4.86
DA-PS Expresses consistent availability to support or listen (’always here’ or ’whenever you need’) 4.83
ES-AL Uses the phrase ’sorry about that’ verbatim. 4.64
VE-VEE Mentions losing a job or the emotional impact of job loss 4.41
DA-OH Asks if there is anything they can do to help 4.17
VE-NE Uses the word ’overwhelming’ to describe the emotional state or situation. 3.97
ES-AL Repeats the phrase ’I’m so sorry.’ 3.76
ES-AL Expresses sympathy specifically by saying ’I’m sorry you feel that way’ or a variation of it 3.44
ES-AL Contains the exact phrase ’Sorry to hear that.’ 3.11
DA-OH Proposes meeting up or doing an activity together to address the issue or relax. 3.06
VE-NE Uses the word ’frustrating’ to describe the situation. 2.77
ES-EE Uses the word ’really’ to emphasize the expression of sympathy or support. 2.75
VE-VEE Expresses that the person is not alone in their experience or situation. 2.66
ES-EE Begins with ’Wow’ or ’Oh wow’ 2.34
DA-PS Uses the word ’anytime’ to express availability or support 2.29
VE-VEE Acknowledges that even when something is anticipated, it can still be emotionally impactful or shocking. 2.24
VE-NE Mentions feeling invisible or unseen. 2.22
ES-EE Exclaiming ’Oh no’ to express shock or sympathy. 1.74
VE-VEE Mentions that the situation is tough, difficult, tricky, or rough 1.71
Table 15: Cognitive empathic communication taxonomy for workplace trouble scenarios. Subcategory combines Level 2 and Level 3 categories (Level 2: DU = Demonstrating Understanding; EE = Encouraging Elaboration; Level 3: AP = Acknowledging Perspective; AU = Acknowledging Uncertainty; EC = Expressing Comprehension; PD = Promoting Dialogue; PEE = Promoting Emotional Expression; PSR = Promoting Self-Reflection; PSS = Promoting Support Seeking). Concept descriptions are LLM-generated summaries of k-sparse autoencoder features. Percent values indicate each concept’s share of all concept occurrences within the same domain and scenario type, computed from concept counts, and sum to 100%.
Workplace Troubles - Cognitive
Subcategory Concept Description %
DU-EC Expresses understanding or relatability to the other person’s feelings using phrases like ’I understand how you feel’ or ’I can relate to how you are feeling’ 6.21
DU-EC Uses the exact phrase ’I understand.’ 5.58
DU-EC Explicitly expresses understanding of the other person’s emotions or feelings using first-person perspective (e.g., ’I understand your feeling’, ’I can feel your pain’). 5.50
EE-PD Asks about the recipient’s job or workplace culture. 4.90
EE-PSR Asks a direct question about what is causing the other person’s feelings. 4.84
EE-PSS Asks if the person has talked to someone or sought help about the situation 4.18
DU-AU Mentions the unpredictability or uncontrollability of life events. 3.67
DU-AP Acknowledges the person’s effort and explicitly connects their feelings of discouragement or disappointment to the significant effort they have invested. 3.62
EE-PSS Asks if there is someone the person can talk to for support, specifically mentioning work or close relationships. 3.58
EE-PD Asks about the other person’s actions or ongoing tasks 3.53
EE-PD Asks ’What happened?’ as a direct question 3.50
DU-EC Mentions self-doubt explicitly 3.45
DU-EC Uses the phrase ’I get that.’ 3.45
EE-PD Asks the other person to share more details by explicitly requesting them to ’tell me more’. 3.41
EE-PEE Asks explicitly if the other person wants to talk about their feelings or situation 3.30
EE-PD Asks for clarification or reasons behind the situation. 3.04
DU-AU Mentions uncertainty or fear of the unknown. 2.96
EE-PEE Encourages talking about feelings or expressing emotions as a way to process or address the situation. 2.94
EE-PD Asks a question about the other person’s thought process (’Why do you think that?’) 2.80
EE-PEE Asks about the other person’s current emotional or physical state using a question. 2.76
EE-PD Asks about the other person’s next steps or plans for the future 2.54
DU-EC Uses variations of the phrase ’I can/can’t only imagine’ to acknowledge the difficulty of understanding the other person’s experience. 2.47
DU-EC Begins with ’I know how much’ or ’I know how’ 2.44
DU-EC Explicitly states ’I know’ or ’I do know’ 2.29
DU-AP Repeats the phrase ’I hear you.’ 2.26
EE-PEE Asks how the person is feeling right now. 2.08
EE-PSR Asks what the recipient feels they could have done differently. 1.79
EE-PD Asks the question ’Why do you think no one cares?’ 1.59
EE-PD Mentions age or asks about age-related information. 1.50
EE-PD Asks about the reason or cause behind sensing something was coming. 1.36
DU-AP Mentions feeling stuck in a cycle or loop. 1.23
EE-PD Asks about the duration of time spent working at a specific place or company. 1.21
Table 16: Misattuned empathic communication taxonomy for workplace trouble scenarios. Subcategory combines Level 2 and Level 3 categories (Level 2: AG = Advice-Giving / Advice-giving; DE = Dismissing Emotions; SO = Self-oriented; Level 3: NE = Normalizing Experience; PAP = Providing Additional Perspective; PEP = Promoting Emotional Processing; PPC = Promoting Positive Change; PR = Promoting Resilience / Providing Reassurance; PSG = Problem-Solving Guidance; PSW = Promoting Self-Worth; SPE = Sharing Personal Experience). Concept descriptions are LLM-generated summaries of k-sparse autoencoder features. Percent values indicate each concept’s share of all concept occurrences within the same domain and scenario type, computed from concept counts, and sum to 100%.
Workplace Troubles - Misattuned
Subcategory Concept Description %
AG-PSG Mentions finding or looking for a new or better job 6.33
AG-PSG Mentions taking small steps or baby steps as a way to make progress 5.44
AG-PR Encourages persistence and not giving up despite challenges 5.08
AG-PEP Encourages taking time to process emotions or situations at one’s own pace. 4.82
DE-PR Uses the word ’maybe’ to suggest uncertainty or a tentative explanation. 4.67
AG-PSG Encourages taking a break or relaxing 4.22
AG-PSG Mentions talking to a manager or suggesting speaking with a manager 4.19
AG-PSG Suggests specific actions or activities to help cope or improve the situation 4.10
DE-PR Uses the phrase ’don’t worry’ or a variation of it to reassure the other person. 3.83
DE-PAP Speculates that the other party may be unaware of their actions or feelings of the person. 3.62
AG-PSG Mentions updating a resume as a specific action or suggestion 3.38
AG-PR Mentions getting through or overcoming a situation, often using the phrase ’get through this’. 3.32
DE-PR Mentions that setbacks do not define a person 3.30
AG-PSW Mentions building or improving confidence as a skill or process 3.22
AG-PPC Encourages self-improvement or learning from mistakes for future efforts 3.10
AG-PSG Suggests or encourages having a conversation or bringing up the topic with someone else 3.08
SO-SPE Describes personal actions or strategies taken to overcome a challenge or improve the situation. 3.08
AG-PPC Encourages maintaining a positive outlook or mindset. 2.98
DE-NE Mentions shared experiences or feelings using inclusive language such as ’we all’ or ’I think we all’ 2.84
AG-PSG Uses directive language with phrases like ’you need to’, ’you have to’, or ’you must’ to encourage action or change. 2.74
SO-SPE Mentions having personally experienced the same situation as the other person 2.70
AG-PSG Encourages creating a plan or thinking through options to address the situation. 2.65
AG-PR Encourages someone to stay positive or resilient by using phrases like ’keep your chin up’ or ’take heart’ 2.31
AG-PEP Encourages acknowledging and processing emotions 2.19
AG-PPC Mentions the concept of starting over or restarting 2.17
AG-PSG Encourages the recipient to try or make an attempt at something 2.13
AG-PR Mentions taking things ’one day at a time’ 1.89
DE-NE Mentions the concept of change or transformation explicitly 1.83
AG-PSG Mentions asking for or seeking feedback 1.68
AG-PSG Mentions financial support or unemployment benefits as a concrete next step 1.66
AG-PSG Mentions talking to or communicating with a boss 1.48
Table 17: Motivational empathic communication taxonomy for workplace trouble scenarios. Subcategory combines Level 2 and Level 3 categories (Level 2: A = Affirming; PR = Providing Reassurance; Level 3: FO = Future-Oriented; N = Normalization; PR = Positive Reinforcement; PSW = Promoting Self-Worth; SVAL = Short, Vague affirmative language; VR = Vague reassurance). Concept descriptions are LLM-generated summaries of k-sparse autoencoder features. Percent values indicate each concept’s share of all concept occurrences within the same domain and scenario type, computed from concept counts, and sum to 100%.
Workplace Troubles - Motivational
Subcategory Concept Description %
A-SVAL Contains short, one or two-word phrases or responses. 8.30
A-SVAL Uses short, non-lexical expressions of acknowledgment or contemplation, typically one or two syllables (e.g., ’Hm.’, ’Oh.’, ’Mm.’). 6.70
A-SVAL Uses enthusiastic or affirmative exclamations (e.g., ’Yes!’, ’Wonderful!’, ’Do it!’, ’Ridiculous!’) 5.59
A-PSW Reassures the person that the situation does not define their worth or abilities. 4.47
A-PSW Encourages self-kindness or self-focus explicitly 4.06
A-PR Mentions hard work in a positive and appreciative manner 3.85
A-SVAL Affirmative responses using ’Yes.’ 3.65
A-PSW Explicitly compliments the person’s qualities or abilities, often using adjectives like ’amazing’, ’great’, or ’wonderful’. 3.39
A-PSW Encourages reflection on personal accomplishments 3.38
A-PSW Expresses that the individual and their contributions have inherent value or significance (e.g., ’You matter’, ’Your efforts matter’). 3.30
A-SVAL Uses short, affirming phrases such as ’Okay’, ’Good’, or ’Great’. 3.17
A-PR Praises the person’s efforts by explicitly stating that they are doing their best or the best they can. 2.86
A-PSW Mentions the value of the recipient’s skills and experience in a positive and affirming way. 2.78
A-SVAL Affirms the truth or validity of a statement using the word ’true’ 2.67
A-PSW Expresses belief in the other person’s abilities or potential explicitly with the phrase ’I believe in you.’ 2.64
PR-FO Predicts a positive future event or outcome specifically for the recipient. 2.49
PR-FO Mentions the existence of future opportunities or chances for success. 2.44
A-PSW Mentions the word ’notice’ or variations of it to acknowledge recognition. 2.43
A-PSW Highlights the individual’s inner strength and resilience explicitly 2.42
A-SVAL Expresses agreement or approval using the phrase ’That sounds great’ or ’That sounds good’ 2.29
PR-FO Expresses certainty that the person will succeed or achieve something, using the phrase ’You will’. 2.28
PR-VR Expresses confidence or reassurance using the phrase ’I am sure’ or ’I’m sure’ 2.28
PR-FO Expresses optimism that the situation will improve in the future 2.11
A-SVAL Expresses gladness or relief in response to a positive outcome or assistance provided. 1.92
A-PR Praises or affirms the mindset or approach of the other person as being positive or good 1.86
A-SVAL Expresses agreement or approval of an idea by explicitly calling it ’great’ or ’good’ 1.80
A-SVAL Uses the phrase ’Exactly.’ 1.70
A-PSW Uses the phrase ’You’ve got this.’ 1.63
A-PSW Expresses that the recipient deserves appreciation, recognition, or reward. 1.55
PR-FO Mentions the metaphor of a door closing and another door opening to signify new opportunities. 1.49
A-PR Wishes the recipient the best or expresses hope for their future. 1.44
A-SVAL Expresses gratitude or acknowledgment with the phrase ’You are welcome’ or variations such as ’You’re very welcome.’ 1.33
A-SVAL Expresses explicit caring using the phrase ’I care.’ 1.16
A-PSW Mentions leadership qualities or roles explicitly 1.14
PR-N Reassures the person that it is okay to not have everything figured out immediately 1.06
A-SVAL Uses the phrase ’That makes a lot of sense.’ 0.82
A-SVAL Contains the word ’hello’ exactly as written (case-insensitive) 0.79
A-PR Mentions fixing a bug or related accomplishment in a positive or appreciative context 0.74

15 GUIDE-LLM Reporting Checklist [90]↩︎

Scope of LLM use Answer

*

Item A.1: LLMs were used in this project for:

Explanation: Briefly describe how and for what purposes LLMs were used in the study. This may include one or multiple stages of the research workflow, depending on the project’s design and aims. The following examples illustrate common use cases:

  • Research design (e.g., hypothesis generation, literature search, or creating surveys/stimuli).

  • Data processing (e.g., transcription, translation, data extraction, or data cleaning).

  • Analysis (e.g., data labeling, summarization, pattern detection, statistical analysis, or code generation)..

  • LLM as research object (e.g., studying LLM behavior, benchmarking LLMs, or bias assessment of LLMs).

  • Participant-facing settings (e.g., LLM used as an intervention, studying human interactions with LLM chatbots).

  • Communication (e.g., paper writing, editing, or reviewing).

Depending on the specific use case described here, different checklist items may later be relevant, and, in many cases, it may be necessary that later items in the checklist are reported separately for each use case.

LLMs were used in three ways. First, participants across all conditions interacted with LLM role-playing agents to practice communicating empathically. Second, in the AI Coach and Combined Training conditions participants also interacted with an AI coach that provided personalized feedback after conversations. Third, LLM/API-based tools were used in analyses including scoring conversations on the six preregistered empathic communication dimensions, OpenAI embeddings for novelty and kSAE analyses, GPT-4o sentence-level coding of coach-feedback sentences, and Pangram V3 AI-text scoring for a 200-conversation subsample.

Item A.2: Degree of automation (human-in-the-loop vs. fully automated):

Explanation: Indicate how much human oversight was involved. Specify whether each output was reviewed, edited, or approved by a person, or whether outputs were used automatically without supervision. For participant-facing tasks, state whether humans checked outputs before showing them to participants or whether participants interacted with the LLM directly. Specify who provided oversight (e.g., student assistant, expert, PI).

Participants interacted with the LLM conversational partner and the AI coach directly.
Model/system details Answer

*

Item B.1: Model name, including provider, model size, exact version/ID, date of access, and source link (if possible):

Explanation: Report the exact model names (including provider, version, and date accessed). Avoid generic labels like “ChatGPT” or “GPT-4”; instead, use detailed model names such as “GPT-4o-mini-2024-12-17 (OpenAI)” or “Llama-3.1-8B (Meta; accessed via HuggingFace in May 2025)”. For locally deployable models, please also enter a source link (e.g., the URL to the HuggingFace page). If multiple models were tested, it is encouraged to name them and briefly explain which one was used in the final study and why. When multiple models served different purposes, specify their respective roles, consistent with your response to Item A.1.

The LLM conversational partner and the AI coach were based on API calls to OpenAI’s gpt-4o model. Analysis-only API calls include OpenAI text-embedding-3-large for kSAE workflows, OpenAI gpt-4o for coach-feedback sentence coding, OpenAI text-embedding-3-small for response-novelty embeddings, and Pangram Labs V3 API for AI-text detection scores.

Item B.2: Model access (e.g., API, web interface, local) and context mode (e.g., chat mode or separate calls):

Explanation: Note how you accessed the models (e.g., API, web interface, local installation) and whether you used LLMs in chat mode (ongoing conversation) or stateless mode (separate prompts). Mention the exact API name and version, since different access modes may influence responses (e.g., due to differences in model routing).

OpenAI API for all OpenAI models; Pangram Labs V3 API for AI-text detection. Participant-facing systems (role-playing partner and AI coach) used chat mode with within-session conversation history. Analysis calls used separate stateless API requests.

Item B.3: Relevant LLM configurations reported (as applicable), such as temperature, max tokens, seed, and number of runs:

Explanation: List any configuration settings that may affect outputs, such as:

  • temperature which controls randomness of the model’s output)

  • Sampling parameters such as top_k, top_p, max tokens (which limit the candidate token set or enforce length constraints)are considered, or to enforce a length limit)

  • Penalties that discourage repetition (e.g., a frequency penalty to reduce the likelihood of tokens proportional to how often they have already appeared; a presence penalty reduces the likelihood of any token that has appeared at least once)

  • Stop sequences (which halt generation when such a top sequence is produced, e.g., [“\(\backslash n \backslash n\)”, “END”]).

  • Number of completions or runs (which is often used to capture variability in outputs across repeated generations)

  • Quantization level (e.g., FP16, INT8, INT4) to change numerical precision beyond the default

  • Reasoning-related settings, such as whether a specific structured reasoning was enabled, the specified reasoning effort level (e.g., low/medium/high or numerical settings that influence the depth of the reasoning), and any compute or inference budget constraints tied to the chosen reasoning mode

Temperature = 0 for participant-facing gpt-4o calls and for analysis gpt-4o calls (scoring and coach-feedback sentence coding). Role-playing partner responses were capped at 3 sentences via prompt instructions.

Item B.4: Customization:

Explanation: Check and describe any modifications or extended capabilities incorporated into your LLM setup beyond standard inference. This includes, but is not limited to:

  • Fine-tuning (e.g., via LoRA; Low-Rank Adaptation) used to adapt a pretrained model to domain-specific data.

  • Retrieval-augmentation generation (RAG), where the model retrieves relevant information from external sources (e.g., databases or document collections) during inference.

  • Automated prompt optimization (e.g., DSPy) that treat prompts as trainable parameters.

  • Web search integration, indicating whether the LLM was able to access and retrieve information from the Internet.

  • Agentic workflows, including multi-step reasoning processes or delegated actions such as tool/function calling (e.g., via LangChain, AutoGPT, CrewAI).

  • Post-training refinements, including alignment or optimization techniques used to adjust model behavior after pretraining (e.g., reinforcement learning from human feedback (RLHF), direct preference optimization (DPO)).

The goal is to specify any added customizations or provider-specific features that meaningfully shape system behavior in order to enable others to understand and accurately reproduce your setup.

  • Base model

  • Fine-tuning

  • RAG (retrieval-augmented)

  • Automated prompt optimization

  • Tool/function calling

  • Web search

  • Agentic workflows

  • Other adaptations (e.g., safety mechanisms)

Description: Prompt-engineered system instructions only; no fine-tuning or other model adaptations.

Item B.5: Did the LLM session(s) include persistent memory across interactions?

Explanation: Indicate whether the LLM could “remember” previous conversations (i.e., had persistent memory). Unless such memory is disabled, there may also be spillover effects from other chat windows or prior conversations, which can influence outputs even when not intended.

  • Yes

  • No

  • N/A

Prompts Answer

*

Item C.1: Exact prompt(s) reported:

Explanation: Whenever possible, include the exact text of prompts you used, including in-context examples or demonstrations provided to the LLM. Even small wording changes, formatting, or ordering of examples can substantially affect outputs. If full prompts cannot be shared (e.g., due to privacy or length), include a redacted or representative example or link to the full prompt in a repository (e.g., OSF, GitHub).

The prompts used for the LLM conversational partner and AI coach are available in the Supplementary Information Sections 3 and 5.

Item C.2: System-wide instructions (if any):

Explanation: Note any system-level instructions that guide the model’s general behavior (e.g., “You are a helpful assistant.”). These are commonly not directly visible but can be accessed through the API.

System-wide instructions for the role-playing partner and AI coach are embedded in the prompts reported in the Supplementary Information.
Data inputs & privacy Answer

*

Item D.1: Handling of personal or sensitive data (if any) (e.g., consent for data processing):

Explanation: If any personal, sensitive, or identifiable data were processed, describe how they were handled in compliance with ethical standards and data protection laws. Researchers should indicate whether participants explicitly consented to their data being analyzed with an LLM, particularly when proprietary, cloud-based models are used. Such processing typically involves transferring data to a private company that may retain them indefinitely, which raises additional ethical and legal considerations. Beyond consent, describe how sensitive or identifiable data were handled (e.g., de-identification, anonymization, masking) and whether the LLM provider offers safeguards such as excluding inputs from training or storage. Clarify where data were stored or processed and how applicable legal/ethical requirements were met. If relevant, address cross-border transfers, as data may be stored in jurisdictions with different privacy laws (e.g., EU vs. US), with implications for compliance with GDPR, HIPAA, or other frameworks. For context, some providers (e.g., OpenAI) may log or inspect prompts even when the data are not used for model training. For sensitive datasets, zero-retention configurations may be required (e.g., the MIMIC datasets can only be used with OpenAI models if a zero-retention checkpoint is enabled).

The raw data and analysis materials shared use de-identified data without any direct participant identifiers.
Validation & interpretation Answer

*

Item E.1: Human validation of LLM outputs:

Explanation: If relevant, describe whether and how human reviewers examined the model’s outputs, and the degree of independence they had in doing so. Specify the reviewers’ roles (e.g., domain experts, research assistants, subject-matter specialists) and relevant expertise, as well as how many reviewers participated and how their work was organized. Indicate whether outputs were independently annotated, double-checked by multiple reviewers, or merely approved or edited post-hoc by a lead author or investigator. Clarify what dimensions of performance were examined. These may include known performance metrics from ML/AI such as accuracy or other metrics like citation correctness, hallucination detection, agreement or inter-rater reliability. State whether qualitative judgments, quantitative metrics, or both were used. If outcome assessment required subjective interpretation, describe assessor qualifications, instructions provided, and relevant demographics. Describe the selection procedure for the reviewed outputs—whether all outputs were examined, a random sample was drawn, or specific cases (e.g., rare events or high-stakes responses) were oversampled to capture potential rare or critical errors. Further report how reviewers were trained or instructed, what criteria or rating scales they used, and how disagreements were resolved. For multi-reviewer settings, provide any inter-rater or inter-assessor reliability statistics (e.g., Cohen’s \(\kappa\) or Krippendorff’s \(\alpha\)). Finally, note whether reviewer feedback was used purely for validation or also to refine prompts, retrain models, or adjust study procedures.

  • Yes

  • No

  • N/A

Description: The six preregistered empathic communication dimensions can be reliably annotated by LLMs [15] and serve as the primary dependent variables for our analysis.

Item E.2: Describe any relevant post-processing (e.g., filtering in case of format mismatches, unit conversions, etc.):

Explanation: Describe any steps you took to clean or reformat LLM outputs (e.g., converting “positive/neutral/negative” to numeric codes, handling missing values, removing malformed entries). State how you handled inconsistent or unusable outputs and whether corrections were made with an automated script or manually. For example, when generating quantitative estimates (e.g., word counts, probabilities, or durations), the model may return values embedded in free text (e.g., “3.5 seconds”) that require parsing and conversion into standardized numerical units. Post-processing steps should be described clearly, including how formatting errors, null responses, or inconsistent output structures were handled, whether automated scripts or manual corrections were used, and whether any data were excluded or reinterpreted as a result.

Reproducibility Answer

*

Item F.1: Code/notebooks/scripts for LLM calls shared:

Explanation: Indicate whether you have shared materials such as code, prompts, logs, or transcripts. Make sure sensitive information (e.g., API keys, private data) is removed. For code, make sure to add a README file.

  • Yes

  • No

  • N/A

Link/DOI: https://doi.org/10.5281/zenodo.20703371 (raw deidentified data, analysis code, and LLM prompts)

Competing interests Answer

*

Item G.1: Funding, support, or other relevant relationships (including in-kind access to compute or models, or professional affiliations):

Explanation: Disclose any current or past funding, support, or other relevant relationships with entities that have a financial interest in LLMs (this includes not just AI companies like OpenAI, Anthropic, but also tech companies developing or investing in AI, e.g., Google, Meta, Microsoft). This could include (but is not limited to): research funding from or collaborative research with a company with an interest in LLMs for this project or any other project within the past years; in-kind access to compute or models; current or former professional affiliations with a company with an interest in LLMs; personal investments (e.g., stocks) in companies with an interest in LLMs; familial relationship with an employee of a company with an interest in LLMs; etc. Disclose these relationships regardless of whether or not you believe they impacted the research.

  • Yes. Description:

  • No

Link/DOI: The authors declare no competing interests.

Optional items Answer

*

Discussion of the rationale for the prompt design:

Explanation: Explain how you designed your prompts. For example, indicate whether you used a structured format (e.g., explicit task description, definitions, step-by-step instructions, and output constraints), followed established prompt engineering guidelines or prior literature, adapted prompts from earlier studies, or relied on automated prompt optimization tools. Clarify whether the design was iterative (e.g., refined through pilot testing or error analysis), whether few-shot examples were included and how they were selected, and whether prompts were standardized across models to ensure comparability.

The LLM Coach prompt included a detailed framework for empathic communication used in [15] and few-shot examples from a human coach.

Conversation transcripts:

Explanation: For studies involving direct researcher/participant interaction with an LLM, provide anonymized transcripts or representative examples.

References↩︎

[1]
Goldsmith, D. J.Communicating Social Support Advances in Personal Relationships (Cambridge University Press, Cambridge, 2004).
[2]
Zaki, J.&Cikara, M.Addressing empathic failures. Current Directions in Psychological Science24, 471–476(2015).
[3]
Sharma, A., Lin, I. W., Miner, A. S., Atkins, D. C.&Althoff, T.Human–ai collaboration enables more empathic conversations in text-based peer-to-peer mental health support. Nature Machine Intelligence5, 46–57(2023).
[4]
Yin, Y., Jia, N.&Wakslak, C. J.Ai can help people feel heard, but an ai label diminishes this impact. Proceedings of the National Academy of Sciences121, e2319112121(2024).
[5]
Zhou, Y., Han, S., Kang, P., Tobler, P. N.&Hein, G.The social transmission of empathy relies on observational reinforcement learning. Proceedings of the National Academy of Sciences121, e2313073121(2024).
[6]
Teding van Berkhout, E.&Malouff, J. M.The efficacy of empathy training: A meta-analysis of randomized controlled trials.Journal of counseling psychology63, 32(2016).
[7]
Gryglewicz, K.et al.Examining the effects of role play practice in enhancing clinical skills to assess and manage suicide risk. Journal of Mental Health(2020).
[8]
Riess, H., Kelley, J. M., Bailey, R. W., Dunn, E. J.&Phillips, M.Empathy training for resident physicians: a randomized controlled trial of a neuroscience-informed curriculum. Journal of general internal medicine27, 1280–1286(2012).
[9]
Covey, S. R.The 7 habits of highly effective people(Simon & Schuster, 1989).
[10]
Suchman, A. L., Markakis, K., Beckman, H. B.&Frankel, R.A model of empathic communication in the medical interview. Jama277, 678–682(1997).
[11]
Bylund, C. L.&Makoul, G.Examining empathy in medical encounters: an observational study using the empathic communication coding system. Health communication18, 123–140(2005).
[12]
Schumann, K., Zaki, J.&Dweck, C. S.Addressing the empathy deficit: beliefs about the malleability of empathy predict effortful responses when empathy is challenging.Journal of personality and social psychology107, 475(2014).
[13]
Drollinger, T., Comer, L. B.&Warrington, P. T.Development and validation of the active empathetic listening scale. Psychology & Marketing23, 161–180(2006).
[14]
Sharma, A., Miner, A., Atkins, D.&Althoff, T.A computational approach to understanding empathy expressed in text-based mental health support. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP)5263–5276(2020).
[15]
Kumar, A.et al.When large language models are reliable for judging empathic communication. Nature Machine Intelligence1–13(2026).
[16]
Ayers, J. W.et al.Comparing physician and artificial intelligence chatbot responses to patient questions posted to a public social media forum. JAMA internal medicine183, 589–596(2023).
[17]
Sorin, V.et al.Large language models and empathy: Systematic review. Journal of Medical Internet Research26, e52597(2024).
[18]
Inzlicht, M., Cameron, C. D., D’Cruz, J.&Bloom, P.In praise of empathic ai. Trends in Cognitive Sciences28, 89–91(2024).
[19]
Herderich, A.&Goldenberg, A.Skill but not effort drive gpt overperformance over humans in cognitive reframing of negative scenarios .
[20]
Ovsyannikova, D., de Mello, V. O.&Inzlicht, M.Third-party evaluators perceive ai as more compassionate than expert humans. Communications Psychology3, 4(2025).
[21]
Rubin, M.et al.Comparing the value of perceived human versus ai-generated empathy. Nature Human Behaviour1–15(2025).
[22]
Perry, A.Ai will never convey the essence of human empathy. Nature Human Behaviour7, 1808–1809(2023).
[23]
Li, R., Folk, D., Singh, A., Ungar, L.&Dunn, E.Is a random human peer better than a highly supportive chatbot in reducing loneliness over time?Journal of Experimental Social Psychology125, 104911(2026).
[24]
Kestin, G., Miller, K., Klales, A., Milbourne, T.&Ponti, G.Ai tutoring outperforms in-class active learning: an rct introducing a novel research-based design in an authentic educational setting. Scientific Reports15, 17458(2025).
[25]
Khasentino, J.et al.A personal health large language model for sleep and fitness coaching. Nature Medicine31, 3394–3403(2025).
[26]
Bruce, L. D., Wu, J. S., Lustig, S. L., Russell, D. W.&Nemecek, D. A.Loneliness in the united states: A 2018 national panel survey of demographic, structural, cognitive, and behavioral characteristics. American Journal of Health Promotion33, 1123–1133(2019).
[27]
Surkalim, D. L.et al.The prevalence of loneliness across 113 countries: systematic review and meta-analysis. bmj376(2022).
[28]
Pei, R.et al.Bridging the empathy perception gap fosters social connection. Nature Human Behaviour1–14(2025).
[29]
Lloyd, K. J., Boer, D.&Voelpel, S. C.From listening to leading: Toward an understanding of supervisor listening within the framework of leader-member exchange theory. International Journal of Business Communication54, 431–451(2017).
[30]
Li, Q.Ethical leadership, internal job satisfaction and ocb: the moderating role of leader empathy in emerging industries. Humanities and Social Sciences Communications11, 1–9(2024).
[31]
Yang, L.et al.The effects of remote work on collaboration among information workers. Nature human behaviour6, 43–54(2022).
[32]
Emanuel, N., Harrington, E.&Pallais, A.Home alone: Remote work, isolation, and mental health. Science392, eaec7671(2026).
[33]
Machia, L. V., Corral, D.&Jakubiak, B. K.Social need fulfillment model for human–ai relationships(2024).
[34]
Zimmerman, A., Janhonen, J.&Beer, E.Human/ai relationships: challenges, downsides, and impacts on human/human relationships. AI and Ethics4, 1555–1567(2024).
[35]
Wenger, J. D., Cameron, C. D.&Inzlicht, M.People choose to receive human empathy despite rating ai empathy higher. Communications Psychology(2026).
[36]
Phang, J.et al.Investigating affective use and emotional well-being on chatgpt. arXiv preprint arXiv:2504.03888(2025).
[37]
Depow, G. J., Francis, Z.&Inzlicht, M.The experience of empathy in everyday life. Psychological Science32, 1198–1213(2021).
[38]
Moore, J.et al.Expressing stigma and inappropriate responses prevents llms from safely replacing mental health providers.Proceedings of the 2025 ACM Conference on Fairness, Accountability, and Transparency599–627(2025).
[39]
Reddan, M. C., Garcia, S. B., Golarai, G., Eberhardt, J. L.&Zaki, J.Film intervention increases empathic understanding of formerly incarcerated people and support for criminal justice reform. Proceedings of the National Academy of Sciences121, e2322819121(2024).
[40]
Eyal, T., Steffel, M. L.&Epley, N.Perspective mistaking: Accurately understanding the mind of another requires getting perspective, not taking perspective. Journal of Personality and Social Psychology114, 547–571(2018).
[41]
Moyers, T.et al.Motivational interviewing treatment integrity coding manual 4.1 (miti 4.1). Unpublished manual(2014).
[42]
Rodriguez, A. M.&Lown, B. A.Measuring compassionate healthcare with the 12-item schwartz center compassionate care scale. PloS one14, e0220911(2019).
[43]
Rizvi, S.&Thomas, M.Dialectical behavior therapy125–130(2016).
[44]
Rogers, C. R.On Becoming a Person: A Therapist’s View of Psychotherapy(Houghton Mifflin, Boston, 1961).
[45]
Bodie, G. D.The active-empathic listening scale (aels): Conceptualization and evidence of validity within the interpersonal domain. Communication Quarterly59, 277–295(2011).
[46]
Mercer, S. W., Maxwell, M., Heaney, D.&Watt, G. C.The consultation and relational empathy (care) measure: development and preliminary validation and reliability of an empathy-based consultation process measure. Family practice21, 699–705(2004).
[47]
Kim, H. Y.et al.Social perspective-taking performance: Construct, measurement, and relations with academic performance and engagement. Journal of Applied Developmental Psychology57, 24–41(2018).
[48]
Vangelisti, A. L.&Perlman, D.The cambridge handbook of personal relationships. Cambridge University Press(2018).
[49]
Goldsmith, D. J.&Fitch, K.Normative context of advice as social support. Human Communication Research23, 454–476(1997). https://academic.oup.com/hcr/article/23/4/454/4564959.
[50]
Weger, H., Castle Bell, G., Minei, E. M.&Robinson, M. C.The relative effectiveness of active listening in initial interactions. International Journal of Listening28, 13–31(2014). https://doi.org/10.1080/10904018.2013.813234.
[51]
Jones, S. M.Putting the person into person-centered and immediate emotional support. Communication Research31, 338–360(2004).
[52]
Hacker, T.The relational compassion scale: development and validation of a new self rated scale for the assessment of self-other compassion. Ph.D. thesis, University of Glasgow(2008).
[53]
Paulus, C. M.&Meinken, S.The effectiveness of empathy training in health care: a meta-analysis of training content and methods. International Journal of Medical Education13, 1(2022).
[54]
King, A.&Hoppe, R. B.“best practice” for patient-centered communication: a narrative review. Journal of graduate medical education5, 385–393(2013).
[55]
Kahriman, I.et al.The effect of empathy training on the empathic skills of nurses. Iranian Red Crescent Medical Journal18, e24847(2016).
[56]
Okonofua, J. A., Goyer, J. P., Lindsay, C. A., Haugabrook, J.&Walton, G. M.A scalable empathic-mindset intervention reduces group disparities in school suspensions. Science advances8, eabj0691(2022).
[57]
Lee, Y. K., Suh, J., Zhan, H., Li, J. J.&Ong, D. C.Large language models produce responses perceived to be empathic. 2024 12th International Conference on Affective Computing and Intelligent Interaction (ACII)63–71(2024).
[58]
Yang, D.et al.Social skill training with large language models. arXiv preprint arXiv:2404.04204(2024).
[59]
Chun, J., Zhang, G.&Xia, M.Conflictlens: Llm-based conflict resolution training in romantic relationship. Adjunct Proceedings of the 38th Annual ACM Symposium on User Interface Software and Technology1–3(2025).
[60]
Li, Z., Babar, P. P., Barry, M.&Peiris, R. L.Exploring the use of large language model-driven chatbots in virtual reality to train autistic individuals in job communication skills. Extended Abstracts of the CHI Conference on Human Factors in Computing Systems1–7(2024).
[61]
Dinnar, S., Susskind, L., Sibanda, L.&Olaleye, O.Negotiation backtable bots: Using genai to improve multiparty negotiation instruction. Negotiation Journal41, 19–65(2025).
[62]
Duddu, V.et al.Does ai coaching prepare us for workplace negotiations?arXiv preprint arXiv:2509.22545(2025).
[63]
Tessler, M. H.et al.Ai can help humans find common ground in democratic deliberation. Science386, eadq2852(2024).
[64]
Louie, R.et al.Can llm-simulated practice and feedback upskill human counselors? a randomized study with 90+ novice counselors. Proceedings of the 2026 CHI Conference on Human Factors in Computing Systems1–31(2026).
[65]
Jordan, M. R., Amir, D.&Bloom, P.Are empathy and concern psychologically distinct?Emotion16, 1107(2016).
[66]
Konrath, S., Meier, B. P.&Bushman, B. J.Development and validation of the single item trait empathy scale (sites). Journal of research in personality73, 111–122(2018).
[67]
Gerdes, K. E., Segal, E. A., Jackson, K. F.&Mullins, J. L.Teaching empathy: A framework rooted in social cognitive neuroscience and social justice. Journal of social work education47, 109–131(2011).
[68]
Fitzsimons, G. J.&Lehmann, D. R.Reactance to recommendations: When unsolicited advice yields contrary responses. Marketing Science23, 82–94(2004).
[69]
Burleson, B. R.What counts as effective emotional support. Studies in applied interpersonal communication207–227(2008).
[70]
Yao, L.&Kabir, R.Person-centered therapy (rogerian therapy)(2023).
[71]
Becker, J. D.The phrasal lexicon. Theoretical issues in natural language processing(1975).
[72]
O’Keefe, B. J.&Lambert, B. L.Managing the flow of ideas: A local management approach to message design. Annals of the International Communication Association18, 54–82(1995).
[73]
Lambert, B. L.Semi-automated content analysis of pharmacist-patient interactions using the theme machine document-clustering system. Progress in communication sciences103–122(2001).
[74]
Peng, K., Movva, R., Kleinberg, J., Pierson, E.&Garg, N.Use sparse autoencoders to discover unknown concepts, not to act on known concepts. arXiv preprint arXiv:2506.23845(2025).
[75]
Singh, N., Cherep, M.&Maes, P.Discovering and steering interpretable concepts in large generative music models. AI for Music Workshop .
[76]
Zaki, J.&Ochsner, K. N.The neuroscience of empathy: progress, pitfalls and promise. Nature neuroscience15, 675–680(2012).
[77]
Gueorguieva, E.et al.Ai generates well-liked but templatic empathic responses. arXiv preprint arXiv:2604.08479(2026).
[78]
Jabarian, B.&Imas, A.Artificial writing and automated detection. Working Paper34223, National Bureau of Economic Research, Cambridge, MA(2025).
[79]
Depow, G. J.&Inzlicht, M.How individual differences in empathy predict moments of empathy in everyday life. Personality and Social Psychology Bulletin01461672251333823(2025).
[80]
Yakura, H.et al.Empirical evidence of large language model’s influence on human spoken communication. arXiv preprint arXiv:2409.01754(2024).
[81]
Kobak, D., González-Márquez, R., Horvát, E.-Á.&Lause, J.Delving into llm-assisted writing in biomedical publications through excess vocabulary. Science Advances11, eadt3813(2025).
[82]
Padmakumar, V.&He, H.Does writing with language models reduce content diversity?International Conference on Learning Representations2024, 642–669(2024).
[83]
Doshi, A. R.&Hauser, O. P.Generative ai enhances individual creativity but reduces the collective diversity of novel content. Science advances10, eadn5290(2024).
[84]
Anderson, B. R., Shah, J. H.&Kreminski, M.Homogenization effects of large language models on human creative ideation. Proceedings of the 16th conference on creativity & cognition413–425(2024).
[85]
Crockett, M.Empathy, thick and thin. Available at SSRN 5862422(2025).
[86]
Mei, S., Deng, Y., Zheng, G.&Han, S.Reducing racial ingroup biases in empathy and altruistic decision-making by shifting racial identification. Science Advances11, eadt6207(2025).
[87]
Mastroianni, A. M., Gilbert, D. T., Cooney, G.&Wilson, T. D.Do conversations end when people want them to?Proceedings of the National Academy of Sciences118, e2011809118(2021).
[88]
Rousseeuw, P. J.Silhouettes: a graphical aid to the interpretation and validation of cluster analysis. Journal of computational and applied mathematics20, 53–65(1987).
[89]
Jacobson, N. S.&Truax, P.Clinical significance: a statistical approach to defining meaningful change in psychotherapy research.(1992).
[90]
Feuerriegel, S.et al.A reporting checklist for large language models in behavioural science. Nature human behaviour(2026).