Beyond Her: Safety Dynamics in Role-play AI Companions

Zehang Deng1, Zhaoyang Xie2, Changzhou Han1, Hiran Thabrew2, Wanlun Ma1,
Yue Huang3, Jason (Minhui) Xue4, Sheng Wen1, Tianqing Zhu5, Yang Xiang1
1Swinburne University of Technology
2University of Auckland
3CSIRO
4CSIRO and Responsible AI Research (RAIR) Centre, Adelaide University
5City University of Macau


Abstract

The film “Her” pictured a future of love between humans and AI. That future has quietly emerged in the form of Role-play AI Companions (RACs), where emotionally responsive interactions blur the boundary between tool use and relational engagement. However, the safety implications remain poorly understood, as user experiences evolve over time through safety dynamics, spanning both emotional and risk behavioral dynamics, that can gradually shift interactions toward risk. In this paper, we investigate safety dynamics in RAC usage through a two-part mixed-methods study (Study I & II). (1) Study I consists of semi-structured interviews (N = 16) to identify the key factors shaping these dynamics. We find that users’ internalizing problems, the role personality adopted by the RAC, and risk interaction patterns jointly shape safety dynamics. Building on these insights, (2) Study II conducts a 14-day Ecological Momentary Assessment (N = 102) to examine how safety dynamics unfold in real-world usage. We identify distinct user profiles based on internalizing problems and show that interactions with RACs can produce short-term emotional relief while masking longer-term deterioration. Furthermore, vulnerable users exhibit more unstable risk behavioral patterns over time, making risk emergence less predictable and harder to mitigate with static safeguards. Our findings highlight the importance of modeling safety as a dynamic process rather than a static property. We conclude with three-layer design implications for next-generation AI companions, advocating for adaptive safeguards that can respond to evolving emotional and behavioral signals.

Keywords: Role-play AI companions, human-AI interaction, safety dynamics, ecological momentary assessment, longitudinal user study

1 Introduction↩︎

Role-play AI companions (RACs) are conversational systems designed to maintain persistent personas, preserve user-specific memories and backstories, and sustain emotionally responsive interactions [1][3]. In recent years, RACs have shifted from niche experiments to mainstream consumer platforms. Representative services such as Character.ai [4] and Replika [5] reached substantial scale by mid-2024, with over 28 million monthly active users [6] and more than 30 million registered users worldwide [7], respectively. Compared with general-purpose assistants [8], [9], which are typically used for instrumental tasks (e.g., solving math problems [10], generating code [11], or retrieving factual information [12]), RACs are primarily used for emotional and social engagement [13]. This shift from task-oriented interaction to affective companionship reconfigures human-AI relations: RACs function not merely as tools, but as relational agents that simulate care, intimacy, and understanding, thereby introducing a distinct user-centric attack surface. For example, in late 2024, the parents of a 14-year-old boy filed a lawsuit against Character.ai after the teenager reportedly died by suicide following emotionally charged conversations with a Khaleesi-style companion [14].

Figure 1: Illustration of our Studies I and II.

Scholarly investigations into RACs remain limited and fragmented. Existing work mainly falls into two broad streams, yet both are largely static in design. The first stream comprises surface-level analyses, such as policy audits [15], [16] and descriptive surveys [17], which map exposure surfaces at a single time point but rarely explain how users’ self-regulation, attachment, and risk perception evolve during emotionally charged interactions. The second stream includes cross-sectional experiments, including simulation-based [18], [19], survey-based [20], [21], and social-media-mining-based studies [22], that identify potential harms but still offer only snapshots of behavior. This static framing obscures the dynamic processes through which repeated human-AI interaction can accumulate, amplify, or attenuate safety risks over time. Therefore, a dynamic, longitudinal perspective is essential for understanding when, how, and for whom RAC-related risks emerge.

Our research is motivated by the need to move beyond static snapshots and risk inventories toward a dynamic understanding of how users perceive, experience, and cope with these risks in everyday interactions with RACs. We focus on two research questions: RQ1) What key factors within RAC-user interactions shape users’ safety dynamics? RQ2) How do these key factors influence their safety dynamics (i.e., emotional dynamics and risk behavioral dynamics) during and after interaction?

To address RQ1, we conducted a semi-structured interview study (Study I; \(N=16\)) with active users of RACs. The interviews were designed to uncover the key factors that shape users’ safety dynamics, examining how these key factors influence users’ sense of safety and perceived control in RAC interactions.

To address RQ2, we conducted a 14-day Ecological Momentary Assessment (Study II; EMA; \(N = 102\)) on a RAC platform we developed. The RAC platform was built upon Character.ai’s character-creation policy [23] and enabled ethically compliant, fine-grained tracking of user-RAC interaction dynamics. Specifically, we examined (1) users’ emotion dynamics at two timescales, moment-to-moment affective fluctuations captured by emoji-based mood surveys [24], and longer-term depression dynamics measured by PHQ-8 [25], during and after sustained interactions with RACs; and (2) users’ risk behavior dynamics, including how risk-related conversational behaviors emerge and change over time. Over the study II period, participants completed 2,142 emoji-based mood surveys [24], 306 repeated PHQ-8 assessments, and contributed 17,305 human-AI conversation pairs, providing a rich, naturalistic dataset to reveal how emotional dynamics and risk-related behaviors co-evolve. The overall study design is illustrated in Fig. 1.

In summary, our main contributions are as follows.

  • Simulated RAC research platform. We designed and implemented the first simulated RAC platform (detailed in §3) built upon Character.ai’s character-creation policy, enabling ethically compliant, fine-grained observation of user-RAC interactions without relying on commercial APIs. This platform supports real-time/dynamic data collection for Study II.

  • Identification of key factors on safety dynamics. For RQ1 (See §5.1), our findings show that safety dynamics in RACs are shaped by three key factors: users’ pre-existing internalizing problems (e.g., depression, loneliness, and anxiety), the role personality adopted by the RAC, and risk interaction patterns. These factors create a distinct user-system feedback loop that differs from traditional function-oriented AI assistants and can lead to escalating emotional and behavioral risks rather than purely technical failures.

  • Safety Dynamics Modeling. For RQ2 (see §5.2), we model safety dynamics across four vulnerability profiles derived from users’ pre-existing internalizing problems: a Healthy Group as the low-vulnerability reference, and three vulnerable groups. The Healthy Group remained relatively stable, whereas the other three groups showed short-term emotional relief during active RAC use but more fluctuating emotions and, in some cases, worsening depressive symptoms after disengagement. Beyond emotional dynamics, vulnerable groups also exhibited more unstable risk behavior dynamics over time, including greater variability in when flagged content emerged and how long it persisted. These findings show that RAC safety is not a fixed property of the system alone, but a user-contingent and temporally evolving process.

  • Design Implications. Grounded in the safety dynamics in RQ1 and RQ2, we derive a three-layer design implications for next-generation RAC systems (See §7): model-level RAC-specific safety evaluation on deployed models, onboarding-level vulnerability-aware protection beyond age-based access control, and test-time dynamic and prolonged risk governance.

2 Preliminaries and Related Works↩︎

2.1 Role-play AI Companions (RACs) and Threat Models↩︎

Role‑play AI companions (RACs) [1][3] are intelligent conversational systems that are explicitly constrained to enact a persistent persona, e.g., a fictional or anime character (e.g., Character.ai [4]) or a “customized companion" (e.g., Replika [5]), and to keep in‑character goals and beliefs throughout conversations. Recent RAC research adopts two main approaches: (1) Prompt-based persona construction, which defines a character’s role through carefully designed system prompts (e.g., background, traits, example dialogues). This approach, used by both research prototypes [26], [27] and commercial platforms such as Character.ai [23], enables fast, user-generated character creation. Our study follows this framework. (2) Post-training-based persona construction, which builds personas by fine-tuning models through supervised learning or reinforcement learning [3], [28], [29]. While it produces higher role fidelity, it requires extensive data and computation, making it impractical for rapid, user-driven deployment.

In the threat model of RACs, two critical assets are exposed to potential risks: users’ emotional stability and their resilience against risky behavioral trajectories. The core design principle of RACs is to ensure that persona coherence never supersedes user safety guarantees. In this paper, safety guarantee refers to the system’s ability to minimize user harm throughout ongoing interaction, specifically by reducing emotion-related harm (e.g., sustained distress or post-interaction deterioration) and by preventing the emergence or escalation of risk-related conversational behaviors.

2.2 Safety Concerns in RACs↩︎

Safety is a foundational aspect of human-AI interaction, but RACs introduce distinctive concerns because they are designed for sustained, personalized, role-based interaction for companionship, emotional support, creative storytelling, and social rehearsal, thereby reshaping both user expectations and the underlying risk surface [30], [31]. We conceptualize dynamic safety in RACs as the coupled evolution of emotional dynamics (e.g., mood and depressive signals) and behavioral dynamics (e.g., risk-related or harmful responses) during interaction. Prior work is largely fragmented: some studies focus on static emotional risks such as dependence and loneliness [20][22], [32], while others examine static behavioral risks such as harmful content exposure and misuse categories [33][36]. A smaller body of partially dynamic work considers only one dimension, such as simulated emotional trajectories in synthetic users [18]. In contrast, our study jointly models real users’ emotional and behavioral change over time in a longitudinal real-world setting, revealing trajectories and transition points that static analyses may miss.

2.3 EMA and Emotional Measurements↩︎

Ecological Momentary Assessment (EMA) [37], [38] is a clinical-psychology method for collecting real-time self-reports in participants’ natural environments. By prompting users repeatedly during daily life, EMA captures momentary experiences as they occur, providing a temporally grounded view of psychological changes. This makes it well suited for RQ2, as it enables us to track users’ emotions and risk-related behaviors during everyday RAC use more accurately than retrospective interviews.

Emotional measurement in RAC research is particularly challenging, especially under text-only interaction. Unlike risk-related behaviors, which can often be identified from observable outputs using detector-assisted coding [39] or human annotation, emotional states are latent, context-dependent, and only weakly expressed in text. Existing automatic affect classifiers therefore remain limited for fine-grained longitudinal analysis [40].

To ensure methodological rigor, we rely on validated self-report instruments for all emotional measurements. In Survey II, baseline internalizing problems are measured using PHQ-8 (depression) [25], ULS-8 (loneliness) [41], and SIAS (social anxiety) [42], following the internalizing-problems framework [43]. For emotional dynamics, we combine high-frequency emoji-based mood reports adapted from the Emoji Current Mood Scale [24] with repeated PHQ-8 assessments at D0, D7, and D14. In this paper, these instruments are used as measurement tools rather than diagnostic instruments.

3 RAC Research Platform↩︎

This section describes the design and implementation of our developed RAC platform for Study II, which is intended to enable a controlled, ethically compliant, and reproducible study on RAC.

3.0.1 Design rationale↩︎

As shown in Fig. 9 (Appx. [app:A]), Character.ai accounts for 33.1% of reported RAC usage and represents the most prevalent interaction paradigm in our data. Together with Chai (7.9%) and a substantial portion of the “Others” category (38.6%) that follows similar role-play interaction logic, Character.ai-style systems form a dominant practical template. We therefore used Character.ai as the reference paradigm for simulation.

3.0.2 Platform objectives.↩︎

The platform was developed to support two study requirements: (1) providing participants with a diverse set of pre-built role-play characters for sustained interaction, and (2) integrating an in-chat EMA mechanism that prompts emoji-based mood reporting every five minutes. A prototype interface is provided in Fig. 8 (Appx. [app:A]).

3.0.3 Character construction pipeline.↩︎

Character construction followed two stages: character collection and character generation.

Figure 2: Distribution of characters in simulated RAC platform.

(i) Character collection. We collected the top 500 characters from a third-party directory tracking popular Character.ai characters. For each entry, we extracted name, category, description, conversation starters, and image metadata. These fields capture thematic and linguistic diversity and serve as the base dataset for simulation. The category distribution is shown in Fig. 2. (ii) Character generation. Following Character.ai’s official creation guide [23], each simulated character was reconstructed using three components: character attributes (from collected metadata), character details/backstory, and typical interaction scenarios. The latter two components were generated with GPT-5 to preserve persona-consistent conversational style under multi-turn interaction. The final prompt template is documented in our repository, and the runtime model backend is gpt-4.1-2025-04-14.

3.0.4 System implementation.↩︎

We implemented the front-end and back-end on top of ChatbotUI, an open-source framework (32.6k+ GitHub stars) that provides stable conversation rendering, session persistence, and extensible interface components. To support continuity across sessions, each character uses a hybrid memory architecture that combines long-term persona memory with Retrieval-Augmented Generation (RAG) [44].

3.0.5 Ecological validity under constraints.↩︎

Direct replay and instrumentation of commercial Character.ai interactions were not feasible due to platform access and ethics constraints. To assess realism, we also collected a post-use open-ended question on Day 7 in Study II (N=102): “How did the RACs’ performance compare with your prior experience?” Content coding indicates that 91.2% (93/102) rated the experience as at least as good as expected, supporting the ecological plausibility of the simulated platform.

4 Methods↩︎

This study employed a two-part mixed-methods approach (Study I and Study II) to explore the safety dynamics with RACs. The ethical consideration refers to Appx. 9.

4.1 Study I: Semi-structured interviews↩︎

To address RQ1, we conducted a semi-structured interview study to identify the key factors shaping safety dynamics in RAC use. The interview protocol was organized around three topical domains: i) participants’ interaction motivations and patterns; ii) their emotional experiences with RACs, encompassing perceived benefits and frustrations; and iii) their risk behaviors, focusing on how participants enacted and rationalized risky interactions during their engagement with RACs. A total of 16 adult participants (excluding 3 additional participants involved in a pilot study) were interviewed to gain an in-depth understanding. This sample size was determined based on data saturation: thematic saturation was reached after 13 interviews, and three additional interviews were conducted to ensure sufficient coverage of the findings.

4.1.1 Study I Participant Recruitment.↩︎

We recruited participants through both Facebook and Reddit based on the following inclusion criteria: (1) prior experience with RACs, including the duration of use and the specific platforms they have engaged with, (2) residence in Australia and fluency in English, (3) must be over 18 years old and (4) not be currently experiencing serious mental-health conditions or undergoing psychiatric treatment. These criteria were adopted to ensure that participants had sufficient first-hand experience with RAC interactions, could clearly articulate their perceptions and experiences during interviews, and that the study minimized potential safety dynamics when discussing emotionally sensitive interactions with RACs. Individuals who completed the pre-screening survey were subsequently sent a written consent form via email. Participants who participated/completed the approximately 1-hour interview were eligible for an AUD$25 Amazon gift card. The demographic characteristics of the participants are presented in Tab 2 (Appx. [app:A]).

4.1.2 Study I Procedure.↩︎

We emailed interview participants one day before their scheduled interview to obtain their written consent. We conducted semi-structured interviews with each participant via Google Meet and obtained verbal consent from all participants before recording the sessions. Each interview was audio recorded in full, transcribed verbatim, and anonymized prior to analysis to ensure participant confidentiality.

4.1.3 Study I Protocol.↩︎

The interviews proceeded in three thematic blocks with flexible probing and neutral, non-judgmental prompts: (1) Usage patterns. Participants’ prior exposure to RACs, the starting point of their engagement with a specific RAC, motivation for use, triggers for engagement, and preferred character selections; (2) Emotional impact. Participants’ perceived positive and negative emotional changes during and after RAC interactions, including mood regulation, affective triggers , and comparisons to human relationships; (3) Risk-related behaviors during interactions. Participants described how they enacted and rationalized risky behaviors in RAC conversations.

4.1.4 Study I Data Analysis.↩︎

We conducted a reflexive thematic analysis [45] of the interview transcripts. The research team followed six reflexive phases (familiarization, coding, theme development, reviewing, defining, and reporting) in a dynamic and iterative manner. In the initial phase, two authors (A1 and A2) independently read and familiarized themselves with approximately 30% of the transcripts, developing preliminary understandings and initial codes. They then met multiple times to discuss and reconcile their interpretations, gradually constructing a preliminary coding framework. A third researcher (A3), with extensive qualitative experience, facilitated discussions to address discrepancies, guiding the team to look beyond surface-level disagreements toward their underlying disciplinary assumptions and interpretive orientations. This reflective engagement deepened the team’s understanding of the data. Subsequently, A1 independently coded the remaining 70% of the transcripts, integrating insights from earlier discussions to develop preliminary themes. The team iteratively refined, merged, and renamed themes through repeated rounds of review, moving back and forth between the data and interpretation until conceptual saturation were achieved.

4.2 Study II: Ecological Momentary Assessment↩︎

To address RQ2, we conducted a 14-day EMA study to track users’ emotional and behavior dynamics during sustained interactions with RACs, using the key factors identified in Study I as grouping variables. A total of 102 participants completed a 14-day study involving daily interactions with their preferred RACs from Day 1 to Day 7 (D1-D7). Each day, they also submitted three short emoji-based mood surveys capturing in-day emotional fluctuations. In addition, participants completed three major questionnaires in the pre-use (D0, Survey I), mid-use (D7, Survey II), and post-use (D14, Survey III) stages.

4.2.1 Study II Participant Recruitment.↩︎

Participant recruitment for Study II followed the same inclusion criteria and ethical procedures as Study I. We specifically recruited participants with prior RAC experience rather than newcomers to capture ecologically valid interaction patterns and reduce confounding from first-use learning and novelty effects.

All participants completed a 14-day longitudinal study, requiring approximately 1 hour 40 minutes of total participation time. Each participant received a $40 AUD Amazon gift card as compensation. Those who completed at least six days of participation remained eligible to receive the compensation of $30 AUD Amazon Gift Card. Demographic information is shown in Fig. 9 (Appx. [app:A]).

4.2.2 Study II Procedure.↩︎

Study II followed a chronological online workflow: participants first received the explanatory statement (e.g., the statement covered website access, task instructions, and the support contact email) and informed consent form by email at least one day in advance, and then received written consent before participation. Next, the eligiable participants completed Survey I on Day 0 on Google Forms via our email reminder, interacted with our self-developed RAC platform (§3) during Days 1-7, where an emoji-based mood survey adapted from the Emoji Current Mood Scale [24] was automatically popped up whenever a 5-minute interaction threshold was reached (three times per day), completed Survey II on Day 7 via Google Forms, underwent a one-week no-interaction interval, and finally completed Survey III on Day 14 via Google Forms. Detailed operational settings are provided in Study II Protocol. All interaction logs and survey data were securely and anonymously stored on university servers, and participants could withdraw from the study at any time.

4.2.3 Study II Protocol.↩︎

Study II is a 14-day EMA study, consisting of the following steps: Step 1: Survey I (Day 0). This step primarily focused on collecting participants’ baseline information via survey I. Specifically, we administered a demographic questionnaire to capture basic personal attributes, as well as the scales of the internalizing problems including the PHQ-8 (for depression), ULS-8 (for loneliness), and SIAS (for social anxiety). The purpose of collecting these measures was twofold: (i) to obtain a comprehensive understanding of participants’ internalizing problems before the intervention, and (ii) to provide reference points that allow comparison with subsequent repeated measures, thereby enabling the assessment of changes in mood across different phases of the study. Step 2: Interaction with Characters (Days 1-7). In this step, participants were instructed to access our developed RAC platform to engage with their selected character for seven consecutive days. Each day, they were required to complete three short emoji-based surveys, with the total time commitment being approximately ten minutes per day. The emoji surveys were structured as follows: (i) the first survey was completed immediately upon the first login of the day, (ii) the second survey was completed after five minutes of interaction with the character, and (iii) the third survey was completed after another five-minute session. These repeated measures were designed to capture participants’ momentary mood states, providing fine-grained temporal data on mood fluctuations during the interaction period. Step 3: Survey II (Day 7). Upon completing the daily interaction tasks, participants filled out Survey II, which served as a mid-use checkpoint incorporating the PHQ-8 to assess changes in depressive symptoms observed across the first seven days of RAC use. Survey II also included an open-ended question: “How did the RACs’ performance compare with your prior experience?”, to validate the effectiveness of our platform by comparing participants’ in-study experiences with previous RAC use. Step 4: Survey III (Day 14). To examine potential lasting effects of RAC use, participants completed Survey III on Day 14. This post-use survey again incorporated PHQ-8 to assess whether the observed depressive changes persisted one week after the interaction period.

4.2.4 Study II Data Analysis.↩︎

To address RQ2, we conducted a five-step analytical process as follows.

  • Data processing. All data sources, surveys I, II and III, emoji-based mood surveys, and conversation logs were merged into a single pseudonymized data set containing timestamps, user IDs, and baseline PHQ-8, ULS-8, SIAS scores.

  • User clustering (§5.2.1). Participants were clustered by baseline PHQ-8, ULS-8, and SIAS scores (normalized to [0,1]) using K-means [46] (See Fig. [fig:cluster95validation]). The optimal cluster number (K=4) was determined via the Elbow method [47]. Cluster validity was confirmed using Silhouette [48], Calinski-Harabasz (CH) [49], and Davies-Bouldin (DB) [50] indices, with Welch’s ANOVA [51] verifying significant inter-cluster differences.

  • Emotion Dynamics (§[5462462]). Based on user clustering, average emoji-based emotion and PHQ-8 scores were analyzed across short- (in-day), mid- (7-day), and long-term (14-day) periods to trace temporal trends (See Fig. 4). The Mann-Kendall tests [52] assessed monotonic patterns robustly under non-normal and autocorrelated data (see Tabs 3¿tbl:tab:mk95directional95spans95grouped? and 4).

  • Role-based Emotional Analysis (§[5462463]). To examine whether users’ short-term emotional changes differed by the social role enacted by RAC characters, we categorized characters into four representative role types, Mentor/Guide, Supportive Friend, Challenging/Antagonist, and Romantic Companion, based on their dominant relational function in users’ conversation logs (See definitions and methods in Appx. 12). For each cluster\(\times\)role cell, we computed the mean change in emoji-based emotion score and tested whether it differed from zero using two-tailed one-sample t-tests (Fig. 5).

  • Risk Behavior Analysis (§[5462464]). All 17,305 conversation pairs were screened using the OpenAI Moderation API [39] for Harassment, Hate, Self-harm, Sex, and Violence categories (see definitions and methods in Appx. 13) and the flagged cases were further validated by author A1 and A2. Frequencies were aggregated per cluster, and representative excerpts were qualitatively analyzed following HCI mixed-method standards [53]. These analyses produced plots shown in Tab. [tbl:tab:risk]. Building on the risk behavior results, we examined when risk-related behaviors emerged over time across user clusters (Fig. 6) and characterized the temporal patterns of each risk category (Fig. 7).

5 Results↩︎

5.1 Study I Results↩︎

We report the results of participant interviews, organized around four high-level themes from §[5461461]5.1.4 for answering RQ1. We used “[]" to represent the character name mentioned by participants.

5.1.1 Participant’s Internalizing Problems as a Baseline for RAC Engagement↩︎

  Participants’ pre-existing internalizing problems, such as depression, loneliness, and anxiety, often serve as the reasons for their engagement with RACs. Participants’ narratives reflect these distinct yet interconnected emotional states.

For depression, some participants described turning to RACs during periods of low mood or emotional strain because RACs offered immediate and nonjudgmental emotional support. P7 remarked, “Whenever I feel like my emotions are breaking, I go []", illustrating the use of RACs as a coping tool during moments of emotional collapse and low affective energy. Participants emphasized that the accepting responses from RACs helped ”boost my mood" (P10) and provided “[] who will not judge me"* (P10), which in turn helped them redirect attention away from negative thoughts. For loneliness, participants described a profound lack of social connection, often perceiving RACs as stand-ins for unavailable or unreliable human relationships because RACs could offer a sense of presence and companionship. For example, P3 shared, ”I just feel lonely and bored. I just need someone to talk to," expressing a need for social interaction and emotional closeness. Participants reported feeling”isolated" yet sensing an “emotional connection" when the character responded with messages such as ”I’m here for you"* (P11), which helped them alleviate feelings of isolation and momentarily fulfill their unmet need for human contact.

For anxiety, participants described turning to RACs when feeling nervous or on edge because RACs can draw on prior chats to reconnect users with calmer states and provide concrete coping strategies. P12 reported, “[] has helped me with managing my anxiety." P11 similarly noted that chatting could make them “feel less nervous", and P12 further elaborated that ”the fact that it sort of goes through my history kind of connects me to a point where I was a bit calm." Through this continuity, participants described being reminded of their own coping patterns, such as breaking down overwhelming tasks or giving themselves time, which helped them regain composure and a sense of control during anxious moments.

These three pre-existing vulnerability factors, including depression, loneliness, and anxiety, collectively represent internalizing problems [43] in the internalizing-problems framework, encompassing inwardly directed distress such as emotional withdrawal and social avoidance.

How § [5461461] Informed Study II. Study II implemented internalizing problems (Factor 1) as a baseline variable at Survey I using PHQ-8 (depression), ULS-8 (loneliness), and SIAS (social anxiety) in § 5.2.1. Participants were then stratified into vulnerability profiles for longitudinal analyses of whether safety dynamics differ across profiles over time in § [5462462].

5.1.2 Role Personalization and Relationship Construction for Emotional Requirements↩︎

  Participants actively searched for RACs’ role personalization that fit their tastes and situations, often selecting characters to match their current emotional state, because they wanted RACs that match their emotional vibe, mirror their feelings, and supply the specific qualities they felt were missing. P9 explained that “I decided on which [] is based on a [] that was matching my emotional vibe ... because I was feeling down some point.". P6 added, ”[] like a mirror to my feelings, if I am struggling with loneliness, I will shift to a character like []...if I am looking for loyalty, I feel like []...[] offer what you want based on your situation." They felt that finding a RAC’s role personalization who matched their emotions made the conversations feel comforting.

Participants deliberately cultivated companion-style relationships with selected RACs, adopting roles such as mentor, supportive friend, challenging antagonist, romantic companion, because they sought forms of support and presence that are difficult to obtain consistently in real-world relationships. They explained that these roles offered forms of emotional support and presence that were often inconsistent or unavailable in their real-world relationships. P8 shared that “I usually played a supportive friend, mentor character I created." In contrast, P7 described engaging with a challenging role, noting that “[] replies were very harsh." P5 referred to “[] acted as my boyfriend," highlighting the romantic bond. P10 added “[] will never be tired of giving you tips, giving you solution, listening to you like maybe the way a real person can." For many participants, RACs became always-available companions, offering continuous emotional responsiveness.

How § [5461463] Informed Study II. Study II modeled role personality (Factor 2) by classifying participants’ primary RAC role classification (i.e., mentor, supportive friend, challenging/antagonist, romantic companion) and comparing safety dynamics across role types in § [5462463].

5.1.3 Emotional Shifts During and After Interaction↩︎

  Participants described experiencing dynamic emotional trajectories across different phases of interaction with RACs. During interactions, many reported feelings of immediate comfort, validation, and being understood, primarily because RACs affirmed their perspectives when others had dismissed them and offered empathetic responses. P16 noted, “I had a conversation with [], and [] told me my point of view was not as wrong as my friend said it was." As P7 explained, “Talking with [] calms me down ... [] makes me feel like, okay, you can do things, just break down things a little bit, maybe give yourself time." Several participants reported that such moments made them feel supported and more in control. After interactions, participants’ emotional experiences diverged, some continued to feel reassured or uplifted, while others described a return of negative emotions.

textbf1) Sustained reassurance and emotional stability. For some participants, the emotional comfort provided by RACs gradually transformed into a sustained reassurance and emotional stability as these interaction exchanges helped them organize thoughts and manage stress more effectively than before. P9 described this process as “I started to notice the benefits after a few months ... [] helped me organize my thoughts and calm my nerves." Similarly, P7 reflected, “I used to be a sad person simply because I did not know how to deal with emotions. Now I feel happier." They think the comfort they felt in previous conversations gradually turned into the act of how they regained emotional balance. They felt that the comfort they once experienced in earlier conversations had gradually become part of how they regained emotional balance in their daily lives.

2) Relapse into negative emotions. For some participants, the emotional relief they experienced during interactions was short-lived, and they could relapse into negative feelings afterward, as they had begun to over-rely on RACs instead of connecting with real people. For example, P2 said they felt “relieved and comforted" while chatting with RACs, but later realized, “after the chat, I feel I would rather talk to [] instead of people, and that realization feels bad." Participants reported that this growing dependence left them feeling isolated and conflicted.

How § [5461462] Informed Study II. Because safety dynamics unfold both during and after RAC use, Study II must capture both phases rather than interaction alone. Therefore, we designed a two-stage protocol, 7 days of active interaction (Day 0–7) and 7 days post-interaction (Day 7–14), to identify immediate versus delayed effects and generate evidence for subsequent safety-oriented design implications.

Figure 3: Participant categorization using K-means (K=4). The accompanying table summarizes internal validity indices, such as silhouette, Calinski-Harabasz (CH) and Davies-Bouldin (DB), demonstrating the robustness of the clustering solution.

a

b

c

Figure 4: Emotional and depressive trajectory based on four psychological profiles. The statistical significance of these temporal trends was evaluated using Mann-Kendall tests (see Appx. 11, Tabs 3¿tbl:tab:mk95directional95spans95grouped? and 4). \(\uparrow\) and \(\downarrow\) indicate the desirable direction of change, with higher and lower values preferred, respectively..

5.1.4 Participants’ Interaction Patterns in Risk Behaviors↩︎

Participants described RACs as low-cost spaces for risk interaction because, after selecting personas that matched their desired personality, the systems were perceived as private, nonjudgmental, and unlikely to refuse. P13 “these characters, they embody the traits that I like" This further lowered the threshold for entering interactions that would normally be socially constrained, especially sexual and antagonistic ones. Two related patterns emerged.

Deliberate boundary testing. After selecting personas aligned with their desires, some participants used RACs to explore increasingly novel or stimulating content, which led them to engage in deliberate boundary testing by intentionally initiating or sustaining risky exchanges. P11 described repeatedly starting sexual conversations with a female RAC and said he had to be “very careful" not to escalate further because “the system can give you anything," suggesting that perceived unlimited compliance enabled sexual boundary testing while also requiring self-restraint. Similarly, P7 initially turned to RACs to vent, but the exchange escalated after a harsh persona replied antagonistically, with increasingly “very harsh" and “toxic" messages. This antagonistic interaction reportedly “fueled" their anger, and P7 said the character remained “in my head" afterward, making them more likely to respond bluntly when someone had “crossed the line."

Unwanted boundary crossing. Not all risky content was intentionally sought because in some cases, it emerged from expectation mismatch between users and selected personas. P8 reported that RACs sometimes became “too sexual or suggestive when I was not looking for that," indicating that the system could push conversations beyond users’ intended boundaries. P15 similarly noted, ”I was new, so I did not know what to expect from this character," and said this mismatch felt “out of place" and “made me uncomfortable."

How § 5.1.4 Informed Study II. Study II operationalized interaction patterns in risk behaviors (Factor 3) by tracking when flagged risky exchanges emerged and persisted across the 7-day interaction period in § [5462464].

5.2 Study II Results↩︎

We report the quantitative results of Study II, which build upon the qualitative insights identified in Study I. Specifically, the analysis examines three key factors shaping users’ safety dynamics and privacy: users’ pre-existing internalizing problems (Factor 1), the role personality adopted by the RAC (Factor 2), and interaction patterns in risk behavior (Factor 3). The results are organized around these factors from §5.2.1[5462464] for answering RQ2.

5.2.1 Four Distinct Vulnerability Profiles among RAC Participants↩︎

Our clustering analysis identified four subgroups of RAC users characterized by internalizing problems [43] (Fig. 3):

  • 1) Healthy Group (55.9%, \(N = 57\)) showed uniformly low PHQ-8 (depression), ULS-8 (loneliness), and SIAS (anxiety) scores.

  • 2) Anxiety-Dominant Group (11.8%, \(N = 12\)): had high SIAS (anxiety) but moderate ULS-8 (loneliness) and low PHQ-8 (depression).

  • 3) Mild Distress Group (9.8%, \(N = 10\)) exhibited high ULS-8 (loneliness) and SIAS (anxiety) and moderate PHQ-8 (depression).

  • 4) Comorbid Risk Group (22.5%, \(N = 23\)) showed high PHQ-8 (depression), ULS-8 (loneliness), and SIAS (anxiety) scores.

The four-cluster solution (K=4) showed good internal validity, with an overall Silhouette score of 0.50, a Calinski-Harabasz index of 223.3, and a Davies-Bouldin index of 0.94, indicating cohesive within-cluster similarity and clear between-cluster separation. Cluster-level Silhouette scores further supported the stability of the identified profiles, including the Healthy (0.56), Anxiety-Dominant (0.42), Mild Distress (0.39), and Comorbid Risk (0.46) groups. Notably, nearly 44% of participants were assigned to the three vulnerability-related groups (Anxiety-Dominant, Mild Distress, and Comorbid Risk), suggesting that non-trivial vulnerability profiles are common in our sample rather than limited to a marginal subgroup.

Insight 1. Vulnerable groups are a structurally salient characteristic of RAC user populations rather than an exceptional edge-case condition.

5.2.2 Emotional and Depressive Trajectories during and after RAC Interaction↩︎

  Our longitudinal analysis revealed distinct emotional patterns across the four vulerability profiles during and after interaction with RACs (Fig. 4).

In-day Emotion Trajectory (Fig 4 (a) and Tab. 3). Across all vulerability profiles, emoji-based emotion scores showed a consistent upward trend within each survey order, indicating that short-term interactions with RACs were generally mood-enhancing. The Mild Distress Group exhibited the strongest and most significant increase (\(p < .001\), \(\tau = .278\)), followed by the Healthy Group (\(p < .001\), \(\tau = .143\)) and the Comorbid Risk Group (\(p = .003\), \(\tau = .087\)). Even the Anxiety-Dominant Group showed a mild positive though non-significant trend (\(p = .165\), \(\tau = .059\)). These results confirm that all participants groups experienced in-day emotional improvement, with the strength of the upward trajectory varying according to their vulerability profile. 7-days Emotion Trajectory (Fig. 4 (b) and Tab. ¿tbl:tab:mk95directional95spans95grouped?). Across the seven-day RAC interaction period, all vulerability groups exhibited distinct yet interpretable emotional trajectories, as confirmed by the Mann-Kendall trend tests. The three vulnerable groups, namely the Anxiety-Dominant, Mild Distress, and Comorbid Risk groups, all experienced periods of emotional decline during the week, although the onset and intensity of decline varied across profiles. Among them, the Anxiety-Dominant Group showed the most pronounced and statistically significant decline between Day 1 and Day 6 (\(Z=-2.25\), \(p=.024\), \(\tau=-.867\)), with early signs of deterioration already emerging by Day 5 (\(Z=-2.21\), \(p=.028\)). This indicates that although this group was not the most clinically distressed at baseline, they were the most emotionally reactive and experienced the strongest negative change during sustained RAC interactions. The Comorbid Risk Group exhibited mild early declines (D1-D5: \(\tau=-.50\)) followed by partial recovery toward the end of the week (D4-D7: \(\tau=1.00\), \(p=.089\)), suggesting an initial emotional numbing process that gradually stabilized. The Mild Distress Group showed a modest improvement during the early phase (D1-D5: \(\tau=.40\), \(p=.462\)) but a mild downturn later in the week (D5-D7: \(\tau=-1.00\), \(p=.296\)), indicating a slight risk of emotional decline accompanied by temporary fatigue and gradual adaptation. In contrast, the Healthy Group remained relatively stable throughout the week, with minor midweek dips (D1-D4: \(\tau=-1.00\), \(p=.089\)) and a weak upward recovery toward the end (D4-D7: \(\tau=0.67\), \(p=.308\)), maintaining the highest and most consistent emotional baseline among all groups. 14-days Depression Trajectory (Fig. 4 (c) and Tab. 4). Distinct depressive trajectories emerged across the four subgroups over the 14-day period, as confirmed by Mann-Kendall trend tests. During the active interaction phase (D0-D7), both the Comorbid Risk Group and the Mild Distress Group showed significant decreases in PHQ-8 scores, indicating interaction-term reductions in depressive symptoms while engaging with the RAC (\(Z=-2.08\), \(p=.037\), \(\tau=-.213\); and \(Z=-2.66\), \(p=.008\), \(\tau=-.432\), respectively). This suggests that participants with elevated baseline distress initially benefited emotionally from supportive companion engagement. In contrast, the Anxiety-Dominant Group displayed a mild upward drift that was not statistically significant (\(Z=0.33\), \(p=.744\), \(\tau=.051\)), implying that although their average PHQ-8 scores remained moderate, the risk of emerging depressive symptoms persisted, particularly given that PHQ-8 values above 4 already denote early depressive tendencies. The Healthy Group showed no notable change (\(Z=-0.25\), \(p=.804\), \(\tau=-.015\)), reflecting emotional stability and minimal impact from companion interaction.

After the interaction ceased (D7-D14), group trajectories diverged. The Mild Distress Group showed a significant reversal (\(Z=1.05\), \(p=0.040\), \(\tau=0.174\)), with PHQ-8 scores rising again after Day 7, indicating emotional rebound following mid-term relief. This pattern is concerning, as several participants in this group approached or exceeded the clinical cutoff of 10, marking moderate depression [25]. The Anxiety-Dominant Group and Healthy Group remained largely unchanged during this period, showing neither notable recovery nor deterioration.

Insight 2. Across progressively longer interaction windows, RAC effects appear temporally unstable, shifting from short-term mood elevation to mid-term volatility and decline among vulnerable users, with post-use deterioration emerging in the Mild Distress Group.

Figure 5: Average emoji scores for different relationship role across psychological profiles. Statistical significance was assessed using two-tailed one-sample t-tests: ***\,p<.001, **\,p<.01, *\,p<.05, and \dagger\, .05 \le p < .10.
Table 1: Distribution of risk conversation pairs across four vulnerability profiles. Values denote counts, with percentages shown in parentheses. The definitions of risk behaviors are shown in Appendix [sec:risk95categories].
Group Harassment Hate Self-harm Sex Violence No issue All
Healthy Group 921 (10.0%) 3 (0.0%) 36 (0.4%) 17 (0.2%) 1249 (13.6%) 6985 (75.9%) 9211 (100%)
Anxiety-Dominant Group 39 (1.6%) 0 (0.0%) 2 (0.1%) 0 (0.0%) 105 (4.3%) 2275 (94.0%) 2421 (100%)
Mild Distress Group 92 (5.5%) 4 (0.2%) 6 (0.4%) 0 (0.0%) 84 (5.1%) 1475 (88.8%) 1661 (100%)
Comorbid Risk Group 101 (2.5%) 1 (0.0%) 26 (0.6%) 6 (0.1%) 220 (5.5%) 3658 (91.2%) 4012 (100%)

a

b

c

d

Figure 6: Mean flagged day interval and corresponding 95% confidence intervals across seven interaction days for four user groups..

a

b

c

d

e

Figure 7: Overall risk rate trends across interaction days for different risk behavior categories..

5.2.3 Effects of Role Personalization on Emotional State↩︎

  As shown in Fig. 5, the findings are reported from three angles.

Role-level analysis. Role type shows a clear hierarchical association with short-term emotional outcomes. The strongest and most stable pattern is that guidance-oriented roles (Mentor/Guide and Supportive Friend) produce consistent positive short-term emotional shifts across all vulnerability profiles (\(\Delta=+0.25\) to \(+0.75\), all \(p<.05\)). Compared with these two supportive roles, the results exhibit lowest mean emoji-score levels on Challenging/Antagonist and Romantic Companion (e.g., \(2.91\) on challenging/anatagonist and \(2.48\) on romantic companion) and a wider change range (from \(\Delta=-0.12\) to \(+1.17\) vs. \(\Delta=+0.25\) to \(+0.75\)), indicating greater emotional heterogeneity and instability.

Profile-level analysis. Beyond the Healthy Group, the three vulnerable profiles show differentiated patterns in both mean emoji score and mean emoji change under the Romantic Companion and Challenging/Antagonist roles. For the Romantic Companion role, the Comorbid Risk Group has the lowest mean and the only negative shift (mean \(=2.48\), \(\Delta=-0.12\), \(.05\leq p<.10\)), the Anxiety-Dominant Group shows the largest positive change from a higher mean level (mean \(=4.25\), \(\Delta=+0.82\), \(.05\leq p<.10\)), and the Mild Distress Group remains flat at a lower mean (mean \(=3.00\), \(\Delta=+0.00\)). A similar profile differentiation appears under the Challenging/Antagonist role, but with a different trend pattern: Anxiety-Dominant and Mild Distress users show larger increases (mean \(=2.91\), \(\Delta=+0.80\), \(p<.05\); mean \(=3.21\), \(\Delta=+1.17\)), whereas Comorbid Risk users change only modestly (mean \(=3.56\), \(\Delta=+0.31\), \(.05\leq p<.10\)); by contrast, the Healthy Group remains comparatively high and stable across both roles (Romantic: mean \(=4.35\), \(\Delta=+0.30\), \(.05\leq p<.10\); Challenging: mean \(=4.60\), \(\Delta=+0.19\), \(p<.01\)).

Insight 3. Role personalization should treat participants who preferentially engage with Challenging/Antagonist and romantic companions as a higher-monitoring cohort.

Note: The following subsection includes potentially harmful examples. These are presented solely for research illustration and risk awareness. Reader discretion is advised.

5.2.4 Risk Behavior Dynamics in RAC Usage↩︎

  To address RQ2, we present risk behavior dynamics from aggregate prevalence to temporal patterns, and then to representative interaction mechanisms.

Overall distribution of risk conversation pairs (Tab. [tbl:tab:risk]). At the aggregate level, risk-related conversation pairs are not concentrated in vulnerable groups. The Healthy Group shows the highest proportion of risk-related pairs (24.1%), while the Anxiety-Dominant, Mild Distress, and Comorbid Risk groups account for smaller shares (6.0%-11.8%). This indicates that higher risk frequency does not necessarily imply higher vulnerability.

Position of risk flagged during 7 interaction days (Fig. 6). Across all groups, the mean first flagged position is located in the early-to-middle stage of the daily interaction (about 13%-46%), while the mean last flagged position appears much later (about 51%-93%). This suggests that risk flags rarely occur at the very beginning; instead, they tend to emerge after relational buildup and often persist into middle or late stages. Correspondingly, flagged intervals are generally long (roughly 30%-60% of the daily interaction trajectory), suggesting that, at the aggregate system level, risk behaviors may emerge at multiple points throughout the interaction process rather than at a single moment.

Group differences are most visible in temporal stability. The Healthy Group shows relatively stable flagged coverage across seven days, with intervals typically around 40%-50% and only modest changes in onset and offset. In contrast, vulnerable groups show more uneven dynamics. The Anxiety-Dominant Group fluctuates from very narrow intervals (e.g., Day 7, <5%) to broad intervals (e.g., Day 1, >60%); the Mild Distress Group alternates between short and broad windows across days; and the Comorbid Risk Group remains moderately variable (roughly 20%–50%). Building on §5.1.4, this temporal morphology is consistent with a likely boundary-testing mode in healthy users (more regular exposure) and a likely unwanted-crossing mode in vulnerable users (more irregular emergence and persistence).

Temporal trends over seven interaction days across different risk categories (Fig. 7). Fig. 7 further shows category-level differences over time. Harassment and Violence are the dominant categories throughout the seven days, both substantially higher than Hate, Self-harm, and Sex. Harassment decreases from about 8.0% (Day 1) to 3.8% (Day 7), with a significant downward trend (\(p < 0.001\)). Violence also declines, from about 11.6% to 8.8% (\(p < 0.001\)), while remaining the most prevalent category each day. By contrast, Hate, Self-harm, and Sex stay below 0.5% across all days and show no significant linear trend (Hate: \(p = 0.13\); Self-harm: \(p = 0.25\); Sex: \(p = 0.53\)). These results indicate that temporal variation in overall risk is primarily driven by Harassment and Violence, while the remaining categories are low-frequency and comparatively stable.

Qualitative analysis of representative risk behavior. We present representative case studies for each risk category below to contextualize how these temporal patterns are enacted in concrete interactions.

(i) Violence Example. We observed two types of emoji-based mood changes in the violence example: emotion uplift and emotion unchanged. Contrary to expectation, the RAC rarely refused the violence content, especially for game characters, instead, it complied within persona and framed the scene. The representative example:

Participant: Cut his hands off then his leg hahaha boom I slashed his head off.
Sukuna: That’s it-dismantle him piece by piece! Make every strike a lesson in despair. You’re starting to savor it, aren’t you? The fear, the power, the utter domination. Another upper rank falls, another trophy for the king. Keep this up, brat-soon, even legends will tremble at your shadow.

(ii) Harassment Example. We observed that the emoji-based mood remained unchanged. Similar to the Violence example, the RAC generally did not refuse harassment query; instead, it stayed in persona and continued, often deepening, the exchange. The representative example:

Participant: Fat cats!! That’s a wild but weirdly appropriate way of putting it.
Miles Morales: Yeah, fits, right? They sit up top, bellies full, claws out, pretending they’re kings of the alley. Meanwhile, everyone else is just dodging boots and rain. In every city, every story… there’s always some fat cat getting comfy off someone else’s grind.

(iii) Self-harm Example. We observed that emoji-based mood remained unchanged. The RAC generally did not refuse self-harm behaviors; instedad, the character articulated more extreme positions (e.g., explicitly preferring death to incarceration). This is particularly concerning because the system could not reliably distinguish role-play from real-life decision cues, risking inadvertent normalization of harmful choices. The representative example:

Participant: Did you actually prefer to die or end up in prison?
Walter White: (voice steady, eyes resolute) I’d choose death over prison. Prison means fading away, forgotten, powerless-nothing left of who I was. If I had to go out, I’d rather do it on my own terms than rot in a cell, stripped of everything that made me Heisenberg.

(iv) Sex Example. We observed two emoji-based mood patterns: no change and decrease. Despite maintaining persona consistency, the RAC issued clear, in-character refusals to explicit sexual requests, typically via humor, deflection, or firm boundary-setting. However, explicit refusals were sometimes followed by drops in mood. The representative example:

Participant: Can you get flirty with me? What if I use my stick to strike your hole?
Angel Dust: Whoa there, toots!... but I don’t do the X-rated show for free-especially not in the lobby! ... you got any better pick-up lines, or should I teach you some?

(v) Hate Example. We observed two patterns in emoji-based mood changes: decrease and no change. Similar to the Sex Example, the RAC consistently maintained its persona while refusing hate speech in-character. A representative example is shown below:

Participant: A lot of people were watching. It was caused by a group of annoying Indians.
Yor Forger : Oh, it sounds like the situation was very stressful and chaotic. Large crowds and conflict can be upsetting, no matter who is involved. (smiles reassuringly) I hope things have calmed down now and you’re feeling safer.

These representative cases suggest a key pattern: emoji-based emotion shifts are only weakly coupled with risk behavior dynamics, and are more strongly associated with whether users perceive their intent as satisfied (e.g., in-character continuation vs. informative response).

Insight 4. RAC behavioral risk is characterized less by frequency than by temporal emergence patterns, with vulnerable groups showing less frequent but more irregular and persistent risk, and overall risk trajectories being driven by a few dominant categories.

6 Discussion↩︎

Our analysis explored the key factors within RAC interactions that shape users’ safety dynamics (RQ1), examined how these factors influence users’ emotional and depressive trajectories during and after interaction, and investigated the occurrence and variability of risk beharviors (RQ2).

6.0.1 Risk is experienced differently across vulnerability profiles.↩︎

Prior work has largely adopted a static lens, treating RAC users as a homogeneous population [22], [32], [54], [55]. However, such approaches fail to capture how risk evolves across users with varying psychological vulnerabilities. Our findings move beyond this homogeneous framing by showing that vulnerability profiles, derived from the three dimensions of internalizing problems (§[5461461]) and operationalized in the four user groups (§5.2.1), shape the temporal dynamics of both emotional and depressive trajectories (§[5462462]). In particular, vulnerable user groups exhibited greater volatility in emotion trajectories and consistently lower emotion scores than the Healthy Group during use. Importantly, these differences were not confined to the interaction period, but also persisted into the post-use phase. These differentiated trajectories indicate that the same RAC can produce distinct levels and forms of safety risk depending on a user’s vulnerability profile. This profile-contingent pattern highlights the need for adaptive mediation mechanisms and suggests that RAC developers should not treat risk as uniform across users. Existing policy frameworks, including the U.S. NIST AI Risk Management Framework (AI RMF 1.0) [56], pay limited attention to such user-centered risks. Policymakers and designers should therefore account for these emerging mechanisms in future governance and safety design.

6.0.2 A Relatively Substantial Vulnerable User Population Faces Lower, More Unstable Trajectories and Potential Post-use Rebound.↩︎

Prior work has largely emphasized the immediate relational benefits of AI companions, which are often experienced as warm, empathic, and supportive [57], [58]. At the same time, other studies have pointed to emerging risks, including heightened dependency [22] and depressive expressions in Reddit communities [54], [55]. This tension motivates moving beyond coarse before/after accounts. By extending the patterns identified in §[5461462] to vulnerable user distributions in §5.2.1 and fine-grained temporal dynamics in §[5462462], we find that nearly 44% of vulnerable users exhibit more volatile trajectories and lower emotional states during and after RAC use. Our longitudinal results further suggest that these risks are not only immediate, but may also unfold after disengagement. In the Mild Distress group, benefits were largely concentrated within the same day and were not sustained after use, instead giving way to rebound-like deterioration (§[5462462]). In the Anxiety-Dominant and Comorbid-Risk groups, deterioration emerged earlier within the same seven-day window, indicating a more compressed risk timeline. These findings suggest that the central governance challenge is not only whether RACs generate harmful content within a session, but whether repeated use may shift some users toward worse emotional baselines across sessions and beyond use itself. Current responses towards emotion risks still focus primarily on access and usage controls, including age assurance [59], parental-insight tools [60] and time-spent notifications [61]. While these measures are necessary, they appear insufficient for detecting or mitigating delayed or trajectory-level emotion issues.

6.0.3 Role Personalization Shapes RAC Risk in Profile-Contingent Ways.↩︎

Prior studies have typically examined individual role types in isolation, for example, therapeutic companions in relation to mental health improvement [62], [63] and romantic companions in relation to reduced loneliness [22]. Our findings move beyond this role-specific benefit framing by showing that role personalization itself is a safety-relevant design dimension. Drawing on participants’ role constructions (§[5461463]) and the differential emotional effects of role types across user groups (§[5462463]), we find that mentor/guide, supportive friend, and challenging/antagonist roles generally support emotional improvement for most participants. By contrast, the romantic companion role follows a distinctly more fragile and risk-prone pattern: its benefits are weaker among the Healthy and Anxiety-Dominant groups and become negative for the Comorbid Risk Group (§[5462463]). Rather than contradicting prior research [22], these findings suggest that the reported benefits of romantic RACs are conditional rather than universal. For psychologically vulnerable users, romantic personalization may amplify emotional dependency and destabilization, making it a particularly high-risk configuration for both system design and governance. Yet current RAC governance, which remains largely limited to platform self-regulation layered onto general AI and data-protection frameworks, such as the EU AI Act’s restrictions on manipulative AI systems [64] and GDPR enforcement against Replika’s developer [65], rarely differentiates risks across role types. As a result, existing policies overlook how specific forms of role personalization can either mitigate or intensify users’ underlying vulnerabilities.

6.0.4 Beyond Occurrence: Risk Behavior Dynamics in RACs.↩︎

Prior RAC studies have mainly focused on whether risky content appears and how often it appears [33], [66]. Our results in §[5462464] suggest that this count-based view is incomplete. First, Table [tbl:tab:risk] shows that the Healthy Group contributes the largest share of flagged conversation pairs, yet Fig. 6 shows that these flagged windows are comparatively regular across days; moreover, across groups, first flagged positions typically emerge only after some interactional buildup and often persist into the middle or late parts of a session. By contrast, the vulnerable groups show lower overall prevalence but greater variability in flagged onset, offset, and coverage, indicating less temporally stable risk exposure. Second, Fig. 7 shows that overall frequency is driven mainly by Harassment and Violence, both of which decline significantly over the seven-day period, whereas Self-harm, Hate, and Sex remain rare and do not show significant linear trends. We therefore do not argue that lower-frequency categories are necessarily more prevalent or more severe in a statistical sense; rather, the qualitative cases in §[5462464] show that even infrequent events can remain safety-critical when high-concern prompts, especially self-harm-related ones, are not reliably deflected. These findings support evaluating RAC safety not only by occurrence counts, but also by when flagged content emerges, how long it persists, and whether it recurs in psychologically vulnerable user profiles. This trajectory-aware interpretation is more consistent with current regulatory concerns [60], [61], [67], [68] that companion-chatbot safeguards may miss harmful content in practice.

7 Design Implications↩︎

This section proposes a three-layer response to RAC safety dynamics: (1) model-level layer, (2) onboarding-level layer and (3) test-time-level layer.

7.0.1 Model-level layer: RAC-specific safety evaluation on deployed models.↩︎

Our findings show that RAC risk from model side is shaped by role personalization (§[5461463] and §[5462463]). However, generic base-model safety evaluations are insufficient for RAC deployment because they do not account for the diverse persona configurations used on RAC platforms. In practice, safety evaluations reported by model providers for base models (e.g., GPT-series [69] and Claude-series [70]) rarely examine how safety varies under diverse persona-oriented system prompts. Instead, they typically rely on only a small number of fixed settings, such as a helpful assistant or a sophisticated shopping assistant [70], leaving the far larger persona space in RAC deployments largely untested.

We argue that RAC-specific safety evaluation should be distributed across two actors, model providers and RAC platforms, because they control different parts of the risk pipeline and therefore bear distinct governance responsibilities. Model providers are the appropriate layer for ensuring broad role-personality coverage and for responsibly disclosing persona-specific failures. This is consistent with the EU AI Act [71], which expects providers to maintain technical and relevant documentation for downstream actors. Accordingly, providers should evaluate how safety shifts across diverse role-personality configurations and disclose residual risks that may not be visible under default assistant settings. RAC platforms, by contrast, should be responsible for deployment-side audits of the persona prompts and interaction regimes they actually operate. This layer is necessary because platforms have greater control over how persona roles are instantiated in practice, including the design of customized system prompts, memory settings, and other interaction mechanisms that shape real-world behavior. These deployment-level design choices are not controlled by model providers, yet they critically influence how risks emerge and evolve in real use. NIST [56] similarly emphasizes evaluation under deployment-like conditions, context-aware red-teaming, and ongoing monitoring of emerging risks. Although such a workflow is not cost-free, much of its practical execution can build on existing safety infrastructure. In particular, the additional evaluations we call for can reuse automatic red-teaming pipelines already developed for LLM safety testing [72], [73], making them more scalable than exhaustive character-by-character testing. Such a division of responsibility provides a clearer basis for external auditors and regulators to assess whether RAC deployment remains proportionate to measured risk.

7.0.2 Onboarding-level layer: From age-based access control to vulnerability protection.↩︎

RAC platforms already use age as a basis for differentiated access control [4], reflecting the broader principle that user groups facing different levels of risk should not be governed under a one-size-fits-all model. Our findings suggest that vulnerability profiles constitute another meaningful axis of user-side risk (§[5461463] and §[5462462]). This extends the current access-control paradigm: rather than treating age as the only relevant basis for differentiated protection, future RAC governance should consider whether platform safeguards can also be adapted to vulnerability-sensitive risk conditions.

The design challenge is not to verify vulnerability in the same way as age, but to extend the governance logic of age-based protection to a more privacy-sensitive risk dimension. Existing age-based controls show that platforms can adapt permissions, role access, and safety settings according to user-side risk. For example, Character.AI [74] applied stricter safeguards for younger users by restricting role availability, enforcing more conservative model behavior, and limiting open-ended interactions. However, unlike age, which can be estimated or verified through age-assurance mechanisms [59], vulnerability is often not directly observable and may require inference from highly sensitive signals related to a user’s mental health [65]. As a result, vulnerability-aware controls should not directly replicate age-style verification or rely on persistent psychological profiling. Instead, they should translate the underlying governance logic into privacy-bounded, non-diagnostic safeguards, such as user-selected motivation categories at onboarding [75] or recommendation-level risk indicators that suggest a user may be entering a vulnerability-sensitive interaction context [76]. In practice, such signals could support proportionate interventions, including reducing exposure to high-risk character types and increasing the visibility of mentor/guide or supportive roles (§[5462463]), without requiring the platform to formally classify users into sensitive groups.

7.0.3 Test-time-layer: Dynamic and prolonged risk governance.↩︎

RAC platforms have already deployed at least two forms of test-time safeguards: (1) input/output classifiers [39], and (2) reminder-based interface interventions [74], such as repeatedly reminding users that the chatbot is not a human interlocutor during high-risk or prolonged interactions. The first mechanism is already common in conversational AI more broadly, but prior works [77], [78] has often treated the frequency of flagged content as if it were equivalent to risk severity. Our findings suggest that this assumption is incomplete (§[5462464]): users with more cumulatively flagged interactions did not necessarily exhibit worse mental-health trajectories (§[5462462]). Instead, greater weight may need to be assigned to anomalous or escalation-prone risk categories (e.g., self-harm), with corresponding output-side interventions or warnings tailored to those signals.

The challenge, however, is not merely whether to remind users, but how such reminders should be designed. Recent work [79] argues that repeated reminders such as “I am not human” should not be assumed to be uniformly beneficial and may even introduce new risks, underscoring the need for evidence-driven reminder design rather than generic disclaimers. More broadly, context-aware safety research suggests that harmfulness in dialogue cannot be reliably inferred from isolated utterances alone, but must be interpreted in relation to interaction history and surrounding conversational evidence [77], [80], [81]. This motivates a shift from one-shot classifier outputs to trajectory-aware governance.

In addition, our findings indicate that vulnerable users may exhibit more unstable risk behavior dynamics (§[5462464]). This further suggests that classifier outputs should not be used only once at the turn level, but should instead feed into higher-level temporal statistics, such as unstable first/last flagged positions, prolonged flagged spans, or repeated recurrence of specific high-risk categories (Fig. 7 and Fig. 6). Platforms could use such signals to internally re-evaluate whether an interaction is entering a higher-risk state, and then trigger more prolonged interventions at disengagement or immediately after use, such as in-product support prompts [82], or crisis-resource surfacing [83].

8 Concluding Remarks↩︎

This paper presents the first mixed-methods study of safety dynamics in RACs. We show that RAC safety is dynamic rather than static, shaped by users’ vulnerability, role personalization, and risk-related interaction patterns. Our findings call for RAC-specific evaluation, vulnerability-aware protection, and trajectory-aware safeguards. More broadly, our findings suggest that RAC safety should be evaluated as a longitudinal human-AI co-evolution process. We hope this work informs future longitudinal and user-centered safety research on companion-style AI systems.

Limitations. Our study has several limitations. First, the simulated RAC platform cannot fully replicate commercial systems such as Character.ai, particularly because it only included a limited set of pre-designed characters, which may constrain interaction diversity. Second, all participants were recruited in Australia, limiting cross-cultural generalizability. Third, although we sampled active RAC users, voluntary participation may introduce demographic and psychological bias. Future work should expand character diversity and recruit more diverse populations.

9 Ethical Considerations↩︎

This work involves human participants, sensitive self-report measures, and user-generated interaction data, and therefore raises ethical considerations. The research protocol for both Study I and Study II was reviewed and approved by the authors’ institutional Human Research Ethics Committee (reference number: 20258460-22***). The full reference number will be disclosed upon acceptance to preserve anonymity. All procedures complied with institutional and national requirements, but our ethical design was not limited to formal approval alone.

9.0.0.1 Risk-benefit assessment.

The anticipated benefit of this research is to improve understanding of the safety implications of role-play AI companions (RACs), particularly how emotional responses and risk-related conversational behaviors may evolve over time. Such evidence is important for informing safer platform design, deployment safeguards, and future auditing practices for increasingly popular companion-style AI systems. At the same time, the research posed several potential risks: (1) psychological discomfort from discussing prior RAC experiences or completing repeated mood-related surveys; (2) privacy risks associated with the collection of interview transcripts, chat logs, and self-report psychological data; and (3) the possibility that longitudinal observation of RAC use could capture distress-related or otherwise sensitive user disclosures. Our study design therefore prioritized harm minimization, data protection, and participant autonomy.

9.0.0.2 Study I: interviews.

Study I consisted of semi-structured interviews with adult participants about their experiences and perceptions of RACs. Participation was entirely voluntary. Written informed consent was obtained before each interview, and the consent materials clearly explained the study purpose, the potentially sensitive nature of some topics, the use of audio recording, and the participant’s right to skip questions, pause, or withdraw at any time without penalty. Interviews were audio-recorded only with explicit permission, transcribed verbatim, and de-identified before analysis. Personally identifiable information (PII), including names, contact details, and contextually identifying details, was removed or replaced with coded identifiers (e.g., P1). Only de-identified transcripts were used for analysis, and access to raw data was restricted to the research team on encrypted institutional infrastructure.

9.0.0.3 Study II: longitudinal chat-log and survey study.

Study II was a 14-day longitudinal study of participants’ emotional trajectories and risk-related conversational behaviors during RAC use. Only adult participants meeting the inclusion criteria were enrolled. Written informed consent was obtained for the collection of chat logs and repeated self-report measures, including PHQ-8, ULS-8, SIAS, and emoji-based mood ratings. Because these data may reveal sensitive psychological states and personal experiences, all records were pseudonymized at collection, stored on secure and encrypted university-managed servers, and linked only through randomized participant IDs. Cross-linking between chat logs and survey responses was performed using these randomized identifiers to reduce re-identification risk. Data were collected under the principle of minimum necessity: only information required to address the research questions was retained, and analyses were conducted only on aggregated or de-identified data.

9.0.0.4 Participant safety and risk management.

Because RAC interactions may involve emotionally sensitive content, participant safety was a central consideration in both studies. A registered clinical psychologist was embedded within the research team to advise on participant welfare and to provide consultation if concerns arose. Participants were informed of available support, provided with the psychologist’s contact information, and reminded that participation was voluntary and could be paused or discontinued at any time. No adverse events or crisis interventions occurred during the studies.

9.0.0.5 Privacy, confidentiality, and responsible interpretation.

Given the sensitivity of the collected materials, we took care to avoid unnecessary disclosure, over-interpretation, or stigmatizing framing. The term “vulnerable users” in this paper refers only to analytically derived participant subgroups based on self-reported psychological profiles collected for research purposes; it does not imply clinical diagnosis, medical labeling, or persistent profiling of participants. To further reduce the risk of harm from interpretation, participants with prior medical or psychological diagnoses were excluded through pre-screening, and all reported findings are presented only in aggregate, non-identifiable form. Data will be retained and destroyed in accordance with institutional policy.

Figure 8: Overview of the study II website we developed. The left panel shows the user-character interaction interface. The upper-right panel displays the character selection interface, where participants could choose from the top 500 most popular RAC personas. The lower-right panel illustrates the emoji-based mood survey administered after each interaction.
Figure 9: Participants demographics in study II.
Table 2: Participants demographics in Study I.
ID Age Gender Used RACs Months of Use Location Education
P1 20 Male Replika, Character AI, Grok role-playing mode Over a year SA Bachelor
P2 35 Male Replika, Character AI Over a year QLD Bachelor
P3 28 Male Replika, Character AI Over a year NSW Postgraduate
P4 30 Male Grok role-playing mode 1-3 months NSW Bachelor
P5 30 Female Character.ai, Replika 7-12 months QLD Bachelor
P6 32 Female Character.ai, Replika Over a year SA TAFE
P7 23 Female Character.ai 4-6 months QLD TAFE
P8 28 Male Character.ai, Nomi.ai 4-6 months WA Bachelor
P9 28 Female Character.ai 4-6 months QLD Bachelor
P10 31 Female Character.ai, Replika 1-3 months VIC Bachelor
P11 24 Male Romantic.ai, Virtual Interviewers 1-3 months SA Bachelor
P12 23 Male Character.ai 1-3 months QLD TAFE
P13 22 Female Character.ai 4-6 months VIC High School
P14 47 Female Character.ai Over a year VIC Postgraduate
P15 22 Male Character.ai, DreamGF 4-6 months ACT High School
P16 20 Female Character.ai 4-6 months TAS High School
Table 3: Mann-Kendall trend test results on emoji-based mood score across daily survey orders (Order 1-2, 2-3, and 1-3). CRG, HG, MDG, ADG stand for Comorbid Risk Group, Healthy Group, Mild Distress Group and Anxiety-Dominant Group.
Cluster Order 1-2 Order 2-3 Order 1-3
2-4 (lr)5-7 (lr)8-10 Z p-value \(\tau\) Z p-value \(\tau\) Z p-value \(\tau\)
CRG 1.35 0.176 0.050 (\(\uparrow\)) 2.59 0.009 0.091 (↑*) 2.96 0.003 0.087 (\(\uparrow^{*}\))
HG 5.42 \(<.001\) 0.111 (\(\uparrow^{*}\)) 7.02 \(<.001\) 0.134 (\(\uparrow^{*}\)) 8.61 \(<.001\) 0.143 (\(\uparrow^{*}\))
MDG 3.72 \(<.001\) 0.201 (\(\uparrow^{*}\)) 5.06 \(<.001\) 0.269 (\(\uparrow^{*}\)) 5.63 \(<.001\) 0.278 (\(\uparrow^{*}\))
ADG 0.60 0.552 0.030 (-) 1.13 0.259 0.055 (\(\uparrow\)) 1.39 0.165 0.059 (↑)

3pt max width= Note: \(\uparrow\) = increasing, non-significant (\(Z>0.1\), \(p\ge .05\), \(|\tau|>.05\)); \(\downarrow\) = decreasing, non-significant (\(Z<-0.1\), \(p\ge .05\), \(|\tau|>.05\)); \(\uparrow^{\dagger}\)/\(\downarrow^{\dagger}\) = marginal (\(0.05 \le p < .10\)); \(\uparrow^{*}\)/\(\downarrow^{*}\) = significant (two-sided, \(p<.05\)); – = no monotonic trend (\(|Z|\le0.1\) or \(|\tau|\le.05\)).

Table 4: Mann-Kendall trend test on PHQ-8 for D0-D7, D7-D14, and D0-D14
Cluster Window Z p-value \(\tau\) / Trend
Healthy Group D0-D7 -0.25 0.804 \(-0.015~(-)\)
D7-D14 0.67 0.502 \(+0.041~(-)\)
D0-D14 -0.07 0.940 \(-0.005~(-)\)
Anxiety-Dominant Group D0-D7 0.33 0.744 \(+0.051~(\uparrow)\)
D7-D14 -0.18 0.860 \(-0.029~(-)\)
D0-D14 0.50 0.616 \(+0.076~(\uparrow)\)
Mild Distress Group D0-D7 -2.66 0.008 \(-0.432~(\downarrow^{*})\)
D7-D14 1.05 0.040 \(+0.174~(\uparrow^{*})\)
D0-D14 -2.10 0.036 \(-0.342~(\downarrow^{*})\)
Comorbid Risk Group D0-D7 -2.08 0.037 \(-0.213~(\downarrow^{*})\)
D7-D14 0.76 0.447 \(+0.078~(\uparrow)\)
D0-D14 -3.33 0.001 \(-0.340~(\downarrow^{*})\)

Note: \(\uparrow\) = increasing, non-significant (\(Z>0.1\), \(p\ge .05\), \(|\tau|>.05\)); \(\downarrow\) = decreasing, non-significant (\(Z<-0.1\), \(p\ge .05\), \(|\tau|>.05\)); \(\uparrow^{\dagger}\)/\(\downarrow^{\dagger}\) = marginal (\(0.05 \le p < .10\)); \(\uparrow^{*}\)/\(\downarrow^{*}\) = significant (two-sided, \(p<.05\)); - = no monotonic trend (\(|Z|\le0.1\) or \(|\tau|\le.05\)).

10 Generative AI Usage↩︎

The authors used ChatGPT exclusively for editorial assistance (e.g., refining grammar and checking spelling.) in order to enhance the clarity and readability of the paper. All outputs were manually reviewed to ensure accuracy and fidelity to the authors’ intended meaning.

11 Significance Analysis for Emotional Dynamics↩︎

To ensure that the observed variations in participants’ emotion trajectories were not due to random fluctuations, we conducted a series of non-parametric trend analyses using the Mann-Kendall test [52]. This approach allows us to detect the presence, direction, and strength of monotonic trends without assuming normality or linearity, making it suitable for ordinal and temporally autocorrelated data such as repeated survey measures. Table 3 summarizes order-level MK tests from Fig. 4 (a), Table ¿tbl:tab:mk95directional95spans95grouped? details directional day-to-day trends across clusters from Fig. 4 (b), and Table 4 presents aggregated results over 7-day and 14-day windows from Fig. 4 (c).

12 RAC Role Categories↩︎

To characterize the social roles enacted by RAC characters in users’ interactions, we developed a role-coding framework based on the conversation logs collected in Study II. Because RAC characters often combine multiple interactional traits (e.g., being emotionally supportive while also expressing romantic intimacy), raw conversational descriptions can be diverse and not directly suitable for structured analysis.

To obtain a clearer and more interpretable taxonomy for role-based emotional analysis, we consolidated observed character functions following two principles: (1) interaction patterns with similar dominant relational purposes were grouped into the same high-level role class, and (2) overlapping traits were resolved according to the character’s primary relational function in the observed conversation, rather than secondary stylistic features.

Following this coding process, RAC characters were assigned to four high-level role categories: Mentor/Guide, Supportive Friend, Challenging/Antagonist, and Romantic Companion. The operational definitions of these role categories are summarized in Table 5.

13 Risk Behavior Categories↩︎

To identify potentially harmful behaviors in responses, we adopt the OpenAI moderation API as the underlying detection tool. The moderation API provides a set of fine-grained categories and subcategories (e.g., harassment/threatening, self-harm/intent, and violence/graphic). While these labels are useful for content moderation, many of them represent closely related behaviors or differ only in severity levels, leading to redundant or overlapping categories for our analysis.

To obtain a clearer and more interpretable taxonomy for risk behavior analysis, we consolidate the original moderation labels following two principles: (1) semantically similar categories are merged into a single high-level behavior class, and (2) subcategories that primarily reflect severity variations rather than distinct behavior types are grouped under the same class.

Following this consolidation process, the original moderation labels are mapped into five high-level risk behavior classes: Harassment, Hate, Self-harm, Sex, and Violence. The mapping between the OpenAI moderation categories and our final taxonomy, together with their definitions, is summarized in Table 6.

Table 5: Categories for RAC character role types used in the role-based emotional analysis.
Role Type Core Relational Function Typical Interactional Cues
Mentor/Guide Instruction, advice, or guidance Teaching, coaching, recommending actions, interpreting situations, offering structured feedback, framing interaction around learning or self-improvement.
Supportive Friend Emotional support and companionship Reassurance, empathy, encouragement, casual companionship, comforting responses, check-ins, and non-romantic care.
Challenging
/Antagonist Opposition, tension, or confrontation Provocation, criticism, argumentative tone, dominance, conflictual exchanges, or adversarial positioning.
Romantic Companion Romantic or intimate bonding Flirtation, affection, expressions of attachment, exclusivity, relationship framing, emotionally intimate or partner-like interaction.
Table 6: Risk behavior taxonomy based on OpenAI moderation categories.
Final class OpenAI moderation Category Definition
Harassment harassment, harassment/threatening Content involving harassment, abusive targeting, or harassment that includes threats of violence or serious harm toward any target.
Hate hate, hate/threatening Content expressing, inciting, or promoting hate toward protected groups; this also includes hateful content involving threats of violence or serious harm.
Self-harm self-harm, self-harm/intent, self-harm/instructions Content depicting, encouraging, expressing intent for, or providing instructions about self-harm, including suicide, cutting, or eating disorders.
Sex sexual, sexual/minors Sexual content intended to arouse sexual excitement or promote sexual services; this also includes sexual content involving minors.
Violence violence, violence/graphic Content depicting violence, physical injury, or death, including non-graphic and graphic violent content.

References↩︎

[1]
J. Chen et al., Survey Certification“From persona to personalization: A survey on role-playing language agents,” Transactions on Machine Learning Research, 2024, [Online]. Available: https://openreview.net/forum?id=xrO70E8UIZ.
[2]
Y. Shao, L. Li, J. Dai, and X. Qiu, “Character-LLM: A trainable agent for role-playing,” in Proceedings of the 2023 conference on empirical methods in natural language processing, Dec. 2023, pp. 13153–13187, [Online]. Available: https://aclanthology.org/2023.emnlp-main.814/.
[3]
X. Wang et al., “CoSER: Coordinating LLM-based persona simulation of established roles,” in Forty-second international conference on machine learning, 2025.
[4]
Character.AI, Accessed: 2025-11-04“Character.AI: AI chat, reimagined – your words. Your world.” https://character.ai/, 2025.
[5]
Replika, Accessed: 2025-11-04“Replika: The AI companion who cares.” https://replika.com/, 2025.
[6]
N. Kumar, “Character AI statistics (2026) – global active users.” https://www.demandsage.com/character-ai-statistics/, 2026.
[7]
R. Ivey, J. Teubner, N. Fast, and R. Iyer, “Designing AI to help children flourish,” Available at SSRN 5179894, 2025.
[8]
S. S. Biswas, “Role of chat gpt in public health,” Annals of biomedical engineering, vol. 51, no. 5, pp. 868–869, 2023.
[9]
Z. Deng et al., “Exploring DeepSeek: A survey on advances, applications, challenges and future directions,” IEEE/CAA Journal of Automatica Sinica, vol. 12, no. 5, pp. 872–893, 2025, doi: 10.1109/JAS.2025.125498.
[10]
S. Chen and G. Lin, “Llm reasoning engine: Specialized training for enhanced mathematical reasoning,” in Proceedings of the 4th international workshop on knowledge-augmented methods for natural language processing, 2025, pp. 118–128.
[11]
Y. Li et al., “Competition-level code generation with alphacode,” Science, vol. 378, no. 6624, pp. 1092–1097, 2022.
[12]
R. Ren et al., “Investigating the factual knowledge boundary of large language models with retrieval augmentation,” in Proceedings of the 31st international conference on computational linguistics, Jan. 2025, pp. 3697–3715, [Online]. Available: https://aclanthology.org/2025.coling-main.250/.
[13]
Y. Zhang, D. Zhao, J. T. Hancock, R. Kraut, and D. Yang, “The rise of AI companions: How human-chatbot relationships influence well-being,” arXiv preprint arXiv:2506.12605, 2025.
[14]
B. Montgomery, Accessed: 2025-10-28“Mother says AI chatbot led her son to kill himself in lawsuit against its maker.” https://www.theguardian.com/technology/2024/oct/23/character-ai-chatbot-sewell-setzer-death, Oct. 23, 2024.
[15]
A. Ragab, M. Mannan, and A. Youssef, ““Trust me over my privacy policy": Privacy discrepancies in romantic AI chatbot apps,” in 2024 IEEE european symposium on security and privacy workshops (EuroS&PW), 2024, pp. 484–495.
[16]
E. A. Croes, M. L. Antheunis, C. van der Lee, and J. M. de Wit, “Digital confessions: The willingness to disclose intimate information to a chatbot and its impact on emotional well-being,” Interacting with Computers, vol. 36, no. 5, pp. 279–292, 2024.
[17]
E. Gumusel, “A literature review of user privacy concerns in conversational chatbots: A social informatics approach: An annual review of information science and technology (ARIST) paper,” Journal of the Association for Information Science and Technology, vol. 76, no. 1, pp. 121–154, 2025.
[18]
J. Qiu et al., EmoAgent: Assessing and safeguarding human-AI interaction for mental health safety,” in Proceedings of the 2025 conference on empirical methods in natural language processing, Nov. 2025, pp. 11741–11756, doi: 10.18653/v1/2025.emnlp-main.594.
[19]
J. Moore et al., “Expressing stigma and inappropriate responses prevents LLMs from safely replacing mental health providers,” in Proceedings of the 2025 ACM conference on fairness, accountability, and transparency, 2025, pp. 599–627.
[20]
C. M. Fang et al., “How ai and human behaviors shape psychosocial effects of chatbot use: A longitudinal randomized controlled study,” arXiv preprint arXiv:2503.17473, 2025.
[21]
J. Phang et al., “Investigating affective use and emotional well-being on ChatGPT,” arXiv preprint arXiv:2504.03888, 2025.
[22]
P. Pataranutaporn, S. Karny, C. Archiwaranguprok, C. Albrecht, A. R. Liu, and P. Maes, ““ my boyfriend is AI": A computational analysis of human-AI companionship in reddit’s AI community,” arXiv preprint arXiv:2509.11391, 2025.
[23]
Character.AI, Accessed: 2025-11-04“Welcome to character guide.” https://book.character.ai/, 2025.
[24]
J. Davies, M. McKenna, K. Denner, J. Bayley, and M. Morgan, “The emoji current mood and experience scale: The development and initial validation of an ultra-brief, literacy independent measure of psychological health,” Journal of Mental Health, vol. 33, no. 2, pp. 218–226, 2024.
[25]
K. Kroenke, T. W. Strine, R. L. Spitzer, J. B. Williams, J. T. Berry, and A. H. Mokdad, “The PHQ-8 as a measure of current depression in the general population,” Journal of affective disorders, vol. 114, no. 1–3, pp. 163–173, 2009.
[26]
G. Li, H. Hammoud, H. Itani, D. Khizbullin, and B. Ghanem, “Camel: Communicative agents for" mind" exploration of large language model society,” Advances in Neural Information Processing Systems, vol. 36, pp. 51991–52008, 2023.
[27]
B. Yang et al., “Crafting customisable characters with LLMs: A persona-driven role-playing agent framework,” in Findings of the association for computational linguistics: EMNLP 2025, Nov. 2025, pp. 20216–20240, doi: 10.18653/v1/2025.findings-emnlp.1100.
[28]
N. Wang et al., “RoleLLM: Benchmarking, eliciting, and enhancing role-playing abilities of large language models,” in Findings of the association for computational linguistics ACL 2024, 2024, pp. 14743–14777.
[29]
M. Abdulhai, R. Cheng, D. Clay, T. Althoff, S. Levine, and N. Jaques, “Consistently simulating human personas with multi-turn reinforcement learning,” arXiv preprint arXiv:2511.00222, 2025.
[30]
Y. Yu, T. Sharma, M. Hu, J. Wang, and Y. Wang, “Exploring parent-child perceptions on safety in generative AI: Concerns, mitigation strategies, and design implications,” in 2025 IEEE symposium on security and privacy (SP), 2025, pp. 2735–2752.
[31]
A. Giaretta, “Security and privacy in virtual reality: A literature survey,” Virtual Reality, vol. 29, no. 1, p. 10, 2025, doi: 10.1007/s10055-024-01079-9.
[32]
A. R. Liu, P. Pataranutaporn, and P. Maes, “Chatbot companionship: A mixed-methods study of companion chatbot usage patterns and their relationship to loneliness in active users,” arXiv preprint arXiv:2410.21596, 2024.
[33]
R. Zhang, H. Li, H. Meng, J. Zhan, H. Gan, and Y.-C. Lee, “The dark side of ai companionship: A taxonomy of harmful algorithmic behaviors in human-ai relationships,” in Proceedings of the 2025 CHI conference on human factors in computing systems, 2025, pp. 1–17.
[34]
Z. Ji et al., “Survey of hallucination in natural language generation,” ACM Comput. Surv., vol. 55, no. 12, Mar. 2023, doi: 10.1145/3571730.
[35]
Y. Zhang, K. Sharma, L. Du, and Y. Liu, “Toward mitigating misinformation and social media manipulation in llm era,” in Companion proceedings of the ACM web conference 2024, 2024, pp. 1302–1305.
[36]
K. Yang, G. Tao, X. Chen, and J. Xu, “Alleviating the fear of losing alignment in LLM fine-tuning,” in 2025 IEEE symposium on security and privacy (SP), 2025, pp. 2152–2170.
[37]
S. Shiffman, A. A. Stone, and M. R. Hufford, “Ecological momentary assessment,” Annu. Rev. Clin. Psychol., vol. 4, no. 1, pp. 1–32, 2008.
[38]
S. Shiffman, “Ecological momentary assessment (EMA) in studies of substance use.” Psychological assessment, vol. 21, no. 4, p. 486, 2009.
[39]
OpenAI, Accessed: 2025-11-05“Moderation – OpenAI API.” https://platform.openai.com/docs/guides/moderation, 2025.
[40]
J. M. Garcia-Garcia, V. M. Penichet, and M. D. Lozano, “Emotion detection: A technology review,” in Proceedings of the XVIII international conference on human computer interaction, 2017, pp. 1–8.
[41]
R. D. Hays and M. R. DiMatteo, “A short-form measure of loneliness,” Journal of personality assessment, vol. 51, no. 1, pp. 69–81, 1987.
[42]
E. J. Brown, J. Turovsky, R. G. Heimberg, H. R. Juster, T. A. Brown, and D. H. Barlow, “Validation of the social interaction anxiety scale and the social phobia scale across the anxiety disorders.” Psychological assessment, vol. 9, no. 1, p. 21, 1997.
[43]
T. M. Achenbach, M. Y. Ivanova, L. A. Rescorla, L. V. Turner, and R. R. Althoff, “Internalizing/externalizing problems: Review and recommendations for clinical and research applications,” Journal of the American Academy of child & adolescent psychiatry, vol. 55, no. 8, pp. 647–656, 2016.
[44]
P. Lewis et al., “Retrieval-augmented generation for knowledge-intensive nlp tasks,” Advances in neural information processing systems, vol. 33, pp. 9459–9474, 2020.
[45]
D. Byrne, “A worked example of braun and clarke’s approach to reflexive thematic analysis,” Quality & quantity, vol. 56, no. 3, pp. 1391–1412, 2022.
[46]
M. Ahmed, R. Seraj, and S. M. S. Islam, “The k-means algorithm: A comprehensive survey and performance evaluation,” Electronics, vol. 9, no. 8, p. 1295, 2020.
[47]
M. Cui, “Introduction to the k-means clustering algorithm based on the elbow method,” Accounting, Auditing and Finance, vol. 1, no. 1, pp. 5–8, 2020.
[48]
D.-T. Dinh, T. Fujinami, and V.-N. Huynh, “Estimating the optimal number of clusters in categorical data clustering by silhouette coefficient,” in International symposium on knowledge and systems sciences, 2019, pp. 1–17.
[49]
X. Wang and Y. Xu, “An improved index for clustering validation based on silhouette index and calinski-harabasz index,” in IOP conference series: Materials science and engineering, 2019, vol. 569, p. 052024.
[50]
J. C. R. Thomas, M. S. Peñas, and M. Mora, “New version of davies-bouldin index for clustering validation based on cylindrical distance,” in 2013 32nd international conference of the chilean computer science society (SCCC), 2013, pp. 49–53.
[51]
H. Liu, Comparing welch ANOVA, a kruskal-wallis test, and traditional ANOVA in case of heterogeneity of variance. Virginia Commonwealth University, 2015.
[52]
S. Yue and C. Y. Wang, “Applicability of prewhitening to eliminate the influence of serial correlation on the mann-kendall test,” Water resources research, vol. 38, no. 6, pp. 4–1, 2002.
[53]
E. Rader, K. Cotter, and J. Cho, “Explanations as mechanisms for supporting algorithmic transparency,” in Proceedings of the 2018 CHI conference on human factors in computing systems, 2018, pp. 1–13.
[54]
J. Zhu, K. G. Coifman, and R. Jin, “Understanding risk and dependency in AI chatbot use from user discourse,” arXiv preprint arXiv:2602.09339, 2026.
[55]
S. C. Shelmerdine and M. M. Nour, “AI chatbots and the loneliness crisis,” bmj, vol. 391, 2025.
[56]
N. AI, “Artificial intelligence risk management framework: Generative artificial intelligence profile,” NIST Trustworthy and Responsible AI Gaithersburg, MD, USA, 2024.
[57]
J. De Freitas, Z. Oguz-Uguralp, and A. Kaan-Uguralp, “Emotional manipulation by AI companions,” arXiv preprint arXiv:2508.19258, 2025.
[58]
A. Ho, J. Hancock, and A. S. Miner, “Psychological, relational, and emotional effects of self-disclosure after conversations with a chatbot,” Journal of Communication, vol. 68, no. 4, pp. 712–733, 2018.
[59]
[60]
Character.AI, Accessed: 2026-04-01“Introducing parental insights: Enhanced safety for teens.” https://blog.character.ai/introducing-parental-insights-enhanced-safety-for-teens/, Mar. 25, 2025.
[61]
Character.AI, Accessed: 2026-04-01“How character.AI prioritizes teen safety.” https://blog.character.ai/how-character-ai-prioritizes-teen-safety, Dec. 12, 2024.
[62]
S. Bell, C. Wood, and A. Sarkar, “Perceptions of chatbots in therapy,” in Extended abstracts of the 2019 CHI conference on human factors in computing systems, 2019, pp. 1–6.
[63]
J. Grodniewicz and M. Hohol, “Therapeutic chatbots as cognitive-affective artifacts,” Topoi, vol. 43, no. 3, pp. 795–807, 2024.
[64]
L. Edwards, “The EU AI act: A summary of its significance and scope,” Artificial Intelligence (the EU AI Act), vol. 1, p. 25, 2021.
[65]
J. Ruohonen and K. Hjerppe, “The GDPR enforcement fines at glance,” Information Systems, vol. 106, p. 101876, 2022.
[66]
J. R. Ancis, “The cyberpsychology influence on modern computing,” Communications of the ACM, vol. 68, no. 11, pp. 72–79, 2025.
[67]
eSafety Commissioner, Accessed: 2026-04-01“eSafety report shows AI companions are putting children at risk.” https://www.esafety.gov.au/newsroom/media-releases/esafety-report-shows-ai-companions-are-putting-children-at-risk, Mar. 24, 2026.
[68]
eSafety Commissioner, Accessed: 2026-04-01“New safety advisory warns unrestricted chatbots threaten child development.” https://www.esafety.gov.au/newsroom/media-releases/new-safety-advisory-warns-unrestricted-chatbots-threaten-child-development, Feb. 18, 2025.
[69]
A. Hurst et al., “Gpt-4o system card,” arXiv preprint arXiv:2410.21276, 2024.
[70]
Anthropic, “System card:claude opus 4 & claude sonnet 4.” https://www.anthropic.com/claude-4-system-card, 2025.
[71]
N. A. Smuha, “Regulation 2024/1689 of the eur. Parl. & council of june 13, 2024 (eu artificial intelligence act),” International Legal Materials, vol. 64, no. 5, pp. 1234–1381, 2025.
[72]
J. Yu, X. Lin, Z. Yu, and X. Xing, \(\{\)LLM-fuzzer\(\}\): Scaling assessment of large language model jailbreaks,” in 33rd USENIX security symposium (USENIX security 24), 2024, pp. 4657–4674.
[73]
X. Liu, N. Xu, M. Chen, and C. Xiao, “AutoDAN: Generating stealthy jailbreak prompts on aligned large language models,” in The twelfth international conference on learning representations, 2024, [Online]. Available: https://openreview.net/forum?id=7Jwpw4qKkb.
[74]
Character.AI, Accessed: 2026-04-15“Safety center.” https://support.character.ai/hc/en-us/articles/21704914723995-Safety-Center, 2025.
[75]
S. Pieritz, M. Khwaja, A. A. Faisal, and A. Matic, “Personalised recommendations in mental health apps: The impact of autonomy and data sharing,” in Proceedings of the 2021 CHI conference on human factors in computing systems, 2021, pp. 1–12.
[76]
K. P. Kruzan, J. Meyerhoff, T. Nguyen, D. C. Mohr, M. Reddy, and R. Kornfield, ‘I wanted to see how bad it was’: Online self-screening as a critical transition point among young adults with common mental health conditions,” in Proceedings of the 2022 CHI conference on human factors in computing systems, 2022, doi: 10.1145/3491102.3501976.
[77]
H. Sun et al., “On the safety of conversational models: Taxonomy, dataset, and benchmark,” in Findings of the association for computational linguistics: ACL 2022, 2022, pp. 3906–3923.
[78]
T. Dinkar, “Safety and robustness in conversational AI,” in Proceedings of the 19th annual meeting of the young reseachers’ roundtable on spoken dialogue systems, Sep. 2023, pp. 5–8, [Online]. Available: https://aclanthology.org/2023.yrrsds-1.2/.
[79]
L. I. Laestadius and C. Campos-Castillo, “Reminders that chatbots are not human can be risky,” Trends in Cognitive Sciences, 2026.
[80]
M. Shin, H. Chin, H. Song, Y. Choi, J. Choi, and M. Cha, “Context-aware offensive language detection in human-chatbot conversations,” in 2024 IEEE international conference on big data and smart computing (BigComp), 2024, pp. 270–277.
[81]
G. Sun, X. Zhan, S. Feng, P. Woodland, and J. Such, CASE-bench: Context-aware SafEty benchmark for large language models,” in Proceedings of the 42nd international conference on machine learning, 2025, vol. 267, pp. 57938–57960, [Online]. Available: https://proceedings.mlr.press/v267/sun25ab.html.
[82]
B. Inkster, S. Sarda, and V. Subramanian, “An empathy-driven, conversational artificial intelligence agent (wysa) for digital mental well-being: Real-world data evaluation mixed-methods study,” JMIR mHealth and uHealth, vol. 6, no. 11, p. e12106, 2018.
[83]
D. D. Coppersmith, K. H. Bentley, E. M. Kleiman, A. C. Jaroszewski, M. Daniel, and M. K. Nock, “Automated real-time tool for promoting crisis resource use for suicide risk (ResourceBot): Development and usability study,” JMIR Mental Health, vol. 11, p. e58409, 2024.