September 22, 2025
Qualitative research offers deep insights into human experiences, but its processes, such as coding and thematic analysis, are time-intensive and laborious. Recent advancements in qualitative data analysis (QDA) tools have introduced AI capabilities, allowing researchers to handle large datasets and automate labor-intensive tasks. However, qualitative researchers have expressed concerns about AI’s lack of contextual understanding and its potential to overshadow the collaborative and interpretive nature of their work. This study investigates researchers’ preferences among three degrees of delegation of AI in QDA (human-only, human-initiated, and AI-initiated coding) and explores factors influencing these preferences. Through interviews with 16 qualitative researchers, we identified efficiency, ownership, and trust as essential factors in determining the desired degree of delegation. Our findings highlight researchers’ openness to AI as a supportive tool while emphasizing the importance of human oversight and transparency in automation. Based on the results, we discuss three factors of trust in AI for QDA and potential ways to strengthen collaborative efforts in QDA and decrease bias during analysis.
<ccs2012> <concept> <concept_id>10003120.10003130.10003134</concept_id> <concept_desc>Human-centered computing Collaborative and social computing design and evaluation methods</concept_desc> <concept_significance>500</concept_significance> </concept> <concept> <concept_id>10003120.10003130.10003131</concept_id> <concept_desc>Human-centered computing Collaborative and social computing theory, concepts and paradigms</concept_desc> <concept_significance>500</concept_significance> </concept> </ccs2012>
Qualitative research is a foundational methodology used across diverse disciplines such as human-computer interaction, sociology, psychology, anthropology, and healthcare [1], [2]. It provides deep insights into complex human behaviors, experiences, and perspectives, often addressing aspects that quantitative methods cannot capture [3]. One of the most widely used methods within qualitative research is thematic analysis, a process for identifying, analyzing, and reporting patterns (themes) within data [4]. In qualitative research, coding involves assigning specific, granular labels to raw data segments to create the foundational building blocks of analysis [4], [5]. These codes are subsequently clustered and synthesized into themes, which are broader patterns and interpretive insights that capture significant meaning within the dataset [4], [5]. While these qualitative data analysis (QDA) methods generate rich and nuanced findings, this process is often socially intensive, frequently involving teams of researchers engaging in discussion, negotiation, and interpretation to reach shared understanding of the data [6]–[8]. As qualitative datasets grow larger and more complex, the collaborative labor involved in coding and interpreting becomes increasingly time-consuming and labor-intensive, and is susceptible to implicit bias due to the subjective nature of the process [9]. Given the extensive nature of this work, qualitative researchers often rely on QDA tools to facilitate organization and streamline workflow [10]–[12]. In addition, researchers have looked into ways to incorporate artificial intelligence (AI) to support QDA [7], [13]–[15]. Recently, QDA tools including NVivo and MAXQDA started automating various tasks, such as coding and theme generation, reducing the manual effort involved in analysis. AI-powered QDA tools offer the potential to streamline workflows, manage large datasets, and improve efficiency [16].
Despite these advancements, the integration of AI into the collaborative practices of qualitative analysis was not immediately embraced by all qualitative researchers [17]–[21]. An earlier study showed that researchers enjoy the discoveries they make through manual coding due to the interpretive nature of qualitative research and resist the notion of AI fully automating this work [18]. More recent studies show a growing openness of qualitative researchers to collaborating with AI, and suggest ways in which AI can or should be used [15]. Yet, while researchers have proposed AI as a novel form of collaborator [7], little is known about how people perceive this AI “collaborator" compared to human collaborators, and how AI might affect or augment the collaborative nature of QDA.
Given the conflicting views on the integration of AI in QDA, this study explores how qualitative researchers perceive the role of AI in their collaborative analysis workflows, using Lubars and Tan’s framework on human perception of task delegation to AI [22]. Their framework divides the degree of delegation into four categories: no AI assistance, the human leads and the AI assists, the AI leads and the human assists, and full AI automation. We adjusted the degrees to describe different types of AI involvement in the QDA process: human-only, human-initiated, and AI-initiated coding.
Human-only coding involves human coders manually assigning codes based on their interpretation of the data, either by hand or using QDA software without any AI involvement.
Human-initiated coding refers to a process where codes are created and applied by a human coder, grouped into themes, and the resulting codes and themes are later reviewed using AI tools.
AI-initiated coding, in contrast, uses machine learning algorithms to automatically identify codes and assign these codes to data segments, or cluster codes into broader themes, with human coders reviewing the results afterward.
We omitted the full automation category given the strong resistance shown in prior work [18]. Based on this framework, we examined the desired role of AI in the QDA process through interviews with 16 qualitative researchers. We investigated the topic through the lens of how their preferences align with existing human collaboration in QDA and how AI might shift or support these collaborative practices. We further explored the topic of bias in QDA and participants’ use of multimodal data (e.g., audio and video), as it represents an underexplored strategy for mitigating bias through capturing nuances like tone, pitch, and non-verbal cues [23]. Through the study, we aimed to answer the following three research questions:
RQ 1. How do qualitative researchers’ preferences for different degrees of AI involvement in QDA (human-only, human-initiated, and AI-initiated coding) vary across different contexts, and what factors shape these preferences?
RQ 2. What are qualitative researchers’ perspectives on bias in qualitative data analysis and the use of AI for mitigating bias?
RQ 3. How do qualitative researchers perceive the value of multimodal data (e.g., audio, video), and how do they incorporate it into their analysis practices?
Our study showed that human-only coding is still the most preferred method, as participants were reluctant to give interpretive authority to AI. Yet, they were open to workflows where AI acts as an assistant, not collaborator or supervisor, either by generating code suggestions that humans can accept or reject, or by automating repetitive tasks. Many expressed discomfort in human-initiated coding, where AI checks human work, as it undermined their sense of ownership of the codes and professional identity as a qualitative researcher. The interviews surfaced a mistrust in the involvement of AI in the QDA workflow due to a lack of contextual understanding and concerns on data confidentiality. Despite recognizing the potential of audio and video to enrich analysis, participants relied heavily on text due to time constraints and inadequate tools, highlighting an opportunity for future QDA systems to integrate multimodal analysis. Through the study, we make the following three primary contributions:
1. Qualitative comparison of AI delegation paradigms in QDA: We explored researchers’ preferences across three degrees of delegation for AI in the QDA workflow (human-only, human‑initiated, and AI‑initiated coding) and show how efficiency, ownership, and trust influence people’s preferred degree of delegation.
2. Design implications that position AI as an assistant rather than a collaborator: We propose concrete CSCW designs that use AI as a method of enhancing human collaboration in QDA by distinguishing authorship and establishing trust.
3. Highlighting opportunities for mitigating bias, including multimodal integration: We address the underutilization of multimodal data in current QDA workflows and propose AI-driven approaches for incorporating non-verbal cues to support richer and more bias-aware analysis.
Historically, manual qualitative data analysis (QDA) presented significant logistical hurdles. The transcription process alone was exceptionally time-consuming; for instance, a mere 45-minute audio-recorded interview could demand approximately eight hours of a trained researcher’s time, resulting in 20 to 30 pages of written text for analysis [24]. This extensive manual effort in data preparation and handling was a major bottleneck in the research workflow.
When the QDA is performed collaboratively as a team, especially one that is geographically distributed or working on complex projects, a distinct set of challenges emerges [18]–[21]. These difficulties can be broadly categorized into three main areas. First, technical infrastructure issues pose significant barriers. These include incompatibilities between different software tools, problems with secure data sharing, challenges related to legacy data, differing file standards, and complexities reminiscent of version control problems [25]–[27]. Second, methodological coordination presents considerable challenges. Teams often struggle with overcoming the absence of standardized analytical procedures, aligning coding schemes, and ensuring and managing inter-coder reliability [18], [21]. Third, team communication and workflow management are often problematic, especially in remote settings. Difficulties in decision-making, effective knowledge sharing, and coordinating efforts across different time zones are key challenges in such contexts [19], [25], [28]–[30].
The advent of Computer-Assisted Qualitative Data Analysis Software (CAQDAS) marked a significant shift in the QDA landscape, addressing many of the initial manual challenges [31], [32]. Leading CAQDAS tools like NVivo [33] and ATLAS.ti [34] began to streamline the research process by offering features such as automated transcription and more efficient management of large datasets [31]. These tools considerably reduced the time required for thematic analysis by allowing researchers to sort data by codes and summarize coded segments across multiple interviews. While some CAQDAS packages offered assistance in theory development and advanced search functionalities, they fell short in areas such as recognizing synonyms, which could hinder the extraction of deeper insights [11], [35]. Other tools focused on strengths like automated concept mapping, robust data storage, and flexible management capabilities [10], [36], illustrating the ongoing technological evolution within the field.
Despite the advancements brought by CAQDAS, limitations persisted, particularly in tasks requiring high-volume pattern recognition and time-intensive data sorting, which are uniquely suited for Artificial Intelligence (AI). Marathe and Toyama [37] identified that many coding practices across various disciplines could potentially be automated. The application of computational methods in qualitative research promised enhanced efficiency and a potential reduction in bias and errors [38]. Consequently, the research community began to explore automation processes [39], [40], and some CAQDAS tools incorporated features like “Automated Insights” in NVivo, which leverage topic modeling [41]. Further research into semi-automated systems such as Cody, which integrates user-defined rules with supervised machine learning, demonstrated improvements in coding efficiency and consistency [13]. Studies by Overney et al. (2024) [14] and van der Riet et al. (2021) [13] have further highlighted AI’s capacity to enhance coding quality and inter-rater reliability through AI-assisted coding.
While these earlier works focused on augmenting individual efficiency, recent HCI and CSCW research has shifted toward exploring how AI could be integrated to enhance the collaborative aspect across the full workflow of qualitative inquiry from dynamic data collection, collaborative coding, to interpretive partnership. First, in the data collection phase, researchers are reshaping how information is acquired by collaborating with AI. Cuevas et al. [42] demonstrated that LLM-based interviewing chatbots can act as collaborative partners in gathering data, fostering new forms of dynamic engagement between researchers and participants that was previously difficult to scale. Lei et al. [43] also explored AI-supported data collection, but through dynamic surveys where LLM is used to cluster responses in realtime and prompt survey respondents to reflect on the answer trends. Their work aimed at designing survey platforms that support both quantitative structure and collaborative interactions. Second, in the coding and consensus phase, Wang et al. [44] explored the feasibility of “LATA” (LLM-Aided Thematic Analysis), highlighting that while AI can mimic the coding process, it introduces challenges in reliability that require human oversight. Consistent with these concerns, Gao et al. [45] investigated the dynamics of AI-mediated collaborative coding through CoAIcoder, and found that while AI suggestions improved efficiency and agreement, they also introduced a risk of “shallow agreement,” where human researchers converged on the AI’s suggestions rather than engaging in deep, independent interpretation. Third, regarding high-level interpretation, literature suggests a shift from automation to “augmentation” of serendipity. Jiang et al. [18] framed the AI not as a tool for automation, but as a system that honors uncertainty and supports serendipity, arguing that the goal of collaboration is to expose researchers to unexpected patterns while maintaining human agency. Feuston and Brubaker [15] expanded on this by using storyboard scenarios to map the “delegation” of tasks, identifying how researchers negotiate shifts in scale and abstraction when collaborating with AI. Finally, moving into theme construction, Kang et al. [46] explored how interactive systems (ThemeViz) can mediate the “sense-making” process, allowing researchers to actively negotiate and refine AI-generated themes rather than simply accepting them.
Despite these collaborative advances, significant challenges remain regarding the depth of interpretation. Qualitative research values subtle differences in meaning and the absence of utterances [47], nuance that current models often miss [48], [49]. This tension between efficiency and interpretive depth remains a central hurdle in Human-AI QDA. Our research extends this groundwork by focusing specifically on the mechanisms of delegation. Building on Lubars and Tan’s framework [22], we investigate whether researchers view the “AI partner” similarly to a human collaborator (requiring negotiation) or a research assistant (requiring supervision), and how these views shape the desired “collaborative contract” in QDA workflows.
Qualitative analysis is inherently interpretive, relying on the researcher’s subjectivity as a lens to make sense of data. However, a distinction must be drawn between subjectivity, which is a necessary feature of qualitative inquiry and bias, which threatens validity. Drapeau explained that unexamined subjectivity may be introduced in qualitative research, particularly where there is a lack of consensus on definitions, leading researchers to code based on internalized prejudices or conflicts [50]. In addition, researchers are often influenced by how the results may be used later on (e.g., in court, by the media) and could lead to an unconscious decision to derive conclusions based on the moral and legal implications rather than the data present [50]. In HCI and CSCW, reflexivity is the standard method for mitigating these risks, requiring researchers to critically examine their own positionality [51], [52]. Without this, unexamined subjectivity can harden into cognitive biases, such as confirmation bias, where researchers select data that fits pre-existing beliefs [53]. This risk is particularly acute when researchers rely solely on text transcripts; the removal or underutilization of multimodal data (e.g., audio tone, hesitation) can strip away the nuance needed to challenge a researcher’s initial textual interpretation, thereby exacerbating confirmation bias.
While human researchers struggle with cognitive bias, AI introduces a different challenge. Given that AI systems are trained on human-provided datasets, societal biases can be encoded into the model, resulting in algorithmic bias. Panch et al. [54] discuss this form of bias and stress the importance of designing AI systems to actively mitigate it. Our study builds on these discussions by exploring qualitative researchers’ views on how these two distinct forms of bias — human cognitive bias and AI algorithmic bias interact and can be mediated. Prior work by Ma et al. on trust calibration (i.e., comparing the user’s trust in AI with AI’s actual capability) showed that to establish an adequate level of trust, we not only have to assess AI’s correctness likelihood but also the user’s correctness likelihood [55]. Similarly, rather than viewing AI as a neutral adjudicator, we investigate how QDA tools might allow researchers to leverage AI to identify their own blind spots, while simultaneously using their human contextual awareness to audit algorithmic limitations.
We define "codes" and "themes" in alignment with established qualitative research practices [4], [5]. Codes are the specific, granular labels (word/short phrase) assigned directly to segments of qualitative data (e.g., text and audio transcripts) and serve as raw building blocks. Themes represent broader, more abstract patterns or insights that emerge from clustering, categorization, or analytic reflection of individual codes. Our study investigates AI’s role in QDA, with the AI assistance primarily directed at supporting the coding phase by generating suggestions for these granular codes. The subsequent process of developing broader themes remains a researcher-led interpretive task.
We recruited individuals across various departments involving qualitative research (e.g., sociology, history, anthropology, and information science) through campus flyers, email outreach, and professional networking platforms such as LinkedIn. Inclusion
criteria for eligible participants included a minimum of one year of experience in qualitative research and prior experience using a qualitative data analysis tool.
Interviews were conducted with 16 participants, aged 23 to 40 (mean = 29.1), with qualitative analysis experience ranging from 2 to 7 years (mean = 3.9). Participants reported their highest degree completed or currently being pursued: 9 were pursuing or
had completed a PhD, and 7 a Master’s degree. Eleven participants were in academia (8 PhD students and 3 MS students), and five participants were user experience researchers in industry. The research expertise of participants ranged from accessibility,
human-machine teaming, behavioral decision-making, paralinguistic elements in captions for the Deaf and Hard-of-Hearing, early childhood education, to HIV counseling. More information is provided in Table 1.
To ground the discussion and give participants an idea of what AI integration in QDA might look like, we developed ChromaScribe as shown in Figure 1, an AI-based QDA webtool prototype. The tool was introduced as an exploratory artifact rather than a fully realized QDA system. It was not designed to adhere to a single AI-delegation type, and it was possible to use the tool for both the AI-initiated and human-initiated workflow. This choice follows established prior work showing that early-stage prototypes are effective for exploration and reflection rather than evaluation, helping participants articulate needs, expectations, and values around emerging technologies [15], [45], [56]. Thus, the prototype served as a concrete prompt to support discussion around AI-supported qualitative analysis, including perceived benefits, concerns, and feasibility.
However, its inclusion of AI-generated themes could have functioned as a framing device. By exposing participants to these themes, the system may have anchored participants’ mental models of the role of AI in QDA. The study findings should be interpreted within this context. The webtool uses a locally hosted interface to ensure the data is not accessed by any third-party cloud servers. Access to the prototype was restricted through a password-protected login, and credentials were shared only with study participants for the duration of the study. To further ensure confidentiality, the access password was changed after each study session. Upon logging in, participants encountered a pre-loaded dataset, which consisted of anonymized interview transcripts and corresponding audio recordings from a study on HIV molecular surveillance involving healthcare workers and individuals from affected communities. The tool automatically identifies and color-codes themes based on transcripts and audio files, generating a set of preliminary themes using natural language processing, which are then presented to users for review. We selected established, non-generative natural language processing (NLP) models over opaque Large Language Models (LLMs). Specifically, we utilized BERTopic [57] to derive initial latent topics and Word2Vec [58], [59] to identify related keywords. This deterministic pipeline allowed us to explain the system’s logic clearly to participants, mitigating the “black box” effect often cited as a barrier to AI adoption [60]. Users can either select the generated themes or manually add their own, as illustrated in Figure 2. A visualization-based interface displays theme distributions over time, indicating where specific themes appear most frequently in the transcript, as shown in Figure 3.
| Participants | Age | Gender | Years of Experience | Highest Degree | Research Area | Tools used for QDA |
|---|---|---|---|---|---|---|
| P1 | 26 | F | 4 | PhD student | Human-AI collaboration, computing education accessibility | Google Docs, Google Sheets, Miro |
| P2 | 27 | F | 2.5 | PhD student | Design, culture, and accessibility | Google Docs, Google Sheets, Miro |
| P3 | 24 | F | 3 | Masters | Accessibility | Google Docs, Google Sheets, Zoom, FigJam |
| P4 | 33 | F | 7 | PhD | User studies on AI technology | Zoom, Webex |
| P5 | 27 | F | 5 | Masters student | Human-Computer Interaction (HCI) | Word, Miro, Microsoft Teams, FigJam, ChatGPT, iOS voice recording app, Figma |
| P6 | 26 | F | 4 | Masters student | Language learning, VR, ASR experiences | Google Sheets, Zoom, Atlas.ti, Otter AI, Whiteboarding tools, AirTable, Descript |
| P7 | 31 | F | 6 | PhD student | Chatbots, job support technologies | Excel, Miro, Microsoft Teams, Otter AI |
| P8 | 40 | M | 2 | PhD student | Captioning for d/Deaf and HoH users | Google Docs, Google Sheets, Miro |
| P9 | 32 | F | 6 | PhD student | Human-Computer Interaction | Google Docs, Zoom, Whisper, NVivo, Atlas.ti, Open source QDA on Github |
| P10 | 26 | M | 3 | Masters student | AI in clinical decision-making, model analysis | Google Docs, Google Sheets, Otter AI, Descript, Deedose |
| P11 | 28 | F | 2 | Masters | Web and cognitive accessibility, digital tech use | Google Docs, Excel, Microsoft Teams, Deedose |
| P12 | 33 | M | 6 | PhD student | AI and media forensics, public sentiment | Google Docs, ChatGPT, Gemini advanced, Claude |
| P13 | 35 | F | 3 | Masters | Early childhood and home education | Excel, Microsoft Teams, Deedose, Covidence |
| P14 | 26 | F | 3 | PhD student | Deepfake detection tool design | Google Docs, RefAI |
| P15 | 23 | F | 3 | PhD student | Accessibility for older adults and neurodivergent users | Google Docs, Excel, Miro, Zoom, Taguette |
| P16 | 29 | F | 3 | Masters | HIV services and relationship dynamics | NVivo |
We conducted a formative study involving a pre-study demographic survey, an exploration of ChromaScribe, a semi-structured interview, and a post-study survey. The study was reviewed and approved by the university’s Institutional Review Board (IRB). The pre-study survey contained questions on demographic details, QDA experience, coding approaches, and the use of multimodal data. Participants were screened based on their research experience and inclusion criteria.
Each interview began with a consent process and included questions on participants’ current QDA workflow, QDA tool use and preferences, and current challenges related to the QDA process and tools. In the next phase, participants were introduced to ChromaScribe. After watching a demo video, participants were given time to explore the webtool independently and then asked to complete four structured tasks including (a) locating a participant who mentioned a selected theme at least three times and (b) playing the audio of a segment where the interviewer mentioned a selected theme. These tasks were designed to prompt participants to explore how the tool supports theme discovery, navigation, transcript-audio alignment, search and to ensure that all participants engaged with the core functionalities of the prototype. Participants were informed that the prototype is preliminary and only meant to facilitate the discussion.
To address RQ1, participants were introduced to three distinct types (i.e. degrees of delegation) of AI involvement in qualitative coding: human-only coding, AI-initiated coding, and human-initiated coding, after the completion of four tasks on ChromaScribe. The three coding types were presented to participants in a written format to elicit preferences and perceptions during the interview. Each participant was given a brief textual description of the three types and allotted two minutes to read them carefully. They were then asked to rank the coding types from most to least preferred. We further prompted them to share the reason for their preference. To minimize bias toward any particular AI delegation paradigm, the prototype was introduced prior to presenting participants with the three AI involvement modes and the participants were encouraged to explore different workflows as they desired. To address RQ2, we then asked about their collaborative practices in QDA, their thoughts on subjectivity and bias in the process, and how QDA tools could address bias. Lastly, to address RQ3, participants were asked about their current reliance on transcripts, audio, and video during analysis and we asked about their use and thoughts on multimodal data in QDA.
Two pilot interviews were conducted in person to refine interview questions and timing. Following this, an in-person focus group with two participants was held in a research lab at a university. Due to scheduling challenges, the remaining 14 interviews were conducted individually via Zoom, lasting approximately 90 minutes each. As an incentive, participants received a $30 gift card upon completing the post-study survey.
All interview sessions were recorded, with individual sessions transcribed through Zoom’s transcription feature and the focus group session transcribed through Otter.AI. Transcripts were then manually cleaned using Google Docs and coded with Atlas.ti. Data were analyzed using axial coding, a process of relating codes (categories and concepts) to each other, often involving a combination of inductive and deductive thinking to develop more abstract conceptual categories from initial codes [61]. Two researchers independently conducted word-level and sentence-level coding to identify initial codes. To develop the initial codebook, both researchers independently coded the same three transcripts, then met to compare their coding decisions, resolve discrepancies, and refine a shared set of codes. The resulting codebook, maintained in a spreadsheet format, detailed codes, thematic descriptions, and relevant examples. Once finalized, one researcher applied the codebook to the remaining transcripts. When new codes emerged during this phase, they were documented and discussed with the second researcher to determine whether they represented new concepts or could be integrated into existing codes, with the codebook updated accordingly. This iterative process ensured that coding remained consistent and reflexive across the dataset. Key themes included participants’ sense of ownership in the QDA process, anticipated benefits and concerns of incorporating AI into QDA, and how it compares with collaborative efforts with other researchers.
RQ1. How do qualitative researchers’ preferences for different degrees of AI involvement in QDA (human-only, human-initiated, and AI-initiated coding) vary across different contexts, and what factors shape these preferences?
| Coding Method | Most Preferred | Least Preferred |
|---|---|---|
| Human-only Coding |
(n = 7) - Comprehensive understanding of the data - Complete control over the process and higher reliability. |
(n = 5) - Requires significant time and effort from the researcher. |
| AI-Initiated Coding | (n = 5) - Reduced time for repetitive coding tasks and minimized mental effort. - Frees time for higher-level analysis and decision-making. |
(n = 4) - Less control over the coding process leading to less reliable results. - AI generated themes may lack depth and context |
| Human-Initiated Coding | (n = 4) - Leverages researchers’ contextual understanding of study domain and setup. - AI supports verifying codes without compromising control of the process. |
(n = 7) - Discomfort and distrust in AI verifying human QDA work. - The additional layer of coding and verification increases researcher’s workload |
A key factor influencing participants’ preference for human-AI collaboration styles was the sense of ownership each style gave. Many participants preferred human-only (n=7) and AI-initiated (n=5) approaches over the human-initiated approach, where AI verifies or revises human-generated codes. Several participants felt uncomfortable with the human-initiated approach, as they perceived a loss of authority or identity as a qualitative researcher. P5 voiced this discomfort, “Because I have to do everything initially... You [AI] are going to make me do it. You are going to verify... I mean, are you doubting me? It is that superiority complex that comes in. Are you doubting my work, like you were born yesterday, right?” Rather than viewing the human-initiated approach as productive collaboration, they felt that it undermined their expertise and professional autonomy. In their preferred (AI-initiated) approach, the human researcher was positioned to accept or reject AI-generated suggestions, not the other way around.
Interestingly, people who preferred the human-initiated approach also mentioned ownership as a primary motivation. As P2 explained:
I would rather myself go through it first and then, like, have AI check my work – or I have not checked that, like I did not miss anything. Because, in a sense, that makes me feel like I own more of the results. That’s what makes me a researcher. As a researcher, if I’m going through a PhD program and doing years of training, to be trained in this, I feel like I should have something to offer if it is an integral part of my job. Like, I’m happy to use AI as an assistant, but I do not think that it should [lead it].
In both cases, ownership emerged as a guiding principle, whether it meant taking the first pass or making the final call. This emphasis on ownership contrasted with how participants viewed collaboration with human peers. Participants acknowledged the centrality of (human) collaboration in QDA, emphasizing the need for QDA tools that support collaborative analysis. As one participant (P10) noted, "I think collaboration is super important. like, we always need to have multiple coders working on it. But it’s also important to know when coders disagree, and to see, like the history of that how the codes evolve over time." Shared authorship in human collaboration was not only accepted but often valued; disagreement and negotiation were deemed natural in the QDA process where codes and themes are co-constructed through discussions and iterative refinement for a rigorous analysis. However, participants were more ambivalent when considering collaboration with AI in this process. While they were open to incorporating AI into their workflows, sharing interpretive authority with AI raised concerns about ownership and even impact on the professional identity of the researcher. Across participants’ preferences for human-only, human-initiated, and AI-initiated coding, we observed a consistent pattern in how researchers conceptualized delegation. Participants were broadly willing to delegate low-effort, repetitive, or exploratory tasks such as scanning large datasets, identifying candidate codes, or surfacing recurring patterns, to AI systems but they resisted delegating interpretive authority, particularly for ambiguous, or context-sensitive coding decisions. Despite these reservations, most were open to workflows with asymmetrical collaboration that ensured that AI played a supporting role. While preferences regarding AI involvement varied across participants, researchers working in industry settings consistently expressed openness toward AI-supported qualitative analysis. Notably, all five industry participants, preferred workflows that involved AI to some degree, selecting either human-initiated or AI-initiated coding over fully human-only approaches. For example, P3, an industry researcher who regularly used AI features embedded in tools such as FigJam, favored AI-initiated workflows for tasks such as summarization and pattern detection, emphasizing the substantial time savings while retaining human oversight. Other industry participants preferred human-initiated approaches that positioned AI as a verification or support mechanism. P16, who worked with large, multi-country datasets, described AI as valuable for confirming themes and reducing human error when managing complex, distributed analyses. Together, these perspectives suggest that industry researchers were generally receptive to AI involvement in QDA, with variation centered not on whether AI should be used, but on how responsibility and control should be distributed between human researchers and AI systems. Participants consistently framed AI’s role as that of an assistant, rather than an equal peer or a “co-pilot.” In this assistant role, AI was seen as valuable for surfacing preliminary insights and improving efficiency, especially when managing and analyzing large datasets, provided that the human researcher remained the final decision-maker. P2 commented on AI’s expected efficiency, predicting that it could reduce task time from four hours to two.
While concerns about ownership and identity shaped attitudes toward AI collaboration, participants also expressed reservations based on trust, or rather mistrust, in AI’s interpretive ability. P8 explicitly expressed this mistrust, warning that “the more agency we give to AI, the less reliable the results are.”
This low level of trust was evident in participants’ views on the human-initiated approach, which was the least preferred option for seven out of sixteen participants. The approach was seen as inefficient as they would have to generate the initial codes, have the AI check the work, and re-verify the results themselves at the end. This perceived need to personally finalize and validate coding outcomes made the added layer of AI verification feel unnecessary or redundant. P6 noted that she would not feel comfortable relying on AI to validate her work, explaining, “I think I prefer the process of getting some kind of baseline.. And if I sort out the process, I just want to finish the process. So I do not tend to want to... start the process, and have the AI check it. Cause I wouldn’t trust it as much, I guess.”
One reason for this need to recheck the AI’s work is that unlike human collaborators, AI was perceived as lacking interpretive depth and nuanced, contextual understanding that researchers bring to complex qualitative data based on their expertise and disciplinary training. As a result, participants raised concerns about the rigor of the analytical process, noting that AI would produce surface-level themes that failed to capture deeper meanings. Several participants noted that AI’s inability to grasp the broader context of a study made it difficult to trust its coding decisions. Interestingly, the second reason for disliking the human-initiated approach was not about AI’s performance, but its perceived role in the process; participants simply did not want AI to take the role of a supervisor. As participant (P7) explained, “So when ending in a human making the decision, checking the AI’s work is really important. Which is why human initiating coding - I do not think I would ever do that. Because I want to check the AI’s work. I do not want the AI to check my work.”
Data confidentiality also emerged as a crucial concern, profoundly influencing how participants selected QDA tools and whether they even used them, especially those involving AI. For instance, P12 stated that his primary reservation in using QDA tools was privacy, and his willingness to use AI tools hinged on the availability of privacy-preserving features: “My only issue would be privacy… All of the tools [Claude, Gemini, ChatGPT] provide privacy mode. So you don’t give them the content you are analyzing.” This sentiment was echoed by P14, who expressed a clear preference for local data processing to prevent sensitive information from being transmitted to third parties: “If this tool is just local, it’s on my desk. It’s not… giving out to a 3rd party, then give me results back, that’s fine. I want something like very like localized.”
This fundamental concern led participants to articulate specific preferences and strategies for interacting with AI tools to safeguard their data. Beyond relying on QDA tool features, some participants described proactive measures to de-identify data before AI interaction. P5, while acknowledging the utility of tools like ChatGPT for thematic analysis, detailed a cautious approach: “Not saying that I literally put everything in ChatGPT as it is, like the whole transcribing file. No. Had to like still mask a lot of details, remove a lot of details.” Similarly, P11 highlighted the risks associated with audio data containing personally identifiable information and suggested modifications:
For example, the audio that you’re saving as a recording, it becomes a PI issue – like personal identifiable information. Because if you’re storing the audio in its exact like, user tone [...] it’s possible for someone to identify who might be speaking based on the context and the voice. So if I were to store audio, I would like disrupt the … or like voice, I’ll make it into a more like robot... Or if I’m storing the voice, then I would make sure the tool has like all the encryptions and privacy protection stuff.
Furthermore, overarching institutional guidelines played a significant role. P3, who works with sensitive data, underscored this by stating the limitations imposed by ethical review boards: “According to IRB, I cannot release my data of what participants said during the interview because [...] it might contain some kind of information which is private.” P16 also emphasized that “data privacy guidelines” was a critical factor in selecting QDA tools, especially with AI involvement, stating, “We still do have a lot of data privacy guidelines on how our data is used and who has access to it, or what... on the Internet has access to it.” The inability to verify what data is stored by AI tools, how it is processed, or whether it is used for training future models heightened their discomfort in using AI-based QDA tools.
These findings suggest that trust in AI for qualitative analysis is not merely about the accuracy of its outputs. It is tied to its perceived capacity to demonstrate contextual understanding and to operate in a way that respects data privacy and researcher control, standards many participants felt current AI systems did not yet meet. Participants did not converge on a single preferred degree of AI involvement. Instead, preferences reflected context-specific trade-offs shaped by researchers’ prior experience with qualitative methods, familiarity with AI tools, collaborative practices, and data sensitivity, illustrating how preferences for human-only, human-initiated, or AI-initiated coding are contingent on the analytical goals and constraints of a given research context. (see Section 4.1.)
RQ 2. What are qualitative researchers’ perspectives on bias in qualitative data analysis and the use of AI for mitigating bias?
Although human-only coding was the most preferred method, some participants recognized the potential of AI-initiated coding as an additional collaborative layer for managing bias in qualitative research. All participants acknowledged the presence of cognitive bias in human-led qualitative analysis and that collaboration is key to mitigating bias.
However, when asked whether QDA tools, with or without AI, could help reduce bias, participants expressed mixed views. Many participants were skeptical of AI’s current ability to address bias without human oversight. More specifically, they were divided on whether AI could act as a meaningful collaborator in this context. For most, addressing bias was seen as a deeply contextual and reflective process, which is fundamentally human and can only be provided by a human collaborator: "I think it’s hard to like tease out the bias part. I think sometimes we just need a 3rd person to tell me this is the actual interpretation" (P14). In their view, AI lacked the contextual sensitivity and critical judgment needed to challenge or validate human assumptions in a meaningful way.
Despite these reservations, participants remained open to the idea of using AI as a supplementary way of checking for bias, either for cases where human collaborators were unavailable or for enhancing current human collaboration. For example, P1 shared, “I use [ChatGPT] as like a check for my own analysis to see, ‘oh, they found this, but I did not find this’ or ‘I identified more themes and they did more.’ If I do not have a collaborator, I use ChatGPT as sort of like that extra check.” Some participants suggested features that would enable QDA tools to better support bias reduction. For example, P3 noted that QDA tools could check for bias by treating transcripts equally, avoiding cherry picking data that could lead to confirmation bias. She further mentioned that QDA tools could help organize and structure data for a more impartial overview of the data, though noting that this alone is not sufficient to reduce bias. Others envisioned how initial theme suggestions or confirmation prompts could assist researchers in minimizing bias by aiding self-reflection. Notably, P5 mentioned that bias is not something that can be entirely engineered out of qualitative research: “the whole point of user interviews has to have that human touch. There is, you cannot deny there is going to be some amount of bias.” Participants further highlighted how current human collaboration addresses the issue bias by making interpretations visible and open to challenge. Participants described aligning interpretations across teams, such as checking whether “the themes that we are finding here correlate with the things that they are finding there” (P8) and resolving discrepancies through discussion, where “[we] combine our codes… [and we] may have discrepancy during the open coding process. There is a discussion in the process of merging the codes together, we’ll be able to reach higher themes out, which is probably a more refined and in detail” (P14). Rather than eliminating subjectivity, collaboration surfaced them, requiring researchers to critically examine their assumptions. As P12 noted, researchers must “separate yourself… and not defend your own work.” Others emphasized that reflexivity also involves accountability for bias in data collection and analysis, such as recognizing and correcting errors: “this was a leading question… can I trim this?” (P5). At a broader level, collaboration was framed as essential for interpretive rigor, where meaning is refined through ongoing discussion of what constitutes “the most rich information… [while] being true to what is being said” (P16). These iterative conversations and shared workflows support the validation and refinement of interpretations.
In all, while participants saw potential in AI to complement their efforts, bias reduction in qualitative analysis was seen as a human-driven process, benefiting most from collaboration and perspectives of multiple researchers. While QDA tools offer procedural support, they would be most effective when used alongside, rather than in place of, human collaboration.
RQ3. How do qualitative researchers perceive the value of multimodal data (e.g., audio, video), and how do they incorporate it into their analysis practices?
The integration of multimodal data (e.g., text, audio, and video) into QDA offers researchers richer insights [62], [63]. This view was commonly shared by participants as well as they highlighted several advantages of incorporating multimodal data into their workflow for capturing information which is often lost in plain text. To assess the actual use of multimodal data in QDA, we asked participants to give an estimate of how much they relied on text (e.g., transcript, observation notes), audio, and video in percentages, adding up to 100%. Despite the recognized benefits of multimodal data, participant responses reveal a strong reliance on text-based transcripts as the primary medium for qualitative analysis. When considering how strongly they relied on text, estimates ranged from 40% to 100% (mean = 76%). (To note, these estimates are intended to illustrate general trends in participants’ use of different data types rather than to provide precise quantitative measures. Our study is primarily qualitative in nature.) Each participant indicated that text transcripts are essential for their analysis, with some (n=6) relying solely on text for coding and review. For example, P16 stated she uses audio only to gather insights on how participants express themselves but otherwise rely entirely on text transcripts. P15 and P9 both emphasized going through transcripts multiple times to ensure thoroughness in coding, suggesting that transcripts are manageable and useful for repetitive review.
Although most participants collected audio data in their studies, audio data usage in QDA averaged 60% (SD = 11%) among participants. Audio data is mainly used to identify change in tone and pitch, especially for detecting nuances in participants’ expressions, such as laughter or hesitation. P9 explained, “I used audio to identify laughter and tone”, while P8 mentioned using audio for observing mood shifts during the interview. For some, audio was the only medium for data collection (i.e., no video recording, photographs, etc.). P16, engaged in a large-scale qualitative study across five countries, noted, “We strictly do audio recording... and then we transcribe those audio recordings and use the transcripts for our analysis.”
Video data was used even less frequently (mean= 7%, SD = 8.9%), though valued for capturing non-verbal cues and contextual information. P8 highlighted video’s role in “tracking mood changes” and its importance in “interpreting gestures in American Sign Language transcriptions.” Similarly, P3 found video indispensable, stating, “Video has been critical... I also transcribe what a person did and what they said,” thereby integrating non-verbal and verbal cues directly into the textual analysis. Other people used video for gathering behavioral data. For example, P6 described using recordings to “review what people were doing or how people were interacting.” For usability testing, P15 mentioned, “We would do a video recording... we’d track how much time the user took, where they went...,” capturing expressions like surprise visually. P15’s workflow closely mirrors the behavior-centric analysis seen in tools like CoUX [64], where visual evidence is essential for explaining user actions. However, for the majority of our participants conducting standard interviews, the text transcript effectively displaced the need for visual verification, leading to the “flattening” of data described earlier. Video also played a role in verification, with P5 sharing, “It’s just much easier when you want to cross-check on some information.”
Researchers described various methods for aligning and navigating through different modalities such as transcripts, observation notes, audio, and video. The primary method was manually adding non-verbal information from audio or video files into the transcript. For example, P12 described incorporating observations from video or audio into notes during scenario-based studies. Another method was using a separate audio or video player to alternate between audio/video data and transcript. P11 illustrated “When I had to run my data analysis, it had to be both in conjunction with, like the team’s transcript that we generated out of what they were speaking, and then the screen that we recorded. So I had to use these two side by side for understanding the context of the transcript along with the screen recording." She did this in conjunction with the first method as she “used Excel to document quotes I saw in the screen recording and then correlate it with the transcript.” While less common, P12 also mentioned occasionally using AI tools to extract themes from multimodal data, hinting at emerging practices. Participants ideally wanted a better way of transitioning across different multimodal data, such as a synchronized view of video and transcript data. P8, a researcher working on American Sign Language (ASL) projects, primarily relied on text transcripts for data analysis. He shared that “If you are working with ASL transcriptions, I think that the video is definitely useful. It is like an important feature if you want to double check the transcriptions." However, the effort required to review it during analysis was a significant barrier. He confessed, “If it was a streamlined process, where it is very easy to look at a video at a given moment or an audio file, maybe I would do it more. But it just takes so much effort to do it right now that I usually do not bother."
The main reason behind underuse of multimodal data despite the recognized value was the time-consuming nature of engaging with non-textual media. The labor-intensive nature of working with these formats, particularly the difficulty of efficiently locating relevant segments, hindered participants from fully leveraging these sources. This finding illuminates a key difference between our study context and prior work on multimodal systems. While CoUX [64] addresses the technical challenge of synchronizing streams to analyze behavior, our results show that in dialogue-centric analysis, the primary barrier is cognitive and logistical: the text is “good enough” to analyze, so researchers actively avoid the friction of accessing the “richer” audio. Whereas analyzing the video is mandatory in usability testing, qualitative analysis of interview typically relies on the transcript, which acts as a convenient but reductive proxy. P10 pointed out logistical hurdles, stating, “We don’t usually analyze the video as much... it’s harder than we just normally use the audio transcript.” P14 elaborated on a situation where she remembered a participant saying something important but had to manually sift through multiple long transcripts or re-listen to audio to find it. She remarked that “sometimes it’s hard to go back and locate what someone said... I wish I could reverse search across transcripts” and expressed a desire for tools that could “pull out quotes I’m looking for without me rereading everything.” This underscores a need for media-integrated systems capable of retrieving relevant moments in the desired modality (i.e., text or audio) based on semantic or contextual information, not just keywords.
Our study revealed a growing openness among researchers to incorporate AI in the QDA process, diverging from earlier work where participants expressed significant resistance to AI involvement [18], [45], [47]–[49]. Our participants appreciated AI’s potential to streamline tasks in the initial stages of coding of QDA, such as grouping data and identifying common patterns for large qualitative datasets. They envisioned that this approach would allow them to focus on the deeper, interpretive aspects of data analysis that they find most engaging. Yet, significant concerns around ownership, trust, and potential biases made them hesitant about adopting AI in QDA. Crucially, we must interpret these preferences through the methodological lens of our study. The introduction of the prototype prior to evaluating delegation modes could have functioned as a framing device. However, this specific framing provided a valuable lens as exposure to AI-generated themes successfully surfaced participants’ deepest anxieties around interpretive labor and yielded nuanced insights expected from qualitative research. This discussion will explore these key themes, drawing on our findings and existing literature to propose directions for future work aimed at designing and presenting AI as an assistant, especially as a method of enhancing human collaboration and shared sense-making in QDA.
Our participants emphasized the importance of maintaining ownership over the interpretive process, which influenced their coding preferences. Among the three methods presented, human-only coding was the most preferred because it gave researchers the strongest sense of control and interpretive depth. Although AI-generated suggestions were considered useful, participants underscored the importance of staying connected with the data and retaining the final say. This is consistent with Jiang et al.’s findings [18], where participants emphasized that direct engagement with data is vital for sense-making and serendipitous insight. They cautioned that over delegating analysis to AI could disconnect researchers from essential sense-making, weakening emotional connection and the researcher-driven nature of qualitative discovery. Therefore, Jiang et al. advocated for a methodology where AI offers code suggestions, leaving the researcher to accept or decline them. Our results further revealed that direct connection to the data also influenced their identity as a qualitative researcher. For instance, P1 shared that performing data analysis herself affirmed her identity as an HCI researcher and gave her a deeper sense of ownership over the findings.
While the findings mainly focused on individual researcher’s QDA workflow, the introduction of AI could reshape collaborative workflow and ownership as well. First, AI could reshape collaboration by enabling more asynchronous workflows. Because collaborators can independently review and respond to AI-generated inputs, the perceived need for real-time discussion may decrease. This can support scalability by allowing larger teams to divide analytic work into smaller tasks with system support the division aspect as well. At the same time, reduced synchronous interaction risks weakening collaborative sense-making, as there would be fewer shared moments for deep, dynamic discussions. Next, the indiscriminate integration of AI risks fundamentally reconfiguring the interpretive labor of qualitative coding. Our findings show that the current analysis workflow is inherently collaborative and reflexive, driven by moments of “productive friction" where researchers question, justify, and reconcile interpretations to produce rigorous findings. However, the introduction of AI-suggested codes could shift the workflow to center around evaluating system-generated suggestions (“Do we accept or reject the AI’s suggestion?”), replacing co-construction with reaching consensus. Consequently, the meaning of ownership could be altered as the act of approving rather than authoring interpretations, and the QDA results may no longer be deeply rooted in human perspective.
Given the importance of the interpretive process, designers of AI-assisted QDA should be careful of these potential impact of AI. First, to support collaborative sense-making rather than shortcut it, AI-generated codes should be positioned as artifacts for discussion, not final suggestions to be individually approved or rejected. For example, the system could require multiple collaborators to independently evaluate codes that the system is less certain about, and for codes that receive different evaluations, the system could explicitly prompt researchers to articulate why they accept or reject a suggestion. Since the researchers are not critiquing each other’s idea but a third party’s idea, it might also lead to more open discussions. This might especially be helpful to less-experienced researchers, who have reported feeling less confident in the process and being more open to AI.
Future interface designs could support this desire for greater researcher ownership and transparency and facilitate robust collaborative sense-making by incorporating features that clarify the origins and evolution of codes. One option is a split-screen layout where researchers can view AI-suggested codes on one side and their own manual codes on the other, enabling easy comparison and informed decision-making. Another option is to tag the origin of each code, such as “AI-generated," the name of the researcher who coded it, or”researcher-approved” for AI-generated code that has been human-approved. This approach could also be used to show the level of agreement among researchers (e.g., “Below 50% agreement" tag) to highlight codes that still need further discussion and distinguish collaboratively refined codes. Beyond transparency, such features enable AI to actively support delegation by signaling where human involvement is required and can flag analytically difficult segments for discussion, allowing AI to offload routine coding while directing human attention to unresolved interpretations.
This would address P10’s interest in QDA tools that let the user know when coders disagree. He further called for tools that show how codes evolved over time. For this purpose, we can draw inspiration from HistoryFlow [65], which visualizes human dynamics in group editing on Wikipedia to make authorship transparent. Similar visualization could be used in QDA tools, through which researchers can see how conflicting codes were resolved and how their input shaped the final results. This would not only be helpful for an individual researcher, but also for fostering shared accountability and enabling teams to trace the provenance of collective interpretations, a key consideration in CSCW research. Recent work by Ye et al., which focused on how AI could be incorporated in the collaborative workflow without harming the deep reflections, took a similar approach by displaying analysis history as well as providing in-situ reflexive prompts and scaffolding collaborative interpretation [66]. Future work could explore whether this type of visualization affects researchers’ perceived control, collaborative efforts, and willingness to collaborate with AI systems. Furthermore, future research can examine how specific collaboration structures, such as power dynamics between lead researchers and research assistants, or interdisciplinary team configurations, shape these delegation preferences and workflows.
Along with ownership, lack of trust in AI-generated codes was one of the main reasons limiting participants’ willingness to adopt AI-assisted QDA tools. Lee and See’s framework [67] provides a useful lens to understand the issue of trust. They operationalize trust as a combination of performance (i.e., reliability in achieving goals), process (i.e., clarity of inner workings), and purpose (i.e., alignment with user goals). We use this framework to organize our discussion and future work directions.
Participants assessed the performance of AI tools in QDA based on efficiency and reliability. Participants’ openness to AI as a supportive assistant was mainly due to anticipated efficiency, as they saw potential for such tools to automate or streamline repetitive tasks such as initial coding. They noted that providing AI with clear prompts and rules, combined with its high computational capacity, could simplify researchers’ roles to verifying outputs. On the other hand, people were largely skeptical of the reliability of the results due to the perception that the AI would not be able to capture nuanced, context-specific data accurately. This skepticism aligns with broader concerns around AI’s inability to fully grasp human meaning-making, raising doubts about the quality and credibility of its outputs.
Participants’ mixed trust in AI performance reflects findings in prior studies, which noted AI’s strengths in improving coding quality but questioned its effectiveness in handling complex interpretive tasks [68]. In prior work by Zhang et al. in the broader domain of AI-assisted decision making, the researchers explored two approaches of increasing trust: displaying confidence level and explaining the decision model [69]. Their findings indicated that users were more inclined to accept AI recommendations when higher confidence scores were shown, while model explanations did not have a significant impact on people’s acceptance. This contrasts our participants’ desire to see model explanations behind the codes that the AI has applied. It is possible that the ideal approach for enhancing trust differs by domains, and that in the context of QDA, model explanations may play a more critical role in building trust. Future work could explore this direction by comparing approaches such as showing confidence scores (i.e., correctness likelihood) when assigning codes, model explanation, or a list of alternative codes, on researchers’ trust and willingness to integrate AI into their workflow. Additionally, integrating mechanisms that allow researchers to customize the extent of AI involvement based on the dataset size or complexity might foster interest in using AI in QDA by giving them more agency.
In prior work, researchers emphasized the importance of transparency in AI’s coding processes and advocated for tools that can support marginalized populations while maintaining ethical integrity [70], [71]. If these tools reinforce dominant narratives or exclude outlier voices, they risk replicating harmful assumptions. Our results further confirmed this need for transparency. A potential way of making the AI analysis process more transparent is through clear communication about the process, such as explicit explanations of code suggestions (e.g., “This excerpt was coded as ‘lack of trust in AI’ as the participant expressed skepticism about AI’s ability to understand context.") while highlighting the relevant words or phrases in the transcript.** This would make the inner workings more transparent, allowing the user to identify and revise questionable AI outputs. This transparency is not only beneficial for incorporating AI, but also for supporting coordination and collaboration between researchers. More specifically, future systems could similarly enhance collaboration by visualizing coding discrepancies between researchers and offering AI-generated interpretations of the differences (e.g.,”Based on the sections coded as ‘trust in AI’, it seems like Researcher 1 interpreted the code as including both trust and mistrust in AI while Researcher 2 focused exclusively on sections showing trust, excluding those on mistrust.") From a delegation standpoint, our findings suggest that AI-assisted QDA should support how analytic work is divided between AI systems and human researchers while expecting systems to explicitly signal segments that require human sense-making, discussion, or judgment.** In this framing, delegation is not about replacing interpretation, but about directing human attention to analytically difficult moments where meaning must be negotiated.
Lastly, it is important to ensure that the system aligns with the user’s goals. Our interviews revealed that for many QDA researchers, these goals include maintaining strict control over data, especially when working with sensitive or private data. A similar finding was reported by Yan et al. [16], who found people’s skepticism toward ChatGPT due to doubts about data security. However, current QDA tools often provide limited information on where and how data is stored and processed. This lack of transparency even led some to avoid QDA tools altogether, resorting to manual coding or general-purpose software (e.g., MS Word) where local storage is ensured. The desire for local storage of research data presents a new challenge for AI-based QDA tools as there is a trade-off between performance and security. Typically, earlier and smaller language models can run locally (e.g., GPT-2), providing users with greater control over their data, which makes them more suitable for handling sensitive information (e.g., in healthcare or vulnerable populations) [72]. In contrast, newer and more advanced models like GPT-4 that offer enhanced performance are often cloud-based, making them difficult or impossible to run locally due to computational demands and proprietary restrictions [73]. This introduces potential risks related to storing and processing data on remote third-party servers, potentially exposing confidential research data to unauthorized access or breaches. As AI-assisted QDA tools continue to advance, decisions around model selection, data storage, and deployment infrastructure must be critically examined to balance performance with researchers’ ethical and privacy needs. At the very least, all related information should be readily accessible and comprehensive – including a clear description of where and how the data is stored (e.g., on a HIPPA compliant cloud server) and how the data gets used (e.g., whether it will be used for LLM training) – enabling researchers to make informed decisions based on their project requirements and institutional guidelines.
Bias is an inherent challenge in QDA, whether introduced by human subjectivity or algorithmic limitations [54], [74]. Participants recognized the inevitability of cognitive bias in human-led qualitative analysis and largely agreed that collaboration, whether among humans or between humans and AI, is key to mitigating bias. Some were already using AI tools like ChatGPT to validate their manual codes, suggesting that AI can act as an additional layer of review to challenge assumptions.
While AI can be used to reduce cognitive bias, concerns about algorithmic bias was a related theme. Algorithmic bias reflects the systematic inaccuracies embedded in AI tools due to flawed training data, limited contextual understanding, or improper algorithm design [54]. A prior finding uncovered the issue of algorithmic bias in QDA by noting ChatGPT’s lack of contextual understanding, which required researchers to upload literature reviews to provide background information [16]. Our participants echoed this concern that AI system’s failure to understand the contextual nuances of a dataset may produce biased outcomes that reflect an incomplete view of the data. For example, in a healthcare study, a participant might say, “I stopped taking my meds because I felt better,” which a general-purpose AI could misclassify as ‘positive health outcome.’ However, a trained medical researcher would recognize and tag this as a ‘potential non-adherence issue,’ an important red flag in chronic illness management due to factors such as mistrust in the prescription or socioeconomic barriers to continued use.
To help researchers detect and address bias, both their own and that of the AI, future QDA systems could incorporate features that encourage reflective practice. One approach is to ask the users to explicitly explain a sample of their codes and themes before the system attempts broader application of these themes to the rest of the transcript. Through this process, it would not only learn the contextual knowledge, but it could surface potential inconsistencies or bias and subtly prompt researchers to re-assess the appropriateness and neutrality of their coding. To combat confirmation bias, systems could also require users to enter their hypotheses and research questions before the analysis process, and summarize the resulting codes (e.g., visualizing the distribution of codes across participants). This could help users evaluate and better understand how the results are aligned or misaligned with their initial beliefs while preventing unconscious bias from influencing the results. This not only aids in addressing bias but could also ensure that all collaborators are aligned on the project’s goals and interpretive framework.
Based on our results and prior work, we hypothesize that AI could also support mitigating bias in QDA by leveraging multimodal data. Previous work has shown that analyzing textual transcripts may be insufficient to convey the full meaning of interview data, yet multimodal data is underutilized in QDA [23]. Our results further corroborate the limited use of multimodal data in qualitative analysis. This is a missed opportunity since audio and video recordings contain non-verbal cues (e.g., pauses, laughter, facial expressions) that offer additional insights even contradicting the spoken words at times. Overlooking these cues may reduce the quality of AI-generated codes as well. For example, an AI might misinterpret a participant’s sarcastic remark (e.g., “Oh great, another meeting”) as a positive sentiment, failing to recognize the tone of voice or facial expression that signals frustration.
We suggest future work to explore the integration of multimodal data within AI-driven QDA tools to help detect and summarize trends that would be too time-consuming for researchers to find manually. For example, as researchers scroll through interview transcripts, the system could highlight passages with emotionally significant cues (e.g., sustained pauses, changes in vocal pitch, or visible discomfort) using icons above the relevant phrases or modified background color. Hovering over a highlighted section could trigger a pop-up that displays relevant descriptions like “long pause,” “sarcastic tone,” or “averted gaze,” with playback options to hear or see the original moment. The system could also group and summarize recurring non-verbal trends (e.g., “increased vocal tension when discussing workplace policies”) across participants, helping researchers quickly spot patterns that might otherwise be missed in text-only analysis. This design would enable researchers to easily cross-reference the transcript with non-verbal cues, validate AI-generated codes, and/or navigate large datasets more efficiently by focusing on emotionally or behaviorally significant moments. This may reduce both cognitive and algorithmic bias by grounding analysis in a richer, more holistic view of participant responses.
This study had several limitations that should be considered when interpreting the findings. First, this study focuses on three degrees of AI involvement in QDA based on a single framework [22], but there are broader and more nuanced ways of involving AI within qualitative research. For example, AI involvement could extend to analytical coaxing (i.e., creative prompting to perform deeper analysis) [75], reflections on researcher’s high-level mental models [76], and supporting language comprehension [16]. Next, the gender distribution among participants was imbalanced, many were in technology-related field, and the sample size was relatively small, limiting the generalizability of the results. In addition, although the study included experienced researchers who have at least one year of experience in qualitative data analysis, their familiarity with QDA tools and AI tools varied, which may have affected their responses. Since our results showed that the desired role of AI in QDA is dependent on the researcher’s context, future study is needed to verify if the findings hold for researchers with a broader range of experiences, backgrounds, and demographic distribution in a setting where participants can flexibly negotiate AI roles throughout the analytic process. Another limitation of our work is that the discussion on the three degrees of delegation was based on textual descriptions rather than actual experience with a working tool. This hypothetical approach may not fully capture their experiences and perceptions in real-world interaction with such systems. Furthermore, we acknowledge that early exposure to the ChromaScribe prototype likely functioned as a framing device that anchored participants’ mental models around AI-generated thematic suggestions. Had participants interacted with a different dynamic, such as an AI that strictly retrieves data segments based on human-defined queries without suggesting its own themes, their ranking of delegation modes and specific anxieties around interpretive labor might have shifted. To note, while the study reports the number of participants who preferred each coding method for transparency, the research was primarily qualitative in nature. As such, the findings should be interpreted with an emphasis on thematic insights rather than statistical generalizability. Lastly, while our findings highlight that context shapes how researchers perceive and adopt AI in QDA, our study did not deeply examine how specific collaboration structures (e.g., solo work with supervision, distributed collaborations) influence delegation preferences. Future research should investigate how different collaboration scale (e.g., individual, small-team, and large-team) and norms (e.g., coordination practices, accountability mechanisms, and reflexivity) influence human-AI delegation in qualitative analysis.
Building on prior QDA studies that focused on AI’s effects on coding efficiency and reliability, our work revealed researchers’ desire for interpretive ownership, reassurance on data confidentiality, and support in collaborative coding and shared sense-making when considering integrating AI into their QDA workflow. While we utilized Lubars and Tan’s delegation framework to compare three interaction paradigms (human‑only, human‑initiated, and AI‑initiated coding), our findings suggest that researchers view AI-supported QDA not as fixed delegation, but as an evolving, hybrid interaction. The findings indicate openness to AI and a slight preference for AI-initiated coding over human-initiated coding, as it provides efficiency while preserving researchers’ ownership and accountability over interpretive outcomes. Participants also highlighted AI’s potential to reduce cognitive biases through complementary guidance rather than full automation, underscoring the value of transparent negotiation of meaning. Finally, although textual data remains the dominant modality for QDA, researchers recognized the importance of multimodal inputs to capture richer emotional and contextual information and sought tools to better integrate multimodal data in their workflow. Based on these results, we propose future work directions for designing AI as an assistant that align with the values of ownership, accountability, and shared sense-making to support collaborative QDA.