July 10, 2026
Recent advances in generative AI tools have significantly changed how software professionals write, evaluate, and interact with code. Generative AI tools such as GitHub Copilot, ChatGPT, and Claude are increasingly being integrated into everyday workflows. Despite the growing adoption of and reliance on these tools, it remains unclear as to how software professionals evaluate the code they generate. To explore this topic, we will conduct a constructivist grounded theory study that incorporates a survey, semi-structured interviews, and laddering interviews. With the initial survey data collection complete, we aim to interview 20–50 software professionals iteratively until theoretical saturation is achieved. This research aims to build a theory of how software professionals evaluate AI-generated code, grounded in their accounts of evaluative practices, perceptions, and preferences.
Professional software developers are increasingly adopting and relying on generative AI-based tools and agents in their work. The early literature on the topic suggested a shift from writing to reading, evaluating, and repairing generated code (e.g., [1]–[3]). These tasks are not always easy:
not being the author of the code may make it more difficult to understand [4]
it can contain bugs that are hidden or different from those authored by humans [3], [5]
it can be more difficult and time-consuming to debug and repair [3], [4], [6], [7]
the expectations of users do not always align with the tool’s capabilities [6], [8]
There are also broader issues such as tendencies to outsource critical thinking and comprehension effort, and the risks of automation bias, over-reliance, skill decay, and impaired learning [4], [6], [7], [9]–[13]. Reading code requires time and mental effort. Direct forms of evaluation require investing time and effort in code comprehension, which developers tend to avoid whenever task completion is a priority [14]. Generative AI provides ways to use more indirect forms of evaluation (e.g., asking AI to critique or test code) and to take shortcuts in comprehension and evaluation by offloading them to AI. In particular, the issue of over-reliance has been frequently reported in previous work [4], [6], [7], [15]–[19].
The use of AI-generated code is becoming increasingly prevalent in development settings shaped by productivity pressure from sources such as deadlines, competition, and rising expectations, alongside other practical constraints, including incomplete understanding, contextual ambiguity, and organisational policies. Software professionals must increasingly decide when to evaluate, rely on, reject, regenerate, or repair AI-generated code under these conditions while preserving the productivity gains the tools promise. However, such gains may be negated when careful evaluation requires dedicating time to reading, understanding, testing, regenerating, and repairing the code. As generative AI-based tools and agents become more capable in software development tasks, it may also become more difficult for software professionals to keep up with evaluating their outputs, decisions, and actions. Together, these conditions may create an incentive to rely on generated code with less attentive evaluation, increasing the risk of unwarranted trust and over-reliance.
These challenges remain important as long as human effort is still required to critically judge AI-generated code, and they extend beyond functional correctness to the code’s quality, alignment with intent and requirements, implementation and design decisions, and how professionals assume responsibility for all of these.
In this study, the research objective is to develop our understanding of these challenges from the perspective of software professionals. To this end, we start with an in-depth investigation of the perceptions, preferences, and practical conditions that shape software professionals’ evaluative practices in AI-assisted programming (assessment of the value of generated code, and potentially the decisions and actions of AI agents that lead to it), as well as the characteristics of these practices, guided by the following initial research question: RQ: How do software professionals approach the evaluation of AI-generated code?
This initial RQ is intentionally broad to keep the direction of the investigation open and let it evolve as the study progresses and questions of greater significance emerge [20]. To answer the RQ, the initial plan is to interview approximately 20–50 software professionals who have substantial first-hand experience in programming with generative AI. The target range is also intentionally broad because we plan to conduct two types of interviews (see Section 2), and the final number of participants will depend on when theoretical saturation is achieved. The initial interview design will be informed by an analysis of responses to a recent survey on software professionals’ use of generative AI in Finland.
Our data analysis and collection are guided by the constructivist grounded theory (GT) [20] as it aligns with our epistemological stance on qualitative research, which holds that knowledge is constructed rather than discovered (see Section 2.5). It also makes us attend more closely to how the research participants themselves construct the present topic.
While we acknowledge the importance of observing and measuring characteristics of evaluative practices, such as assessments of code quality, we place more importance on understanding how software professionals construct and reason about the evaluation of AI-generated code as the technology and its use are rapidly evolving. For example, what counts as sufficiently careful evaluation under practical constraints such as deadlines or incomplete understanding, or who is seen as responsible for problems originating from AI-generated outputs (and whether improvement is expected from the tool, the developer, or the surrounding review practices) are questions that cannot be answered without attending to the subjective experiences of those involved. Hence, using in-depth interviews as the primary method from a constructivist perspective seems suitable for answering the research question.
Before interviews, we conducted a non-exhaustive review of the existing literature to identify research gaps. The review was enough to reveal a research gap, or rather, a research need to do a more in-depth investigation. The closest prior work, and an inspiration for our topic, was Grounded Copilot: How Programmers Interact with Code-Generating Models [6] where Barke et al. identified programmers’ validation strategies, including “code examination, code execution (or testing), relying on IDE-integrated static analysis (e.g. a type checker), and looking up documentation”, as well as validation by “pattern matching”. These strategies form part of their theory regarding programmers’ interaction modes with GitHub Copilot. We also identified a few other grounded theory studies in adjacent topic areas (e.g., [21], [22]). For example, Li et al. [21] conducted a socio-technical grounded theory (STGT) study that uncovered individual and organisational motives, relationships, and challenges impacting software professionals’ adoption of AI tools.
Existing studies offer useful insights, but there remains much more richness in the phenomenon to be uncovered, as well as a need in the literature for a cohesive, evidence-based theory of evaluation practices themselves.
This section outlines the approach to data collection and analysis processes. Our approach is guided by the grounded theory method, specifically the constructivist variant [20]. We identified interviews as an appropriate method to start addressing the research question. The future direction of this study, including the choice of methods and research sites, is shaped by the emerging grounded theory analysis [20] (see Section 2.4), although prioritising interviews is preferred.
To begin this investigation, a survey was conducted targeting software professionals working in Finland. The survey addressed topics such as the extent of generative AI adoption and perceived effects. It also included questions about the professionals’ evaluation-related practices and trust. The full survey results will be reported separately; here, we focus on the parts of the survey that form the base for exploration in the present study (Section 2.1).
To gain a deeper understanding of how software professionals approach the evaluation of AI-generated code, we continue by investigating how they construct and reason about their evaluative practices. For this purpose, laddering interviews provide a suitable starting point, as they allow us to investigate why professionals prefer or avoid certain practices, how they perceive their consequences, and what value structures or goals underlie these perceptions and preferences (Section 2.2). We plan to complement the laddering interviews with semi-structured interviews to support more open-ended exploration and to pursue emerging directions that laddering interviews cannot effectively capture as the analysis develops.
We recently conducted a survey on the use of generative AI by software professionals in Finland, with 163 responses collected from December 2025 to February 2026. The questionnaire consisted of five parts: (i) generative AI tools in software engineering activities, (ii) generative AI in programming, validation practices for generated code, and trust, (iii) effects of generative AI on work, (iv) background questions, and (v) closing questions.
Part (ii) serves as the basis for our investigation. It included the following three optional open-ended questions:
In general, do you validate AI-generated code differently from human-written code? How? (For example, are there some common weaknesses you often look for?) (51 responses)
What practices help you stay productive with generative AI while managing the risks related to code quality? (79 responses)
Since you started using generative AI tools, has your trust in them changed in any way? If so, why? (103 responses)
These open-ended responses will be analysed following the constructivist GT approach (see Section 2.4). Part (ii) included ten additional closed-ended questions concerning, for example, the scope of generation (size and comprehension effort), validation practices (such as validation methods and timing), and trust in code generation tools. Although these closed-ended responses are not the main focus of the analysis, their quantitative analysis may nonetheless provide useful insights for the GT analysis.
Other parts of the questionnaire may similarly be used when relevant. For example, those who reported using generative AI in code reviews in part (i) were asked to briefly describe how. In part (v), respondents were asked about their broader thoughts and predictions about the future role of AI in software development, and some responses touched on evaluation. The other parts of the questionnaire also provide contextual information to aid the analysis, including demographic information, such as the years of experience in software development and roles or job titles (see Table 1), as well as information on the frequency of using generative AI and AI-generated code in software development, and the programming languages used with generative AI.
The questionnaire was available in both Finnish and English versions. It was designed and piloted iteratively until the pilot participants reported no difficulties in understanding or answering the questions or issues in translation quality. We recruited participants through convenience and snowball sampling using personal and professional social networks, LinkedIn advertisements, and various email lists and chat groups as recruiting channels, and by encouraging our contacts to forward the survey invitation to potential participants among their acquaintances.
| Characteristic | \(\mathrm{n}\) | % | |
|---|---|---|---|
| Years of experience | |||
| < 10 | 79 | ||
| 10–19 | 40 | ||
| 20–29 | 23 | ||
| 30+ | 21 | ||
| Role (job title) | |||
| Full-Stack Developer | 37 | ||
| Software Architect | 15 | ||
| Back-End Developer | 12 | ||
| DevOps Developer | 11 | ||
| Data Analyst / Data Engineer / Data Scientist | 10 | ||
| Development Manager | 7 | ||
| CEO / CIO / CTO | 7 | ||
| Front-End Developer | 7 | ||
| QA / Test Engineer | 6 | ||
| Product Owner / Product Manager | 5 | ||
| Project Manager | 4 | ||
| UI / UX Designer | 3 | ||
| Business Analyst | 2 | ||
| Enterprise Architect | 2 | ||
| Other | 35 |
A note on terminology: In the initial survey, the term validation was used instead of evaluation. The term was adopted from the Grounded Copilot study [6], where the authors used it “broadly, to encompass any behavior meant to increase user’s confidence that the generated code matches their intent.” The same definition was used in our questionnaire. In a related large-scale survey [7], the questionnaire items based on the findings of Grounded Copilot study used the term evaluation instead.1 However, validation was found in survey piloting to be more easily understood, as it conveys more strongly the idea of confirming that the generated code is valid.
In this GT study, our interest is initially broader and directed at all practices that software professionals use to evaluate or assess the value of AI-generated source code-related outputs or actions (e.g., their functional correctness, quality, alignment with intent or requirements, and trustworthiness) without limiting our interest to practices confirming their validity/conformance to user needs (via validation), or veracity/conformance to specifications (via verification) in any strict sense. In this study, we will use the term evaluation, although as the GT emerges, some different term may later prove more appropriate.
This study will use traditional, semi-structured interviews in conjunction with laddering, which is an in-depth interview technique [23] used as an initial data collection method. The method has been used in consumer research to elicit consumers’ “ways of thinking” regarding a product or service category [23] (see [24] for an example). It involves the use of repeated why questions as probes to move from concrete attributes (typically of a product), or rather, in this context, means (e.g., using/not using AI to review AI-generated code), to the higher-level consequences of those means. For example, asking “Why is it important to you not to use AI for reviewing AI-generated code?” can give insights into the consequences perceived by the interviewee (e.g., reduced code awareness). Ultimately, this technique uncovers the personal relevance of these consequences in the form of values or goals which in this example could be professional integrity (a value is reached when the interviewer notices that the why-questions have reached an “end”, which is sometimes better described as a “goal” rather than a value).
In this study, the identified attribute-consequence-value chains (i.e., means-end hierarchies [23]) serve as the basis for theorising using the GT method. With an initial focus on how practices (means) relate to values (ends), we may uncover interesting patterns, e.g., where do programmers prefer to invest evaluation effort? or how do programmers approach situations in which careful evaluation conflicts with demands for increased productivity? The identified values and goals can be assumed to be somewhat stable underlying drivers of practices. Practices can also be shaped by company policies, goals, and values, which can likewise be identified using laddering. Although concrete instances of evaluation practices vary across contexts, there may be value in knowing why programmers tend to prefer or avoid specific approaches. This can, for example, be useful for predicting how these practices change as generative AI-based tools improve.
The interviews will be conducted in Finnish and English. They are expected to last an average of 40 minutes. Each interview starts with a list of stimuli, which includes a few scenarios of what generative AI can do, from straightforward code generation to more autonomous actions. The scenarios are used to examine how participants construct and reason about evaluation under varying practical conditions, initially focusing on scenarios where it is challenging for humans to keep up. The findings of our initial survey can be used as a basis for the scenarios (see Section 2.1). If the survey data are insufficient for developing sufficiently relevant or varied scenarios, the list may also be compiled by synthesising responses to a pre-interview survey sent when contacting potential participants.
At the beginning of each interview, the stimuli are ranked by the interviewees according to their relevance. For the highest-ranking stimulus, the interviewee is asked to relate the stimulus to concrete experiences from their own work, and to come up with a few scenarios they consider representative. For each scenario, we can ask questions such as: “What would you evaluate in this?”, “What kinds of evaluation strategies would you use?”, or “How do you make sure you can trust the output?” The actual questions must be refined by piloting the interview guide. The answers are the attributes/means, which are then followed by a series of why-questions. Attribute-consequence-value chains are created for all attributes listed by the interviewee. These chains will be integrated into the constructivist GT analysis by coding the attributes, consequences, values, and links between them, comparing chains within and across participants, and using these comparisons alongside semi-structured interview transcripts and the initial survey responses to develop GT categories.
The interview guide is expected to be revised as the analysis continues (cf. [25]). Although the stimuli list does not change in a typical laddering study, it may change in this study if new directions emerge during the grounded theory analysis; developing the theory is the priority, not comparability and compatibility of individual attribute-consequence-value chains.
The laddering is followed by a ranking of the chains and background questions. The chain ranking is based on the relevance of the chains for the interviewees and provides useful information for analysis. This also gives the interviewees an opportunity to provide feedback on their accuracy. Background information such as software development experience and familiarity with AI tools will provide contextual information for the analysis. Each interview concludes with a question about the interviewee’s interest in being contacted for more questions about the topic or participating in future research.
During the laddering interview, the interviewer writes down the interviewee’s statements in a condensed and structured format to facilitate the interview. An audio recording device is used to verify the accuracy of the written information after the interview. In addition, the transcripts can be used in data analysis if they are found to contain insights that cannot be captured by the chains. All files will be uploaded either to NVivo2 or Quirkos3 for analysis.
Initially, we plan to recruit participants from the 66 respondents who expressed openness to being contacted for additional discussion on the survey topic by providing their email addresses. 44 potential interviewees remain after excluding the software professionals working in academia (19 out of the 163) and those not using (or only rarely using) AI-generated code in their software development work. Since all survey participants were from Finland, we aim to maximise the demographic diversity of the initial interview sample during the first round of data collection. Following the GT approach, the subsequent rounds are more concerned with what is useful for the emerging theory than with generalisability. That is, the main concern of theoretical sampling is to saturate emerging theoretical categories and their properties, not statistical generalisability. Finding potential interviewees from the survey sample is the priority, but the sampling frame is not restricted to it.
This study follows the constructivist grounded theory method [20]. The core features of the GT method include simultaneous data collection and analysis, theoretical sampling, theoretical sensitivity, coding, memo-writing, constant comparison, and memo sorting, with the intention of reaching theoretical saturation and cohesive theory [26]. Data collection will be performed until theoretical saturation is achieved, i.e. until no new themes or insights emerge from iterative data collection as the emerging theory is sufficiently supported by the collected data [26]. Charmaz’s constructivist reinterpretation of the method is one of the available variants, most notably Glaser’s GT [27], [28] and Strauss and Corbin’s GT [29].
Survey data, laddering chains, and interview transcripts provide the data that is analysed using the constructivist GT method. The first step is initial coding, followed by focused coding and theoretical coding. Initial coding refers to analytically examining and labelling data word-by-word, line-by-line, or incident-by-incident, while remaining open to the theoretical possibilities suggested by the data (not injecting the researcher’s assumptions, biases, or motivations) [20], [26]. In focused coding, the most significant or frequent initial codes are used to categorise the data.
Constant comparison and memo-writing are also integral parts of the GT data analysis. Constant comparative methods are used to make comparisons at each level of the analysis process, to find similarities and differences between data (e.g., statement-to-statement comparisons within the same or different interviews), codes, memos, and categories [20], [26]. Memo-writing refers to writing analytic notes (or diagrams, sketches, etc.) about the codes, data, and theoretical categories as they emerge throughout the research process [20], [26]. They serve diverse purposes, including capturing important ideas and hunches as they appear, documenting the analysis, discovering gaps in data collection, demonstrating connections between categories, and providing material for writing the research article.
The main reason for choosing the constructivist version is that we tend to agree with Charmaz’s stance, including the view that grounded theories are constructions of reality rather than exact pictures of it. The researcher is not a passive, value-free, and neutral observer that “discovers” facts from data—“their values shape the very facts that they can identify” [20]. Therefore, the researcher is challenged to be more sensitive about what they inevitably bring to the research setting and how this influences their findings.
However, for this study, we do not fully commit to social constructivism, at least in any strong sense. Charmaz also encourages flexible engagement with the method rather than enforcing rigid prescriptions, e.g., with respect to epistemological stances [20]. We see the value of the constructivist lens in acknowledging how meanings and analyses are constructed, and we let it inform the epistemological stance for this study. The ontological stance for this study is more aligned with that of critical realism. To be more transparent about our position, our research interest extends beyond these constructions to the underlying mechanisms. Although it is difficult to say at this point where the theory will lead, our preliminary aim is to study a reality that is (at least partially) independent of these constructions, yet suggested by them—even though the resulting view of this reality will inevitably be a construction itself, imperfect and probabilistic.
The first author is a PhD researcher whose understanding of software engineering and generative AI has been developed through formal education, research, and personal projects. His use of generative AI tools, while frequent, remains cautious and selective. As he has not worked as a software professional in industry, he begins the investigation of the phenomenon from an informed and curious but somewhat distanced position. Although he acknowledges that applications of generative AI can have clear benefits in software engineering, he treats all narratives and claims about generative AI as suspicious. He is particularly concerned with how AI is used uncritically to displace, devalue, or obscure human skill, judgement, and autonomy, and with the threat of professionals and organisations becoming dependent on AI providers. This positionality makes him pay special attention to the distinctions between actual and ostensible benefits, short-term and long-term consequences, and the narratives that users tell each other and themselves.
The second author is a PhD researcher who has performed extensive research on code quality measurement, their standards, and how these measurements are validated. She also has research on challenges and issues faced while applying code quality metrics to assess code quality which could influence her choice of questions in the interview guide, as she could be unintentionally comparing the practices and challenges faced in human code evaluation to those faced during agentic code evaluation. Her experience in the software development industry is limited and dates back approximately a decade. Her exposure to AI code generation has primarily been through hobby projects and the generation of small, preliminary code snippets within larger projects. She believes that, despite the advantages of AI assisted development especially in terms of time saved, AI cannot replace human written code. While she considers the use of AI for generating small code snippets acceptable, she would not trust or feel comfortable relying on AI for project level code generation. Consequently, she approaches AI generated code with caution and subjects it to careful scrutiny.
The first two authors are responsible for the collection and analysis of the data, while the remaining authors provide advice and feedback throughout the process. Individually, we are committed to seeking interpretations and perspectives that challenge our own thinking and preconceptions. As a team, we will manage reflexivity by having frequent meetings to review each other’s transcripts, codes, and memos, to discuss the coding process and emerging theory, and to scrutinise them for any biases related to our preconceived opinions and experiences on the topic. In addition, as recommended by Charmaz [20], we plan to use methodological journaling to document “methodological dilemmas, directions, and decisions”, as a way to engage in reflexivity.