What’s Beyond Copilot?
22 AI Systems Developers Want Built

Looking Beyond Copilot:
22 AI Systems Developers Want Built

You Shall Not Pass!
Where and Why Developers Draw The Line on AI Autonomy


1 Introduction↩︎

Generative AI-powered development tools (hereafter AI tools) are transforming how software gets built. AI tools such as GitHub Copilot [1], Claude [2], and in-house assistants today write code, fix bugs, and generate tests, and their capabilities keep expanding. The question is no longer whether AI can do the work; it is should it, which parts, and how much?

This quandary is not new. Long before AI, every wave of automation forced the same reckoning over what to hand to the machine [3][5]. What is different now is the pace and the reach. AI moves faster and reaches deeper into developers’ work, so a single decision to defer can have downstream impacts. If we are not deliberate about this decision and let it settle based on AI tool capabilities, the result can be a set of serious compounding problems  [6][10].

As AI absorbs tasks, developers can miss the work that builds skill and grounds judgment [8], [9], [11], leaving them to “rubber-stamp” work they no longer understand well enough to evaluate. Pressure to ship leads them to offload more to keep pace [6], and as code generation grows cheaper, the same pressure plays out across the pipeline [12]: review, testing, and operations absorb rising volumes of machine-generated work. By the time defects surface, that work has passed through many hands, which makes it costlier to fix [13], [14] and accountability harder to assign [15], [16].

Prior work has examined what drives AI adoption [17], [18] (e.g., trust [19][22]), which tasks suit automation [23], [24], and where developers want support [25], [26], but not where they want control to remain human. Without that boundary, we cannot design work, or the tools that structure it, that keeps developers meaningfully engaged. Building on Endsley’s levels of automation [27], we distinguish two boundaries that mark it: the action boundary, where developers let AI act on their behalf rather than produce work they review, and the decision-making boundary, where they let AI make act on their behalf.

We posit that supporting meaningful work in settings where AI plays a role requires understanding how developers cognitively appraise their work. Drawing on cognitive appraisal theory [28], [29], work design theory [30], [31], and automation theory [27], we examine how developers’ task appraisals along dimensions of relevance, identity congruence, accountability, and demands inform the autonomy they grant AI. More specifically, we investigate:

RQ1.

What level of AI autonomy do developers accept across software engineering tasks, and what predicts where they draw the line?

RQ2.

What predicts when developers let AI cross the action and the decision-making boundaries in their work?

We answer these questions through a survey of 448 professional developers, who informed where they did and did not want AI involved across 20 software engineering tasks spanning the development lifecycle. We classified 1,535 open-ended responses onto a five-level autonomy scale adapted from Endsley and Kiris [27], mapped the accepted AI autonomy levels across Software Development Life Cycle (SDLC) categories, and modeled the accepted level and its two critical transitions (action and decision-making boundaries) against developers’ task appraisals and their traits. We interpret these results through the lens of meaningful work design and frame that design as a flight of Cascading Locks: a sequence of gates that hold or cede autonomy depending on whether the conditions for meaningful work are met. This framing lets us identify patterns and anti-patterns for meaningful work design.

2 Related Work↩︎

As AI tools become standard in software engineering (SE) workflows, a growing body of work has examined what drives their adoption [17], [18], [32].

Studies based on technology-acceptance models found that workflow compatibility and habitual use are strong drivers of adoption. Trust also emerged as a recurring factor, shaped by tooling capabilities [19], [21], [22], individual dispositions [18], [25], and team factors [6], [33]. More recent studies moved from overall adoption to task-level differences. Lambiase et al. [24] found higher AI receptivity for artifact-manipulation and information-retrieval tasks and lower receptivity in collaborative contexts. Pereira et al. [26] observed stronger adoption for code-intensive work and more limited use in creative aspects. Khemka et al. [23] reported strong demand for AI in testing, debugging, documentation, and compliance, while Kumar et al. [34] showed that toil-heavy activities such as documentation and environment setup were disproportionately viewed work to minimise and thus strong candidates for AI support.

Closest to our work, Choudhuri et al. [25] investigated how developers cognitively appraise different aspects of their work, along dimensions of value, identity, accountability, and demands, to explain where they want or resist AI support across the software lifecycle. That work showed that desired AI involvement was shaped not only by tool capability or trust, but also by what the work meant to developers, and that human oversight remained important even where AI support was welcome. Wanting AI support, however, is not the same as granting AI autonomy. Accepting suggestions, accepting AI-produced artifacts, and allowing AI to act without approval represent different forms of involvement. Work on delegation made this distinction explicit: ceding work to AI is a decision about authority rather than use [35]. Ulloa et al. [36] similarly found that product managers were less willing to hand over work they identified with or felt accountable for. Even though these studies suggest that work appraisals shape delegation decisions, they do not characterize the level of autonomy developers grant once AI starts to be used.

The automation literature characterizes autonomy as a graded allocation of authority between humans and machines. Endsley and Kiris [27] and Parasuraman et al. [37] organized automation into levels defined by who decides and who acts, ranging from advisory assistance to full automation. Higher levels have been associated with risks including automation complacency, reduced situation awareness, and out-of-the-loop failures. This perspective underpins machine-in-the-loop design, task-delegability frameworks [38], and analyses of work better suited to automation than augmentation [39]. Within SE, autonomy has primarily been studied within limited settings. Ghorbani et al. [40], for example, found that developers preferred lower-autonomy tools, although more experienced developers were more receptive to higher-autonomy ones. These framings, however, largely described what tools could do rather than the autonomy developers were willing to grant them.

Building on these foundations, we investigate where and why developers draw the line on AI autonomy across SE work. We characterize the level of autonomy developers grant AI, examine how those preferences vary across the software lifecycle, and model how developers’ cognitive appraisals and individual characteristics predict both accepted AI autonomy and the transitions between autonomy levels.

3 Theory and Hypotheses↩︎

When working with machines, humans have to decide how much of the work to keep and how much to cede.

Levels of automation: The literature treats this division of control between humans and machines as a spectrum [27], [37], [41]. We adopt the five-level scale from Endsley and Kiris [27], ordered by who decides and who acts, to characterize where developers draw the line on AI autonomy (Table 1).

Table 1: The five-level autonomy scale, adapted from Endsley and Kiris [27]. Each level reflects who decides and who acts; AI autonomy increases down the rows.
Level of Automation Developer AI
L1  None Decide, Act
L2  Decision Support Decide, Act Suggest
L3  Consensual Decide, Act Act
L4  Monitored Veto Decide, Act
L5  Full Automation Decide, Act

We chose it over frameworks that decompose automation across information-processing stages, such as Parasuraman et al. [37], which assigns separate levels to acquisition, analysis, decision, and action. A single ordered scale gives one level of human–AI control per task, which keeps levels comparable across the heterogeneous tasks in SE work.

Table 1 shows these five levels, ordered by how much AI decides and acts without the developer. At L1, the developer decides and acts, and AI has no role. At L2, AI suggests or flags while the developer decides and produces the work. At L3, AI may decide and act, but the result takes effect only after the developer signs off. At L4, AI decides and acts by default, and the developer retains an optional veto. At L5, AI decides and acts on its own with minimal human involvement.

The two transitions on this scale that impact work design revolve around the kind of control a developer cedes. At the action boundary (L2\(\to\)L3), AI moves from advising to producing the artifact or carrying out the action. At the decision-making boundary (L3\(\to\)L4), the developer’s sign-off moves from required to optional. Hypotheses: Predictors of levels of automation: The goal of our work is to investigate where and why developers cede autonomy to AI and to what extent. While trust is an obvious precursor to whether developers use AI, it concerns the “can AI do it” question, which prior work has addressed [19], [20], [22], [42]. In this paper, we investigate how developers answer the “should AI do it and to what extent” question.

Humans are meaning-makers; we seek significance and value in our work [31]. At work, we implicitly evaluate a task by asking: Is this important to me? Does it align with what I want to do? Am I responsible if it fails? Can I handle its demands? Cognitive appraisal theory [28], [29] formalizes these judgments across dimensions of importance, congruence with one’s motivations or identity, accountability, and cognitive demand. These appraisals shape how people cope with work [43] and predict engagement, persistence, and discretionary effort [44]. Work-design research [30] adds that motivational, social, and contextual job characteristics explain much of the variance in work engagement and productivity.

Choudhuri et al. [25] showed that these appraisals predict developers’ openness to and use of AI. Building on this, we use the appraisals to predict how much autonomy developers cede to AI in SE tasks. Specifically, we study four appraisals: Value and Identity are motivational appraisals [31], Accountability a social one [45], and Demand a contextual one [46].

Task Value is the perceived importance of a task to project success and to personal goals [47]. Valuable work raises both the benefit of finishing it faster and the cost of getting it wrong; developers treat the work that matters most as the work they most want to keep under their own control. H1. Higher task value lowers the AI autonomy developers accept. We expect developers to retain autonomy for tasks they consider valuable.

Task Identity is the degree to which a developer identifies with a task and enjoys it for its own sake [48]. Identity-defining tasks confer ownership and a sense of craft [49], [50], which we hypothesize developers protect by keeping them under their control. H2. Higher task identity lowers the AI autonomy developers accept.

Task Accountability is the perceived responsibility a developer feels for a task’s outcome [51]. Anticipated answerability raises vigilance and the wish to oversee the work before signing off as their own [45], [52]. The more autonomy they grant AI, the less they can check, so we hypothesize: H3. Higher task accountability lowers the AI autonomy developers accept.

Task Demand is the cognitive effort a task imposes [46]. Demanding work strains a developer’s resources and increases the likelihood of offloading work [7], [46], so developers should give AI higher autonomy where demand is high. H4. Higher task demand raises the AI autonomy developers accept.

Task type. Appraisals are tied to the type of work one performs. For example, human-facing work such as mentoring, designing, and stakeholder communication carries the meaning, judgment, and relationships that make work feel one’s own [30], [31], whereas well-scoped or rote system tasks such as test generation, environment setup, and documentation carry less of it. Where a task lies along that range should relate to how much autonomy a developer is willing to give AI. We therefore expect the autonomy developers accept to differ across the SDLC categories. H5. The AI autonomy developers accept differs across task types.

Crossing the boundaries. Letting AI draft work developers sign off on (action boundary) and letting AI act unless developers stop it (decision-making boundary) are very different levels of autonomy, as defined in Table 1. Both are conditions that tie to meaningful work design: people experience work as meaningful when they author it and when they hold the decisions that govern it, and that sense of ownership is what sustains engagement and intrinsic motivation [30], [31], [48], [49]. Each boundary removes one of these conditions, so we model them separately and expect the appraisals to act on each as they do on the overall level. H6. Higher task value, identity, and accountability lower the odds that developers cross each boundary; higher task demand raises the odds.

Controls. We control for developers’ SE and AI experience, since both shape attitudes toward AI [53], [54]. Experience can calibrate these judgments: time spent building software and time spent working with AI tools both temper what a developer expects AI to handle and how readily they delegate to it. Individual traits can also condition how appraisals translate into the autonomy a developer accepts. Prior work has shown that individuals’ risk tolerance and technophilic motivations [20], [55] are associated with stronger AI-adoption dispositions. We expect these traits and experience measures to moderate the hypothesized appraisal effects.

4 Method↩︎

Our goal was to investigate where and why developers drew the line on AI autonomy across their daily SE tasks. We studied professional software developers across Microsoft, an AI-forward technology company employing more than 50,000 developers worldwide, spanning a diverse set of products, teams, roles, processes, and stakeholder contexts. The scale, combined with participants’ exposure to both emerging and mature AI tools, makes it a rich setting for our study.

4.1 Study Design↩︎

We surveyed developers about how they cognitively appraised different aspects of their daily work and where they did or did not want AI involved, following the appraisal-survey design of [25]. We classified the open-ended responses onto a five-level autonomy scale (Section 4.2) and modeled the result against task appraisals and developer traits, characterizing what set the autonomy level developers accepted and the transitions where they granted AI action and decision-making authority. Our study was approved by the company’s IRB.

4.1.1 Survey instrument design↩︎

We followed Kitchenham’s survey guidelines [56] and drew on established theoretical frameworks and validated scales (Table 2) to capture how developers appraise their daily work and how much autonomy they would grant AI on it. To represent that work, we adopted the grounded taxonomy of software engineering (SE) tasks across five SDLC categories (Table 3) from [25]. We refined the survey instrument through iterative pilots with researchers and developers. The suervey was structured as follows.

Table 2: Theoretical constructs and instruments [25]
Construct Instrument
Value Job Characteristics Model [57], [58]
Identity Self-Determination Theory [48]
Accountability Felt Accountability Scale [51]
Demands Job Demands-Resource Model [46]
Levels of Automation Endsley & Kiris Framework [27]
Risk Tolerance, Technophilia Cognitive Style Facet Survey [20], [25]

After providing informed consent, developers reported their SE and AI tool experience and their dispositions toward AI (risk tolerance and technophilia), then selected the 2–3 categories from Table 3 that best reflected their work. To reduce fatigue, the meta-work category (applicable to all developers) was not a default option and appeared only for participants who selected two other categories, ensuring that no participant answered more than three category blocks.

Within each category block, participants rated how they appraised each task along four dimensions: value, identity, accountability, and demands [25], which served as our model predictors. We used single-item measures to keep the survey tractable and reduce participant fatigue, as these retain psychometric validity for well-scoped constructs [59].

Table 3: Grounded SE task taxonomy from [25].
SDLC Category Tasks
Development Coding, Bug fixing, Performance optimization, Refactoring, AI integration
Design & Planning System design, Requirements engineering, Project planning & management
Quality & Risk Testing/QA, Code review, Security & compliance
Infrastructure & Ops DevOps (CI/CD), Environment setup & maintenance, Infrastructure monitoring, Customer support
Meta-work Learning, Research, Documentation, Stakeholder communication, Mentoring

Each block then posed two open-ended questions about where developers wanted and did not want AI: “Where do you want AI to play the biggest role for your [category]-heavy activities over the next 1–3 years?” and “What aspects of your [category]-heavy activities do you not want AI to handle and why?” The autonomy preferences are contingent on task type and contextual constraints, and the open-ended questions let preferences emerge naturally. We subsequently classified responses onto the five-level autonomy scale (Section 3).

We administered the survey in Qualtrics [60]. Closed-ended items used a 5-point Likert scale with a sixth “I’m not sure” option. The survey took 10–12 minutes to complete. To support data quality and reduce response bias, we included attention checks and randomized question order within blocks.

We refined the instrument in pilot rounds with 40 developers, checking clarity and task coverage. These pilot responses were excluded from the analysis. The survey instrument is in [61].

4.1.2 Data collection↩︎

We emailed the survey to a random-stratified sample of 8,000 developers across varied roles, teams, and regions, with one follow-up a week later. Participation was voluntary and anonymous. The survey returned 1,193 responses, a 14.9% response rate consistent with prior large-scale SE surveys [18], [62]. We dropped incomplete (\(n{=}152\)), straight-lined or otherwise patterned responses (\(n{=}59\)), those that failed an attention check (\(n{=}98\)), and ones from developers reporting no AI experience (\(n{=}24\)).

The autonomy classification depended on the open-ended responses, so we also dropped participants whose responses were unsubstantive (\(n{=}412\)). A field counted as substantive when it held content beyond an empty string or a set of filler tokens (e.g., “n/a,” “yes,” “no,” “x,” “nil,” “nope,” “idk,” or “no comment”). One researcher applied this rule to every open-ended field. The remaining 448 developers provided 1,535 task-responses (we describe the decomposition in Section 4.2.1), of which the 1,476 with complete predictor data feed the regression models in Section 5. Most respondents were based in North America and self-identified as men, spanning a wide range of SE and AI experience. Demographics are in [61].

4.2 Data Analysis↩︎

Our analysis proceeded in two steps. We first classified each open-text task response onto the five-level autonomy scale (Table 1), then modeled the classified responses against task appraisals and developer traits (Section 3). Then, we conducted reflexive thematic analysis of the same responses to characterize which aspects of work developers wanted to retain at each autonomy level.

4.2.1 Autonomy level classification↩︎

Each open-text response set recorded, in participants’ own words, the role they wanted AI to play in a task-category (Want) and the role they refused to grant it (Resist). The pair defines the ceiling, the highest level of autonomy a participant tolerated before Resist blocked more. To classify the responses, we used a rule-based algorithm (Algorithm 1) that classified each response to an autonomy level. When the text was ambiguous between two adjacent levels, fuzzy labeling assigned a boundary level rather than forcing one. The algorithm comprised the following components:

Figure 1: Autonomy-level classification of a (participant, category) response. Tasks(\cdot) splits each field into per-task rows; returns the corresponding autonomy-level.

Marker and hyponym tables. The algorithm used two tables in the classification process (see [61]). The hyponym table mapped surface phrases onto canonical tasks (e.g., write code, programming, code-related all resolved to the Coding task), which the algorithm then used to pair Want and Resist responses when both named the same task and split them otherwise. The marker table was a set of autonomy features, each a closed set of case-insensitive substrings: Refusal (AI does no part of the task), Suggestion (AI suggests), Action (AI acts), Artifact (AI produces an artifact), Approval (the human signs off before AI’s work takes effect), Veto (AI decides and acts, the human can override afterward if needed), Monitoring-only (AI acts, the human only watches), and No-human-control (AI acts without notable human involvement). We synthesized both tables from the pilot responses, then finalized them against the full response set to capture forms the pilot missed. To guard against circularity, the external coder (described below) saw neither table during development.

Unit of analysis. The unit of analysis was the triple (participant, category, task). Participants often described several tasks within a single response set (per category) and named different tasks across the two fields. Tasks(\(\cdot\)) therefore separated each field into one row per task using the hyponym table. We then paired Want and Resist when both named the same task and split them otherwise, setting each row’s source to Want, Resist, or Both. This analysis yielded 1,535 task-responses.

Level assignment. For each row, the classify() subroutine applied a fixed sequence of binary tests ordered by increasing AI autonomy, stopping at the first marker that fired (see Algorithm 1). The tests capped each row at the highest autonomy the participant tolerated: any refusal or approval requirement was reached before higher-autonomy markers. Around 8% of rows received a boundary label (e.g., L2-L3, L3-L4). For the descriptive distribution (Figure 2), we conservatively collapsed each to its lower adjacent level (full distribution in [61]). The regression models instead entered each as an interval-censored observation between its two levels, preserving its encoded uncertainty.

AI-council-based classification. We ran algorithm 1 with an AI-council comprising an ensemble of three frontier Large Language Models from distinct families (gpt-5.4, gemini-3-flash, claude-sonnet-4.6; high-reasoning mode [63], [64]). We selected the three from different providers to reduce model-specific blind spots, inductive biases, and failure modes [65][67]. Agreement across model families provided stronger convergent evidence than agreement within one, where shared data and optimization objectives could produce shared errors [68]. No model saw the others’ outputs. We reconciled the three label sets by majority vote: a label held when at least two coders agreed.

Human validation and IRR. To verify the classification, two researchers independently re-coded a random sample of 67 participant-response sets from the 448 participant pool (\(\pm10\%\) margin at 95% confidence). One researcher was not involved in developing the algorithm and acted as an external check against circularity. We measured inter-rater reliability (IRR) with Krippendorff’s \(\alpha\) complemented by pairwise quadratic-weighted Cohen’s \(\kappa\) and three-rater percent agreement [69]. Researcher agreement was \(\kappa = 0.94\), human–council agreement \(\kappa = 0.95\) and \(0.92\), and overall three-rater agreement \(\alpha = 0.93\), indicating strong agreement for the analyses that follow.

4.2.2 Mixed-methods analysis↩︎

We modeled the coded task-responses against participants’ cognitive appraisals and individual characteristics (recall Section 3). Because the design involved repeated measures within participants and across tasks, we used mixed-effects regression [70]. For RQ1, we fit a Cumulative Link Mixed Model (CLMM) on the five-level ordinal scale against z-standardized predictors, accommodating the four fuzzy boundary labels (e.g., L2–L3) as interval-censored observations [71], [72]. For RQ2, we fit a logistic Generalized Linear Mixed Model (GLMM) at each of the two autonomy boundary transitions (action: L2\(\to\)L3, decision-making: L3\(\to\)L4), modeling whether a participant’s accepted autonomy crossed that boundary. All Variance Inflation Factors (VIFs) were \(<2\) [73], indicating no multicollinearity concerns. Model specifications appear in Section 5.

To understand how participants reasoned about levels of AI autonomy across aspects of their work, we conducted reflexive thematic analysis [74], [75] of the classified open-ended responses. We inductively coded participants’ explanations and developed themes characterizing the work aspects tied to each autonomy level. Two researchers iteratively refined the codes and themes, with the research team resolving disagreements through negotiated agreement. The resulting themes organize Figure 2 and explain participants’ autonomy preferences in Section 5. Participants are referenced as P1–P448 hereafter.

4.3 Threats to Validity↩︎

Construct validity. We measured constructs with self-reported items grounded in established theory. Still, surveys can introduce bias or misinterpretation. We mitigated this by involving practitioners in survey design, piloting, randomizing blocks, adding attention checks, and screening patterned responses. We elicited autonomy preferences as open-ended text rather than fixed ratings, since the level a developer deems appropriate is contingent on contextual constraints that cannot be reliably assessed psychometrically. Finally, to mitigate forced categorization, we introduced fuzzy labels for boundary-ambiguous responses, treating them as interval-censored data in the primary model, and confirmed that the conclusions were robust under alternative operationalizations (Section 5.1).

Internal validity. As a cross-sectional study [76], we report associations, not causation. Self-selection is possible, since developers with stronger views about AI may be more likely to respond. As with all survey-based work, our results reflect self-reported perceptions. To guard against classifier bias, we ran an AI council of three models on a fixed, rule-based algorithm, reconciled by majority vote [65], [68], with an external recoder checking against circularity, with strong agreement throughout (Section 4.2.1). Still, inferring a level from brief free text is imperfect. We strengthened validity by triangulating the quantitative results with qualitative data and aligning with theory. The company’s pro-AI culture, if anything, biases against us, making the boundaries developers still refuse a conservative read of where they want to retain autonomy.

External validity. All participants were from one AI-forward multinational organization, spanning diverse teams, roles, geographies, domains, and processes. This limits generalizability to smaller organizations or open-source contexts. Despite this limitation, we argue our findings are of value to the broader community, because it has been shown that single-case studies make meaningful contributions to scientific discovery [77] that complement research on broader populations [78].

5 Results↩︎

Section 5.1 presents where and why participants drew the line on AI autonomy, followed by Section 5.2 that presents what predicted whether participants crossed the action boundary (L2\(\to\)L3) and the decision-making boundary (L3\(\to\)L4).

5.1 RQ1: Accepted Autonomy Levels and their Predictors↩︎

We classified all 1,535 participant responses on the five-level Endsley–Kiris autonomy scale (Table 1), coding each response as the highest autonomy a participant accepted for the task (L1 = no AI involvement through L5 = full automation). Figure 2 places each SDLC category across the five levels, reporting the share of responses at each level, the proportion of autonomy participants kept versus ceded, and representative aspects of work named at each cell.

Figure 2: Participant response distribution (n{=}1{,}535 responses; 448 developers) across the five autonomy levels. Rows group responses by SDLC category, with the top row pooling all categories; columns order the levels L1 to L5, defined by who decides and who acts. Each row reports the share of responses at each level, and the proportion of autonomy participants kept versus ceded to AI, split at each category’s median accepted level; each cell lists representative aspects of work mentioned for that category and level.

5.1.1 Autonomy Level Ceilings↩︎

Finding 1: Most participants accepted AI-produced work without ceding decision-making to AI (\(\le\)L3). The median accepted level was L3, where AI produces the artifact and it takes effect only after the developer’s explicit approval. 74% of responses fell at or below L3, whereas 26% let AI decide and proceed without required approval (L4, optional veto: 10%; L5, full automation: 16%). Participants accepted AI across a broad range of work, including inline refactoring, authoring tests, CI/CD pipeline setup, creating user stories from requirements, and authoring docs from PRs (Figure 2, \(\le\)L3). As P219 put it: “I’d love for [AI] to generate the code, but in a way that makes it easy for me to follow along and correct it” (P219, L3).

Participants wanting to retain oversight is intuitive and consistent with HCI work on user control over intelligent assistive systems [79] and a preference for human supervision of automation [38], [80]. We now turn to the nuance: where and why participants placed that oversight, and how it varied with the task type, participants’ cognitive appraisals, and individual dispositions towards AI support in work.

Finding 2: In system-facing work, most participants delegated bounded execution and verification tasks, while retaining oversight on context-sensitive judgments (L3). The three system-facing categories (Development, Quality & Risk, and Infrastructure & Ops) all had a median accepted autonomy level of L3 (Figure 2): most participants positioned AI primarily as a collaborator operating under their approval. As P106 explained, “I would like to remain a code-reviewer for my AI-generated code and basically steer it in the right direction” (P106, L3).

Within the Quality & Risk category (26% of responses \(>\)L3), participants were most willing to delegate work they characterized as routine verification or operational toil. They accepted high levels of autonomy (\(\ge\)L3) for automatically generating tests, detecting security risks, configuring adversarial or chaos-testing environments, creating pull requests from security reports, and performing security and compliance scans. One participant envisioned “more automated tests built by AI that even a PM could specify what they want to test and AI handles it” (P36, L5). In contrast, participants insisted on retaining oversight and approval authority for activities requiring contextual evaluation, such as code review, pre-merge security assessment, and test-selection decisions. P196 noted, “I would be happy for AI to take over code review if it’s capable. Testing is the same. Both would require mandatory human post-review” (P196, L3). The distinction was not necessarily between technical and non-technical work. Rather, participants differentiated between tasks they viewed as automatable toil and those requiring interpretation, prioritization, or judgment to guard against complacency and automation bias [81].

Finding 3: In design and human-facing work, most participants limited AI to decision support (L2), retaining action and decision-making. These categories had a median accepted level of L2: participants welcomed AI suggestions but wanted to remain the primary actor and decision-maker, citing concerns about AI’s judgment, vision, empathy, and accountability. P276 explained, “I don’t want AI to handle final judgment calls on ambiguous or high-stakes decisions, because these often require human intuition, contextual awareness, and accountability that AI can’t fully replicate or own” (P276, L2). This preference was strongest for activities participants viewed as inherently human. Within Meta-work, mentoring, sensitive client communication, learning, and creative thinking were often retained entirely at L1: “Mentoring and onboarding shouldn’t be done by an AI. It’s a human thing that shouldn’t be handed off to technology” (P101, L1). Participants described these activities as opportunities to develop expertise, cultivate others, and exercise interpersonal judgment.

At the same time, they readily accepted AI assistance for drafting and proofreading communications, organizing work items, gathering learning resources, and brainstorming (L2). P125, for example, described using AI to support learning: “when learning new technologies, I generally have very specific questions that I might feel embarrassed to ask a more knowledgeable person...and get an [AI] expert and specific response back is transformative in the learning process” (P125, L2). Documentation was delegated further (L3), reflecting more comfort with AI codifying and organizing existing knowledge than with decisions that shape people, strategy, or growth.

5.1.2 Predictors of accepted AI autonomy↩︎

The descriptive analysis identified where participants drew the line on AI autonomy; we next examined what factors predicted that line and why participants facing the same task diverged. To investigate this, we fit a cumulative-link mixed model (CLMM) with a logit link [82] over the \(n{=}1{,}476\) (of 1,535) task responses with complete predictor data. The model included z-standardized appraisal and trait predictors (recall Table 2), SDLC category as a fixed effect (Development as the baseline), and a participant-level random intercept to account for repeated measures. The ordinal outcome was the accepted autonomy level, with the four fuzzy boundary labels treated as interval-censored observations [72]. All predictors were estimated within a single model. For appraisal and trait measures, we report standardized latent-scale coefficients (\(\beta\)), representing the change in accepted autonomy associated with a one-standard-deviation increase in the predictor. For SDLC categories, we report odds ratios (OR) relative to Development (baseline). Model fit is summarized using marginal and conditional \(R^2\) [83], [84]. Results are in Table ¿tbl:tab:clmm?.

@lcccc@ Predictor & \(\beta\) & 95% CI & \(f^2\) & \(p\)

Task Value & 0.06 & [-0.07, 0.18] & 0.03 & 0.373
Task Identity & -0.13 & [-0.24, -0.02] & 0.07 & 0.017*
Task Accountability & -0.07 & [-0.19, 0.05] & 0.04 & 0.276
Task Demand & 0.05 & [-0.06, 0.17] & 0.03 & 0.373

SE Experience & -0.10 & [-0.33, 0.13] & 0.06 & 0.393
AI Experience & 0.26 & [0.07, 0.45] & 0.14 & 0.007**
Risk Tolerance & 0.33 & [0.07, 0.60] & 0.18 & 0.013*
Technophilia & 0.02 & [-0.22, 0.25] & 0.01 & 0.898

& OR & 95% CI & \(f^2\) & \(p\)
Design & Planning & 0.42 & [0.31, 0.57] & 0.48 & \(<\)\(\textbf{.001***}\)
Quality & Risk & 1.67 & [1.18, 2.38] & 0.28 & 0.004**
Infrastructure & Ops & 0.79 & [0.58, 1.19] & 0.13 & 0.053
Meta-work & 0.33 & [0.24, 0.46] & 0.61 & \(<\)\(\textbf{.001***}\)
&

Finding 4: Among the four appraisals, only Task Identity predicted accepted AI autonomy, and negatively. Task Identity was negatively associated with accepted autonomy (\(\beta{=}{-}0.13\), \(p{=}0.017\)), supporting H2: the more participants identified with a task, the less autonomy they ceded to AI. Task Value, Accountability, and Demand were not significant predictors (no support for H1, H3, or H4). Participants retained autonomy over the tasks that defined them as professionals even when they accepted that AI could do it: “I do not want to hand AI a design and have it write all the code because I enjoy coding and the challenges of solving problems in programming.” (P176, L2). This finding aligns with self-determination theory, under which the activities people pursue for their own sake are the ones that satisfy the needs for autonomy and competence [48], aspects contributing to one’s sense of craft, source of achievement, and self worth [49], [50].

Finding 5: Participants’ AI experience and Risk Tolerance were positively associated with their accepted AI autonomy level. AI Experience (\(\beta{=}0.26\), \(p{=}0.007\)) and Risk Tolerance (\(\beta{=}0.33\), \(p{=}0.013\)) were both associated with higher accepted autonomy. Participants with more experience using AI tools and those more comfortable with uncertainty were more willing to let AI act with less oversight. As P288 (“very experienced” with AI and reporting high risk tolerance) put it: “Development in general should be taken over by AI. If the human can correctly word out what needs to be done, the actual implementation shouldn’t require too much oversight” (P288, L5). Neither SE Experience nor Technophilia predicted the accepted level: years of building software did not by itself predict ceding autonomy to AI, nor love for new technology.

Finding 6: Task type predicted the participants’ accepted autonomy level. Task type made a difference. Relative to Development (baseline), participants accepted lower autonomy for Meta-work (OR\(=0.33\), \(p<0.001\)) and Design & Planning (OR\(=0.42\), \(p<0.001\)). Meta-work (mentoring, communication, onboarding, and learning) had the strongest resistance. Participants preferred to retain control over activities requiring interpersonal judgment, intent, and relationship management: “I don’t want [AI] to help mentor people. Relationships are important” (P85, L1). Another explained, “I want AI suggestions for communication with external partners...I don’t want wholesale retranslation” (P182, L1). Participants also limited autonomy in Design & Planning, viewing these activities as highly consequential and costly to reverse. P10 remarked, “AI can support but shouldn’t lead, as misalignment here can derail entire projects” (P10, L2).

In contrast, Quality & Risk had higher odds of accepted autonomy than Development (OR\(=1.67\), \(p=0.004\)). Participants were more willing to delegate testing and review activities, often framing AI as an additional safeguard for identifying “bugs, regressions, [performance] bottlenecks, and potential security issues early” (P241, L3). Although participants typically retained final oversight (Finding 3), verification was already embedded within these activities, reducing the perceived cost of AI action. Infrastructure & Ops did not have statistically signigicant differences from Development. These findings support H5, indicating that accepted autonomy varies across SDLC categories, with significant differences observed for Meta-work, Design & Planning, and Quality & Risk.

Robustness check. We evaluated two alternative statistical operationalizations as robustness checks for our findings. First, a linear mixed-effects model reproduced the same directional patterns as the CLMM: AI Experience (\(\beta = 0.12\), \(p = 0.009\)) and Risk Tolerance (\(\beta = 0.16\), \(p = 0.014\)) were associated with higher accepted autonomy, whereas Identity was associated with lower accepted autonomy (\(\beta=-0.07\), \(p = 0.023\)). SDLC-category effects were also preserved, with lower accepted autonomy for Design & Planning (\(\beta = -0.48\)) and Meta-work (\(\beta = -0.56\)), and higher accepted autonomy for Quality & Risk (\(\beta = 0.19\)), all \(p < 0.05\). Second, a binary logistic mixed model that dichotomizes responses at the decision-making boundary (\(\le\)L3 vs.\(>\)L3) reproduces the CLMM in direction and significance: Task Identity reduced the odds of accepting autonomy above L3 (OR = 0.85, \(p = 0.018\)), whereas AI Experience (OR = 1.29, \(p = 0.010\)) and Risk Tolerance (OR = 1.41, \(p = 0.013\)) increased those odds. All three SDLC-category contrasts were preserved: Design & Planning (OR = 0.40), Meta-work (OR = 0.32), and Quality & Risk (OR = 1.69), all \(p{<}0.05\), relative to Development. Overall, the findings were robust to treating the outcome as ordinal, binary, or continuous. Full coefficients are reported in [61].

5.2 RQ2: Crossing the Action and Decision-Making Boundaries↩︎

RQ1 identified the predictors of developers’ overall accepted AI autonomy. In RQ2, we investigate whether the same predictors operate when AI gains authority to act (the action boundary, L2\(\to\)L3) and decide (the decision-making boundary, L3\(\to\)L4). To do so, we fit a GLMM at each of these transitions (Table ¿tbl:tab:transitions?) with the same predictors as RQ1.

6pt

@lcc@ Predictor & Action Boundary & Decision-Making Boundary
& (L2\(\rightarrow\)L3) & (L3\(\rightarrow\)L4)

Task Value & 0.95 [0.79, 1.13] & 1.05 [0.87, 1.27]
Task Identity & 0.89 [0.74, 1.07] & 0.78* [0.64, 0.95]
Task Accountability & 0.82* [0.70, 0.98] & 0.91 [0.76, 1.09]
Task Demand & 0.97 [0.83, 1.14] & 1.20* [1.05, 1.42]

SE Experience & 0.92 [0.68, 1.24] & 0.84 [0.62, 1.14]
AI Experience & 1.27 [1.00, 1.61] & 1.29* [1.01, 1.65]
Risk Tolerance & 1.40 [1.00, 1.96] & 1.46* [1.03, 2.07]
Technophilia & 0.91 [0.67, 1.23] & 1.06 [0.78, 1.45]

\(R^2_m, R^2_c\) & 0.103, 0.574 & 0.065, 0.543

Finding 7: Accountability was negatively associated with the action boundary, lowering the odds of crossing it. Task Accountability lowered the odds of crossing the L2\(\rightarrow\)L3 boundary (OR = 0.82, \(p{=}0.027\)). Put simply, when participants felt highly accountable for a task, they were less willing to let AI produce work artifacts, even while remaining open to AI suggestions: “I do not want AI to ‘handle’ any of it, given that I am accountable for it. But I welcome AI systems that provide feedback on my design and implementation” (P174, L2). As P120 explained in the context of code review, “marking my approval puts my name on it and makes me partially responsible for it” (P120, L2). This pattern is consistent with prior work showing that accountability increases vigilance and the desire to oversee outcomes before committing to them [45], [52]. The effect emerged precisely at the transition where AI outputs become reusable work artifacts and the consequences of automation errors become attributable to the individual, making automation complacency a more salient concern [81].

Finding 8: At the decision-making boundary, Task Identity and Demands pulled in opposite directions. By L3, participants had accepted AI as a collaborator capable of producing work artifacts under oversight. Crossing into L4 meant accepting AI’s decisions by default with an optional veto, shifting the developer role from decision-maker to overseer. Task Identity was associated with lower odds of crossing this boundary (OR = 0.78, \(p{=}0.01\)), whereas Task Demand was associated with higher odds (OR = 1.20, \(p{=}0.03\)).

While accountability predicted participants’ willingness to let AI act (Finding 7), the L3\(\rightarrow\)L4 boundary concerned who retained final judgment once AI had acted. Participants who viewed a task as central to who they were as professionals were more willing to accept AI-generated work than to relinquish the judgment surrounding it. As one participant explained, “I wouldn’t want AI to handle final decision-making in high-stakes risk scenarios, especially where ethical judgment, human intuition, or deep domain context is critical [...] the responsibility should remain with experienced professionals” (P10, L3). Prior work links identity to ownership, autonomy, and authorship over valued work outcomes [48][50]; our findings suggest that this attachment is expressed most strongly at this boundary.

Task Demand exhibited the opposite pattern. Participants facing heavier demands were more willing to delegate decision-making, suggesting that the cost of exercising oversight can outweigh the value of retaining it. Consistent with job demands–resources theory [46], workload increased the appeal of transferring not only execution but also approval and coordination effort. One participant remarked, “As much as I love coding, there is simply too much work to do to want to keep doing it myself” (P56, L5). Together, these results suggest that crossing the decision-making boundary reflects a trade-off between preserving decision-making on identity-relevant work and offloading it when task demands become burdensome.

Overall, H6 was partially supported. Task Accountability was associated with lower odds of crossing the action boundary (L2 \(\rightarrow\) L3), whereas Task Identity and Task Demands were associated with lower and higher odds, respectively, of crossing the decision-making boundary (L3 \(\rightarrow\) L4).

Finding 9: AI Experience and Risk Tolerance were positively associated with crossing the decision-making boundary. AI Experience (OR = 1.29, \(p{=}0.044\)) and Risk Tolerance (OR = 1.46, \(p{=}0.034\)) were associated with higher odds of crossing the L3\(\rightarrow\)L4 boundary. While Finding 8 showed that task appraisals predicted whether participants retained decision-making authority, these traits distinguish how readily participants were willing to grant it. For example, at one end, P276 saw little reason to withhold autonomy: “[…] I have been using AI since the start of any incubating idea to shipping the product” (P276, L5, Risk Tolerance = 5). At the other end, more risk-averse participants held back: “The more technical the task is, the more I feel the need to have personal confidence in its correctness. I don’t usually get that confidence when leaning heavily on [AI]” (P368, L2, Risk Tolerance = 1).

Overall, the action and decision-making boundaries answered to different appraisals: accountability to whether AI may produce the artifact, identity and demands to whether AI may decide without sign-off, with AI experience and risk tolerance easing that crossing.

6 Cascading Locks of Meaningful Work Design↩︎

Designing meaningful work means understanding how developers appraise it. Individuals are meaning-makers; we seek significance and value in what we do [31], and decades of work-design research [30], [46] show that how we appraise a task, along its relevance, identity, accountability, and demand, accounts for much of the variance in work satisfaction and productivity [28][30]. Two of these appraisals hold the line on AI autonomy. A developer who feels accountable for a task keeps AI at suggestion and produces the artifact themselves (Finding 7). A developer for whom a task is central to their identity does not relinquish their decision-making (Finding 8).

We picture this as a flight of cascading canal locks. The vessel is AI’s autonomy on a single task. (A developer may stay at L2 for a security-critical task, but cede autonomy at the L4 level for rote fault-isolation scans). The locks raise it one gate at a time. Accountability is the gate of the first lock, at the action boundary. Identity is the gate of the second, at the decision-making boundary. The locks cascade because they run in order: the vessel reaches the second gate only after the first has let it through, so a developer faces the identity decision only once accountability has cleared AI to produce the artifact. The water that lifts the vessel is AI fluency and risk tolerance. Both raise the autonomy developers accept and the odds of clearing each gate (Findings 5, 9). Demand adds a second inflow at the upper gate (Finding 8): a heavier workload makes developers readier to hand off decision-making to AI.

Viewed through the lens of cognitive appraisal, our findings point to anti-patterns (see Table ¿tbl:tab:antipatterns?) that appear when tool defaults set these boundaries. Each theme below discusses how these appraisals can inform meaningful work design.

6pt

@p0.13 l p0.64@ Theme & Anti-pattern & Description
& Deferred answerability & Review steps are automated; responsibility offloaded until no one owns the result.
& Right-shifted quality & Quality gates migrate to verification; errors surface late and cost more to fix.
& Throughput stampede & Incentives reward velocity; workload metric conflate toil and challenge as one quantity causing work shedding that once defined self-worth and identity.
& Shrinking fence & Craft core defined by AI capability gaps, which shrinks with each release.
& Hollow orchestrator & Role title remains, but person cannot follow what they need to approve. Veto becomes nominal.
& Expertise commoditization & AI does the underlying work; the expertise that distinguished the role stops being scarce.
& Thought homogenization & Developers defer to the same AI for framing; judgment converges on the model shrinking independent thinking and innovation.
& Severed pipeline & Entry-level work automated to increase throughput; juniors lack scaffolding before facing hard problems.
& Cognitive debt accrual& Juniors offload learning work to AI; skill gaps hide until AI fails.
& Mentoring-by-bot & Onboarding automated; relationships needed to pass tacit knowledge never form.

Accountability retention: keep answerability where the artifact is shaped. Most system-facing work is at L3 today: AI produces and a developer approves before the work takes effect. At L4 only an optional veto remains, and answerability moves from before the commit to after the failure. This is the deferred answerability anti-pattern: the developer sees the artifact when verification reports a problem, not when it is built. The accountability literature explains the cost. A developer who expects to answer for a decision beforehand reasons carefully and self-critically; one who answers for it only after the result is known defends the result instead [45], [52]. Two design choices can help. First, the L4 veto has to require real engagement at the commit, so approval is more than just a “rubber stamp”. Second, developers should answer for how they reviewed the work, not only for the shipped product. Both keep answerability concrete and attributable to a person [51], and both belong in how the role and workflow are built, and not delegated to a call for individual diligence [36], [85].

Demand differentiation: separate toil from work that builds judgment. Demand was the one appraisal that pushed toward more AI autonomy at the decision-making gate (Finding 8), and whether that push is safe depends on the type of demand. The job-demands literature distinguishes work that only drains effort from work that drains it while building skill and mastery [46], [86]. Handing AI the repetitive, well-scoped work that consumes attention and teaches nothing costs a developer nothing. The cognitively hard work that builds judgment is the work human oversight depends on. A workload metric that counts only volume cannot tell the two apart, so under a deadline it offloads both. This is the throughput stampede anti-pattern. The design choice is to measure the two demands separately: route toil to AI on purpose, and protect the hard, judgment-building work from the same automation pressure rather than let a volume target decide its fate.

Identity evolution: anchor the protected core in judgment. A developer’s sense of craft can rest on one of two things: the work AI cannot do yet, or the standard of judgment the role demands. The first is a moving target. Each model release does part of what defined the core, the shrinking fence, until the title remains but its real work is gone, the hollow orchestrator. It fails in the labor market too: once AI does the underlying work, any developer can direct it, and the expertise that once set a developer apart is no longer scarce, the expertise commoditization anti-pattern [87]. The same deference to AI homogenizes a team. Finally, as developers lean on AI for framing and first drafts, they start from the same suggestions and reason from the same defaults, so their judgments converge on the model’s, dissent thins, and the team loses the range of independent thinking that open-ended work depends on [55].

Self-determination theory [48] explains why an identity built on what AI cannot yet do is unstable. Ownership comes from work a developer values for its own sake, not from defending tasks against obsolescence, and once a tool closes the gap, that basis for identity is gone [48], [50]. Identity economics adds the market half: an identity has value when it ties a person to a standard others recognize, and a standard of judgment persists even as capabilities become widely available [88].The design choice is to define the core by the judgment the role requires and to write that into role definitions and promotion criteria, where it is organizational rather than a personal bet on what AI will not reach, and where developers reason from their own judgment rather than defaulting to the model’s.

Identity formation: protect the work that trains judgment. Judgment is built by doing the work, which is embodied in the entry-level work that seems most delegable to AI. Automate those tasks away and the skills needed to reach senior levels never form, the severed pipeline. Two failures feed it. Juniors who hand such work to AI are unaware of skill gaps, the cognitive debt [9]; and a bot that manages onboarding delivers explicit knowledge but none of the tacit knowledge that only passes between people, mentoring-by-bot. The competence and ownership that ground a developer’s identity come from doing work pursued for its own sake [48], [49], so the design choice is to keep a share of entry work as practice rather than throughput. The developers who need to exercise judgment over AI-mediated systems learn it on exactly the work most tempting to automate away.

7 Conclusion↩︎

Our study of 448 developers found that the two boundaries around AI autonomy answer to different appraisals. Task accountability predicts whether AI may act and produce work artifacts. Task identity and demands predict whether it may decide on the developers’ behalf. These appraisals reflect how developers relate to their work, so the accepted autonomy level can shift as AI capability, developer experience, and the work itself co-evolve. AI experience and risk tolerance already move it (Finding 5, 9), and the next model release may move it again. What is likely to hold steady underneath is the need for work that sustains meaning, competence, and identity.

A design response pinned to a list of tasks to withhold from AI is therefore brittle, whereas one grounded in meaningful work design endures. We offer the pairing of cognitive appraisals and our analysis instruments as a living framework that organizations can rerun as roles and tools change, to see where and why their developers draw these boundaries and whether adoption is breeding the anti-patterns in Table ¿tbl:tab:antipatterns?.

We have a choice to set these boundaries deliberately, through how roles, workflows, and jobs are designed, or let tool defaults reset them with each release. The first path keeps the work meaningful as automation advances.

References↩︎

[1]
Microsoft.2024. Copilot. https://copilot.microsoft.com.
[2]
Anthropic.2025. Claude. https://claude.ai/.
[3]
Paul M Fitts.1951. (1951).
[4]
Joost CF De Winter Dimitra Dodou.2014. . Cognition, Technology & Work16, 1(2014), 1–11.
[5]
David H Autor.2015. . Journal of economic perspectives29, 3(2015), 3–30.
[6]
Courtney Miller, Rudrajit Choudhuri, Mara Ulloa, Sankeerti Haniyur, Robert DeLine, Margaret-Anne Storey, Emerson Murphy-Hill, Christian Bird, and Jenna L Butler.2025. . arXiv preprint arXiv:2507.21280(2025).
[7]
Zixuan Feng, Sadia Afroz, and Anita Sarma.2025. . arXiv preprint arXiv:2510.07435(2025).
[8]
Margaret-Anne Storey.2026. . arXiv preprint arXiv:2603.22106(2026).
[9]
Rudrajit Choudhuri, Christopher A Sanchez, Margaret Burnett, and Anita Sarma.2026. . SSRN(2026).
[10]
Sadia Afroz, Zixuan Feng, Tyler Menezes, Katie Kimura, Bianca Trinkenreich, Igor Steinmacher, and Anita Sarma.2026. . ACM International Conference on the Foundations of Software Engineering.
[11]
Zixuan Feng, Reed Milewicz, Emerson Murphy-Hill, Tyler Menezes, Alexander Serebrenik, Igor Steinmacher, and Anita Sarma.2025. . ACM Transactions on Software Engineering and Methodology(2025). https://doi.org/10.1145/3789210.
[12]
Rudrajit Choudhuri, Christian Bird, Carmen Badea, and Anita Sarma.2026. . arXiv preprint arXiv:2604.07830(2026).
[13]
Barry Boehm, Victor R Basili, et al2005. . Foundations of empirical software engineering: the legacy of Victor R. Basili426, 37(2005), 426–431.
[14]
Frederick Brooks H Kugler.1987. No silver bullet. April.
[15]
Christian Bird, Nachiappan Nagappan, Brendan Murphy, Harald Gall, and Premkumar Devanbu.2011. . In FSE. 4–14.
[16]
Foyzur Rahman Premkumar Devanbu.2011. . In Proceedings of the 33rd international conference on software engineering. 491–500.
[17]
Daniel Russo.2024. . ACM Transactions on Software Engineering and Methodology(2024).
[18]
Rudrajit Choudhuri, Bianca Trinkenreich, Rahul Pandita, Eirini Kalliamvakou, Igor Steinmacher, Marco Gerosa, Christopher Sanchez, and Anita Sarma.2025. . arXiv preprint arXiv:2505.17418(2025).
[19]
Brittany Johnson, Christian Bird, Denae Ford, Nicole Forsgren, and Thomas Zimmermann.2023. . In 2023 IEEE/ACM 45th International Conference on Software Engineering: Software Engineering in Practice (ICSE-SEIP). IEEE, 409–419.
[20]
Rudrajit Choudhuri, Bianca Trinkenreich, Rahul Pandita, Eirini Kalliamvakou, Igor Steinmacher, Marco Gerosa, Christopher Sanchez, and Anita Sarma.2025. . In 2025 IEEE/ACM 47th International Conference on Software Engineering (ICSE). IEEE, 1691–1703.
[21]
Adam Brown, Sarah D’Angelo, Ambar Murillo, Ciera Jaspan, and Collin Green.2024. . In Proceedings of the 1st ACM International Conference on AI-Powered Software. 1–9.
[22]
Ruotong Wang, Ruijia Cheng, Denae Ford, and Thomas Zimmermann.2024. . In Proceedings of the 2024 ACM conference on fairness, accountability, and transparency. 1475–1493.
[23]
Mansi Khemka Brian Houck.2024. Commun. ACM67, 11(2024), 42–49.
[24]
Stefano Lambiase, Gemma Catolino, Fabio Palomba, Filomena Ferrucci, and Daniel Russo.2025. . arXiv preprint arXiv:2504.02553(2025).
[25]
Rudrajit Choudhuri, Carmen Badea, Christian Bird, Jenna Butler, Robert DeLine, and Brian Houck.2026. . In Proceedings of the 48th International Conference on Software Engineering: Software Engineering in Practice (ICSE-SEIP).
[26]
Guilherme Vaz Pereira, Victoria Jackson, Rafael Prikladnicki, André van der Hoek, Luciane Fortes, Carolina Araújo, André Coelho, Ligia Chelli, and Diego Ramos.2025. . In 2025 IEEE/ACM 47th International Conference on Software Engineering: Software Engineering in Practice (ICSE-SEIP). IEEE, 330–341.
[27]
Mica R. Endsley Esin O. Kiris.1995. . Human Factors37, 2(1995), 381–394. https://doi.org/10.1518/001872095779064555.
[28]
Richard S Lazarus.1991. Emotion and adaptation. Oxford University Press.
[29]
Ira J Roseman Craig A Smith.2001. . Appraisal processes in emotion: Theory, methods, research(2001), 3–19.
[30]
Stephen E Humphrey, Jennifer D Nahrgang, and Frederick P Morgeson.2007. Journal of applied psychology92, 5(2007), 1332.
[31]
Marjolein Lips-Wiersma Lani Morris.2009. . Journal of business ethics88, 3(2009), 491–511.
[32]
Christian Bird, Denae Ford, Thomas Zimmermann, Nicole Forsgren, Eirini Kalliamvakou, Travis Lowdermilk, and Idan Gazit.2022. . Queue20, 6(2022), 35–57.
[33]
Ruijia Cheng, Ruotong Wang, Thomas Zimmermann, and Denae Ford.2023. . ACM Transactions on Interactive Intelligent Systems(2023).
[34]
Sukrit Kumar, Drishti Goel, Thomas Zimmermann, Brian Houck, Balasubramanyan Ashok, and Chetan Bansal.2025. . In 2025 IEEE/ACM 47th International Conference on Software Engineering: Software Engineering in Practice (ICSE-SEIP). IEEE, 12–22.
[35]
Jobin Alexander Strunk, Leonardo Banh, Anika Nissen, Gero Strobel, and Stefan Smolnik.2024. . (2024).
[36]
Mara Ulloa, Jenna L Butler, Sankeerti Haniyur, Courtney Miller, Barrett Amos, Advait Sarkar, and Margaret-Anne Storey.2025. . arXiv preprint arXiv:2510.02504(2025).
[37]
Raja Parasuraman, Thomas B Sheridan, and Christopher D Wickens.2000. . IEEE Transactions on systems, man, and cybernetics-Part A: Systems and Humans30, 3(2000), 286–297.
[38]
Brian Lubars Chenhao Tan.2019. . Advances in neural information processing systems32(2019).
[39]
Yijia Shao, Humishka Zope, Yucheng Jiang, Jiaxin Pei, David Nguyen, Erik Brynjolfsson, and Diyi Yang.2025. . arXiv preprint arXiv:2506.06576(2025).
[40]
Amir Ghorbani, Nathan Cassee, Derek Robinson, Adam Alami, Neil A Ernst, Alexander Serebrenik, and Andrzej Wąsowski.2023. . In 2023 IEEE/ACM 45th International Conference on Software Engineering (ICSE). IEEE, 1405–1417.
[41]
Thomas B Sheridan William L Verplank.1978. . (1978).
[42]
John D Lee Katrina A See.2004. . Human factors46, 1(2004), 50–80.
[43]
Tavis S Campbell, Jillian A Johnson, and Kristin A Zernicke.2020. . In Encyclopedia of behavioral medicine. Springer, 486–487.
[44]
John P Meyer Natalie J Allen.1991. . Human resource management review1, 1(1991), 61–89.
[45]
Philip E Tetlock.1983. Journal of personality and social psychology45, 1(1983), 74.
[46]
Arnold B Bakker Evangelia Demerouti.2007. . Journal of managerial psychology22, 3(2007), 309–328.
[47]
J Richard Hackman Greg R Oldham.1976. . Organizational behavior and human performance16, 2(1976), 250–279.
[48]
Richard M Ryan Edward L Deci.2000. American psychologist55, 1(2000), 68.
[49]
William A Kahn.1990. . Academy of management journal33, 4(1990), 692–724.
[50]
Richard Koestner, Natasha Lekes, Theodore A Powers, and Emanuel Chicoine.2002. Journal of personality and social psychology83, 1(2002), 231.
[51]
Angela T Hall, Dwight D Frink, and M Ronald Buckley.2017. . Journal of Organizational Behavior38, 2(2017), 204–224.
[52]
Jennifer S Lerner Philip E Tetlock.1999. Psychological bulletin125, 2(1999), 255.
[53]
Jenna Butler, Jina Suh, Sankeerti Haniyur, and Constance Hadley.2025. . In 2025 IEEE/ACM 47th International Conference on Software Engineering: Software Engineering in Practice (ICSE-SEIP). IEEE, 319–329.
[54]
Kevin Crowston Francesco Bolici.2025. . Information Research an international electronic journal30, iConf(2025), 1009–1023.
[55]
Andrew Anderson, Jimena Noa Guevara, Fatima Moussaoui, Tianyi Li, Mihaela Vorvoreanu, and Margaret Burnett.2022. . ACM Transactions on Interactive Intelligent Systems(2022).
[56]
Barbara A Kitchenham Shari L Pfleeger.2008. . In Guide to advanced empirical software engineering. Springer, 63–92.
[57]
Yitzhak Fried Gerald R Ferris.1987. . Personnel psychology40, 2(1987), 287–322.
[58]
Bianca Trinkenreich, Fabio Santos, and Klaas-jan Stol.2024. . ACM Transactions on Software Engineering and Methodology33, 8(2024), 1–45.
[59]
Russell A Matthews, Laura Pineault, and Yeong-Hyun Hong.2022. . Journal of Business and Psychology37, 4(2022), 639–673.
[60]
Qualtrics.2026. Qualtrics Survey Platform. https://www.qualtrics.com. .
[61]
[n. d.]. Supplemental Package. https://zenodo.org/record/21060399.
[62]
Margaret-Anne Storey, Thomas Zimmermann, Christian Bird, Jacek Czerwonka, Brendan Murphy, and Eirini Kalliamvakou.2019. . IEEE Transactions on Software Engineering47, 10(2019), 2125–2142.
[63]
Zackary Okun Dunivin.2024. . arXiv preprint arXiv:2401.15170(2024).
[64]
Angjelin Hila Elliott Hauser.2025. . arXiv preprint arXiv:2507.14384(2025).
[65]
Moaath Alshaikh, Tasneem Alshaher, Ricardo Vieira, Beatriz Santana, Clelio Xavier, Jose Amancio, Glauco Carneiro, Julio Leite, Savio Freire, and Manoel Mendonca.2026. . In Proceedings of the 1st International Workshop on Prompt Engineering for Software Engineering (PROMPT-SE 2026). .
[66]
Zijian Li, Luzhen Tang, Mengyu Xia, Xinyu Li, Naping Chen, Dragan Gašević, and Yizhou Fan.2026. . In Proceedings of the 16th International Conference on Learning Analytics and Knowledge (LAK ’26). .
[67]
Julian Ashwin, Aditya Chhabra, and Vijayendra Rao.2026. . Sociological Methods & Research55, 3(2026), 795–839.
[68]
Shuai Wu, Xue Li, Yanna Feng, Yufang Li, Zhijun Wang, and Ran Wang.2026. . arXiv preprint arXiv:2604.02923(2026).
[69]
Kilem L Gwet.2014. Handbook of inter-rater reliability: The definitive guide to measuring the extent of agreement among raters. Advanced Analytics, LLC.
[70]
Andrew Gelman Jennifer Hill.2007. Data analysis using regression and multilevel/hierarchical models. Cambridge university press.
[71]
Donald Hedeker Robert D Gibbons.1994. . Biometrics(1994), 933–944.
[72]
Minge Xie, Douglas G Simpson, and Raymond J Carroll.2000. . Biometrics56, 2(2000), 376–383.
[73]
Joseph F Hair.2009. . (2009).
[74]
Virginia Braun Victoria Clark.2006. . Qualitative research in psychology3, 2(2006), 77–101.
[75]
Virginia Braun Victoria Clarke.2022. Qualitative Psychology9, 1(2022), 3.
[76]
Klaas-Jan Stol Brian Fitzgerald.2018. . ACM TOSEM27, 3(2018).
[77]
Bent Flyvbjerg.2006. . Qualitative inquiry12, 2(2006), 219–245.
[78]
Victor R Basili, Forrest Shull, and Filippo Lanubile.2002. . IEEE transactions on software engineering25, 4(2002), 456–473.
[79]
Hannah Limerick, David Coyle, and James W Moore.2014. . Frontiers in human neuroscience8(2014), 643.
[80]
Lisanne Bainbridge.1983. . In Analysis, design and evaluation of man–machine systems. Elsevier, 129–135.
[81]
Raja Parasuraman Dietrich H Manzey.2010. . Human factors52, 3(2010), 381–410.
[82]
Rune Haubo Bojesen Christensen.2023. ordinal—Regression Models for Ordinal Data. ://CRAN.R-project.org/package=ordinal.
[83]
Shinichi Nakagawa Holger Schielzeth.2013. . Methods in Ecology and Evolution4, 2(2013), 133–142. https://doi.org/10.1111/j.2041-210x.2012.00261.x.
[84]
Richard D. McKelvey William Zavoina.1975. . Journal of Mathematical Sociology4, 1(1975), 103–120. https://doi.org/10.1080/0022250X.1975.9989847.
[85]
Dwight D. Frink Richard J. Klimoski.1998. . In Research in Personnel and Human Resources Management, Gerald R. Ferris(Ed.). Vol. 16. JAI Press, Greenwich, CT, 1–51.
[86]
Jeffery A. LePine, Nathan P. Podsakoff, and Marcie A. LePine.2005. . Academy of Management Journal48, 5(2005), 764–775. https://doi.org/10.5465/amj.2005.18803921.
[87]
Upol Ehsan, Samir Passi, Koustuv Saha, Todd McNutt, Mark O. Riedl, and Sara Alcorn.2026. . In Proceedings of the 2026 CHI Conference on Human Factors in Computing Systems(CHI ’26). ACM. https://doi.org/10.1145/3772318.3791081.
[88]
George A. Akerlof Rachel E. Kranton.2000. . The Quarterly Journal of Economics115, 3(2000), 715–753.