June 27, 2026
Despite near-universal adoption of LLMs in undergraduate academic writing, the field lacks validated frameworks for distinguishing qualitatively different forms of AI reliance. Existing instruments treat reliance as a single frequency construct, generating outcome measures that, as this study demonstrates, systematically reward cognitive dependency over intellectual authorship. Grounded in the AI Literacy Framework, Expectancy-Value Theory, and Biggs’s Presage-Process-Product Model, this sequential explanatory mixed-methods study (\(N = 382\) undergraduates; 14 interviews; 396 qualitative survey responses) at a public minority-serving research university identified and validated four empirically distinct reliance types: Strategic (34.3%), Instrumental (30.9%), Dialogic (30.4%), and Dependent (4.5%). Two separate predictive systems govern these types: Expectancy-Value beliefs predict how intensely students rely on LLMs (\(\beta = .630\), \(\Delta R^2 = .256\), \(R^2 = .722\)), while AI literacy predicts which type of reliance they adopt (\(OR = 0.11\), Dependent vs.Strategic, \(p < .001\)), requiring separate, non-interchangeable interventions. The most consequential finding, Strategic users scoring lowest on every outcome measure, was reinterpreted through qualitative evidence as a measurement artifact: existing instruments capture AI-mediated attainment rather than writing quality, penalizing students exercising the greatest cognitive agency. Reflexive thematic analysis (1,435 coded instances; 7 themes) confirmed the typology’s construct validity and introduced a Two-Model Architecture distinguishing students who negotiate AI use through expectancy-value calculus from the 13% who abstain on principled ethical grounds, a population outside the explanatory scope of every existing reliance framework. Findings carry direct implications for AI literacy curriculum design, outcome measure validity, and equitable LLM policy at minority-serving institutions.
In November 2022, the public release of ChatGPT disrupted undergraduate academic writing at a speed no institutional framework could match. To provide context, ChatGPT is a generative artificial intelligence (AI) model founded by OpenAI that produces human-like responses to user prompts. Within three academic years since 2022, survey data suggest that approximately 92% of college students have used large language models (LLMs) for academic work [1], yet the pedagogical, assessment, and policy frameworks governing that use have remained largely unchanged. Research kept pace with adoption statistics, producing a rapidly expanding literature on prevalence rates, academic integrity risks, student perceptions, and the ethics of AI-generated text [2], [3]. What that literature has not produced is a theoretically grounded account of what it means, in qualitatively distinct terms, to rely on LLMs during academic writing.
The limitation is structural. Existing research treats LLM reliance as a single, continuous frequency construct: students are classified as users or non-users, frequent or infrequent, and outcomes are measured with items that assess how much AI contributed to students’ work. This operationalization obscures an educationally consequential distinction. Consider two students submitting the same research paper. The first uses AI to verify a citation, test the coherence of a paragraph she has already drafted, and challenge a claim she is uncertain about. The second opens a prompt window at the start of the assignment and builds the paper’s intellectual architecture from the AI’s output outward. Under every published measurement framework currently available, both students and users are differentiated only, if at all, by frequency counts on a Likert scale. However, their cognitive engagement, developmental trajectories as writers, metacognitive regulation, and academic integrity practices differ categorically. Reducing reliance on frequency cannot make that distinction.
The empirical consequences of this gap are already visible. A growing body of experimental evidence documents significant cognitive costs associated with AI-assisted writing: students who use LLMs exhibit reduced neural connectivity and are unable to recall passages from essays they have just produced, a pattern characterized as accumulating cognitive debt [4]; AI-assisted writers show significantly fewer metacognitive activities than unaided peers, a phenomenon termed metacognitive laziness [5]; and frequent AI use is negatively associated with critical thinking through increased cognitive offloading [6]. These are important findings. They share, however, a common structural limitation: each compares users against non-users without any framework for distinguishing qualitatively different forms of use. Cognitively preserving AI engagement, the kind that maintains metacognitive oversight and authorial control, is analytically indistinguishable from the kind that produces cognitive debt, unless the measurement instrument is designed to make that distinction.
This study provides that distinction. Drawing on the AI Literacy Framework [7], Expectancy-Value Theory [8], and Biggs’s [9] Presage-Process-Product Model, and using a sequential explanatory mixed-methods design with 382 undergraduates at a public minority-serving research university, we identify and validate four empirically distinct LLM reliance types, Strategic, Instrumental, Dialogic, and Dependent, and model the individual-difference predictors that determine which type students adopt and how intensely. The study makes three contributions to the literature. First, it introduces a theoretically grounded, empirically validated typology that distinguishes qualitatively different forms of LLM reliance, filling the conceptual void that frequency-based frameworks cannot. Second, it establishes that two separate predictive systems govern reliance: AI literacy determines type membership; Expectancy-Value beliefs determine intensity, systems that require separate, non-interchangeable interventions. Third, it exposes a fundamental validity problem in existing LLM outcome measurement, the resolution of which is a necessary condition for the field’s interpretive progress.
That third contribution warrants a direct statement at the outset. The most theoretically consequential finding in this study is not the typology itself. It is the discovery that Strategic users, the most deliberate, critically evaluating, and academically principled reliers in the sample, score lowest on every outcome measure. A naïve reading would suggest that principled AI restraint underperforms uncritical dependency. The correct reading is structural: every published outcome instrument in the LLM literature employs items of the form “I use generative AI tools to achieve X,” measuring AI throughput rather than writing quality or cognitive growth. Students who deliberately limit AI’s role in their work score low on these items by design, not because their writing is weak, but because their principled restraint produces low AI-attributed attainment. Qualitative evidence in this study confirmed this reinterpretation with precision. What the field has been calling “outcomes” is, at best, evidence of AI contribution to student work, and, at worst, evidence-based that systematically rewards dependency and misrepresents the most cognitively sophisticated students as the weakest performers.
The institutional context of this study sharpens these stakes. Conducted at a public R1 minority-serving institution (MSI) and Asian American and Native American Pacific Islander-Serving Institution (AANAPISI) in the mid-Atlantic, where over half of undergraduates identify as members of minority groups and roughly a third are first-generation college students, this study examines LLM reliance in precisely the institutional context where its consequences are most acute and most unevenly distributed. At MSIs, where preparation heterogeneity is structural rather than incidental, AI literacy disparities compound existing inequities: students who arrive with stronger prior AI literacy are best positioned to use LLMs in ways that preserve cognitive development. In contrast, those with limited literacy may accept AI output uncritically, producing superficially polished work that obscures rather than develops their underlying capabilities. First-generation status emerged as a significant predictor of reliance intensity in this study, consistent with a compensation dynamic in which AI supplements academic resources that more advantaged peers access through other channels. These equity implications are not peripheral to the study’s findings. They are among its most important.
The paper proceeds as follows. The Literature Review synthesizes research on LLMs in academic writing, writing development theory, equity dimensions of the AI literacy divide, and academic integrity in the post-plagiarism era. The Theoretical Framework section integrates the three explanatory frameworks and specifies the four-type taxonomy. The Method section describes the sequential explanatory mixed-methods design. Results address each research question in turn. The Discussion situates the findings within the broader literature and develops their practical and theoretical implications.
This study integrates three complementary frameworks that operate at two distinct explanatory levels. At the predictive level, the AI Literacy Framework [7] explains which type of reliance students adopt, the qualitative direction of engagement, while Expectancy-Value Theory (EVT; [8]) explains how intensively students rely on LLMs, the quantitative magnitude of engagement. At the structural level, Biggs’ [9] Presage-Process-Product (3P) Model provides the architectural scaffolding connecting student characteristics (presage), reliance behavior (process), and writing outcomes (product) within a single causal system. Together, they constitute a two-tier predictive architecture, separate but interlocking explanatory systems, whose integration is displayed in Figure 1. The central claim this architecture supports is that AI literacy and EVT beliefs are not redundant predictors: a student may hold high EVT beliefs and adopt any of the four reliance types; a student with high AI literacy may still rely intensively but will do so strategically. Different levers govern different dimensions of engagement, and interventions that address only one leave the other untouched.
Long and Magerko’s [7] AI Literacy Framework defines AI literacy as the integrated set of knowledge, critical evaluation capacities, and ethical dispositions required to engage productively with AI systems. [3] identified four dimensions: knowing and understanding AI, using and applying AI, evaluating and creating AI, and navigating AI ethics, each representing a qualitatively deeper level of engagement. [10] extended this framework to educational contexts; their multidimensional account can be distilled, for the purposes of the present study, into three conditions under which LLM engagement is strategic rather than uncritical: sufficient conceptual knowledge of how LLMs function and fail; developed skills for evaluating and verifying AI-generated outputs; and ethical awareness of the consequences of academic AI use. Critically, higher AI literacy does not suppress engagement; it redirects it. [11] found that declining cognitive ability was associated specifically with students who had limited AI literacy; more literate students used the same tools in ways that preserved and extended cognitive engagement. This redirecting function is the mechanism through which AI literacy predicts reliance type: students with higher literacy are more likely to adopt Strategic or Dialogic reliance and less likely to adopt Dependent reliance. In this study, AI literacy is operationalized across two subscales, Core AI Literacy (6 items: knowledge, evaluation, and ethical awareness; \(\alpha = .775\)) and Prior Exposure (7 items capturing breadth of experience; \(\alpha = .664\)), following Allen and Kendeou’s [10] educational extension.
Expectancy-Value Theory [8] proposes that behavioral engagement is governed by two psychological constructs: expectancy for success, the student’s belief that they can perform effectively, and subjective task value, the degree to which the task is perceived as useful, enjoyable, or important. Applied to LLM engagement, EVT predicts that students who believe they can use AI effectively and who assign high utility, attainment, or intrinsic value to AI assistance will rely on LLMs more broadly and persistently than those with weaker expectancy or value beliefs. The theory’s cost dimension, encompassing effort costs, opportunity costs, and psychological costs such as task anxiety, is particularly relevant here: students who perceive the cognitive cost of AI engagement as low will rely more broadly, regardless of their literacy level. Fan et al.’s [5] experimental finding that time pressure was the dominant situational trigger for uncritical AI use is consistent with EVT’s cost-perception mechanism. This differentiates EVT’s predictive role from AI literacy’s: EVT explains how much students rely on AI; AI literacy explains how they do so. EVT is operationalized using five items that capture expectancy for success, efficiency value, intrinsic enjoyment, attainment value, and utility value (\(\alpha = .923\)).
Biggs’s [9] Presage-Process-Product (3P) Model frames learning as a dynamic interaction among presage factors (student characteristics, prior knowledge, motivational beliefs), process behaviors (strategies and approaches), and product outcomes (learning results). The model provides this study’s structural architecture: demographic variables, AI literacy, and EVT beliefs constitute the presage layer; LLM reliance type and intensity constitute the process layer; and self-reported writing outcomes constitute the product layer. Two properties of the 3P Model are especially consequential here. First, its recursive feedback property, products modifying presage characteristics over time, provides the developmental scaffolding for the qualitative finding that students describe reliance trajectories rather than stable reliance states, narrating movement from Dependent toward Strategic reliance as AI literacy and EVT recalibrate through experience. Second, the model explains the measurement artifact finding: when outcome instruments are structured around AI-attributed attainment rather than independent writing quality, they capture the process layer (AI engagement) rather than the product layer (learning), systematically misrepresenting the gap between the two.
From this two-tier architecture and grounded in writing process theory [12], [13] and preliminary pilot data, four reliance types were derived, each defined by a characteristic relationship between the student and AI’s cognitive role in the writing process. The types vary on two theoretically grounded dimensions: the degree of authorial control the student retains over ideational content, and the degree of critical evaluation applied to AI-generated outputs. Table 1 presents the full typology with defining characteristics, cognitive role of AI, and theoretical grounding. The key distinction that prior frequency-based frameworks cannot capture is that Strategic and Dialogic reliance can involve substantial AI engagement while maintaining, or even developing, cognitive agency, whereas Dependent reliance involves the same surface-level engagement while systematically displacing it.
The integration of LLMs into undergraduate academic writing has been rapid, structurally disruptive, and empirically under-theorized in equal measure. Since ChatGPT’s public release in November 2022, LLM tools have diffused from novelty to infrastructure within three academic years, with adoption rates exceeding 92% among college students by 2025 [1] and institutional frameworks still catching up [2]. Research has kept pace with adoption statistics but lagged on the more fundamental question: what does it mean, in qualitatively distinct terms, to rely on LLMs in academic writing? Students employ these tools across all stages of writing with varying purposes and intensities. During pre-writing, LLMs support ideation and the formulation of research questions, reducing the cognitive barriers that lead to writer’s block [14]. During drafting, students request paragraph generation, argument development, and structural scaffolding, with LLMs achieving moderate alignment with human evaluators on organizational coherence but significantly weaker performance on ideational depth [15]. Editing-stage use, grammar correction, and stylistic polishing represent, simultaneously, the least cognitively demanding and the most empirically prevalent applications [2].
Three primary benefits emerge consistently: enhanced efficiency, improved accessibility, and personalized support. [16] found that access to ChatGPT reduced task completion time by approximately 40% and improved average output quality by roughly 18% in a preregistered experiment. For multilingual writers and students with writing-related disabilities, LLMs provide sophisticated language support extending well beyond rule-based error detection, enabling academic-standard writing despite linguistic barriers. However, these advantages introduce a structural concern: homogenization. As many writers come to depend on a small set of shared models, individual gains in fluency may aggregate into a narrowing of expressive diversity, a phenomenon described as algorithmic monoculture [17]. Across three preregistered studies of 2,200 college-admissions essays, each additional human-written essay added roughly two to eight times more to a corpus’s collective semantic diversity than each base GPT-4 essay did, a gap that narrowed but did not close when GPT-4’s output was enhanced through prompt and parameter modifications [18]. A controlled experiment likewise found that composing with an instruction-tuned model significantly reduced content diversity relative to unaided writing [19]. These findings suggest that LLM assistance may shift academic prose toward statistically common rather than rhetorically distinctive expression.
These benefits exist in tension with documented cognitive costs. The field confronts what might be called a performance paradox: AI assistance improves immediate task outcomes while simultaneously producing process-level disengagement that undermines the learning those tasks were designed to produce. Fan et al.’s [5] randomized controlled study (\(N = 117\)) found that AI-assisted writers outperformed other groups on essay scores but showed no advantage on knowledge transfer and demonstrated what the authors term “metacognitive laziness,” significantly fewer planning, monitoring, and self-evaluation activities than unaided learners. [6] found a significant negative correlation (\(r = -.68\)) between AI tool frequency and critical thinking performance, mediated by cognitive offloading, in a mixed-methods study of 666 participants. Kosmyna et al.’s [4] EEG study found that LLM users exhibited up to 55% reduced neural connectivity compared to those writing independently, and that 83% were unable to recall passages from essays they had just written, a pattern the authors characterize as “cognitive debt.” Critically, all three studies compare users to non-users without any theoretical framework to distinguish qualitatively different forms of use. The present study addresses precisely this gap.
Understanding why the type, rather than the frequency, of LLM reliance is educationally consequential requires grounding in foundational writing development theory.
Flower and Hayes’s [12] cognitive process model reconceptualized writing as recursive, goal-directed problem-solving rather than a linear transcription of pre-formed ideas. Planning, translating, and reviewing are not merely steps in a writing process; they are the cognitive operations through which writing produces learning. When students delegate these operations to LLMs, they produce a written product without the learning process the model describes.
Bereiter and Scardamalia’s [13] distinction between knowledge-telling and knowledge-transforming writing identifies the most educationally consequential dimension of LLM reliance. Specifically, knowledge-transforming involves bidirectional interaction between the problem space of content and the problem space of rhetoric: writers use the act of writing to discover, test, and refine what they think. LLM assistance poses categorically different risks depending on the mode in which it is applied. When students use AI for knowledge-telling tasks, organizing understood content, or polishing grammar, the cognitive impact is relatively contained, as knowledge-telling was already producing minimal learning. When applied to knowledge-transforming tasks, generating arguments, synthesizing sources, or developing analytical frameworks, students bypass the recursive problem-solving through which both the written product and the writer’s understanding develop. Dependent reliance, as defined in this study, maps precisely onto this distinction: students who outsource ideation to LLMs may produce polished knowledge-telling texts while systematically underdeveloping the knowledge-transforming capacity that advanced academic work requires.
Graham and Harris’s [20], [21] Self-Regulated Strategy Development (SRSD) framework provides a developmental account of why this matters: expert writing depends on strategically deployed metacognitive skills that must be practiced to become automatic. When LLMs substitute for planning, monitoring, and evaluation, they may prevent the accumulation of practice that helps students develop strategic writing capacity. Fan et al.’s [5] experimental evidence of metacognitive laziness is precisely what this framework predicts when external scaffolding supplants, rather than supports, internal strategy development.
The digital divide has expanded beyond hardware and connectivity to encompass AI literacy as a defining dimension of educational stratification [22]. As conceptualized by [7] and extended by [3] across four dimensions: knowing and understanding AI, using and applying it, evaluating and creating, and navigating AI ethics, and further developed for educational contexts by [10], AI literacy shapes not whether students use LLMs but how that engagement unfolds. Students with higher AI literacy use these tools strategically; those with limited literacy accept AI-generated content uncritically, incorporating errors and biases that undermine academic quality [11]. [23] identified three student profiles differentiated primarily by AI literacy level, with lower-literacy students expressing uncertainty and fear while higher-literacy students used generative AI productively. This AI literacy gap intersects with socioeconomic stratification in ways that compound existing inequities.
[24] describe a critical dynamic: LLMs provide language scaffolding that benefits multilingual writers, yet writing competence is the prerequisite for leveraging that scaffolding productively. Students who already write well use AI to refine and accelerate; students who lack foundational skills often accept AI output uncritically, producing superficially polished texts that obscure rather than develop their underlying capabilities. At an MSI of this kind, where preparation heterogeneity is a defining institutional feature, this paradox has direct force: the present study’s finding that Strategic reliance was predicted by higher prior AI literacy is consistent with this mechanism.
[25] document that 68% of faculty report their institutions have not prepared them to use AI in instruction; a parallel survey of institutional leaders a year earlier found that 81% expected generative AI to widen digital divides [26]. In the absence of structural AI literacy development, students who arrive with the most prior preparation are best positioned to use LLMs productively, deepening the inequities that MSIs are expressly designed to address.
LLM integration has destabilized academic integrity frameworks built on source-matching detection. LLMs generate novel text existing in no retrievable source, rendering conventional detection technically obsolete [27]. Dedicated detectors produce systematically higher false-positive rates among non-native English speakers, raising equity concerns as multilingual students face disproportionate scrutiny [28]. [29], [30] characterizes the present moment as a “postplagiarism era” in which the integrity challenge has fundamentally shifted from detecting copied text to developing students’ capacity to engage ethically with AI. [31] found that students often perceived AI-generated text as superior to their own prose, with delegation driven not only by efficiency but by genuine self-doubt about academic capability, a finding that maps directly onto the self-efficacy dynamics of dependent reliance documented in the present study. The most enduring institutional response is not technical surveillance but the development of principled academic authorship through explicit instruction and AI literacy curricula.
The present study integrates three frameworks that address distinct but complementary dimensions of LLM reliance. The AI Literacy Framework [3], [7], [10] explains why AI literacy predicts which type of reliance students adopt: higher literacy redirects engagement toward critical, evaluative use rather than suppressing it. Expectancy-Value Theory [8] explains why motivational beliefs predict reliance intensity: students who believe they can use AI effectively and who assign high value to AI assistance engage more broadly and persistently. Biggs’s [9] Presage-Process-Product Model provides a structural architecture that maps individual characteristics (presage) onto engagement strategies (process) and outcomes (product), with the critical implication that deep versus surface approaches are contextually variable rather than fixed traits. Their integration produces a two-tier predictive architecture: AI literacy predicts direction; EVT predicts intensity. These are separate systems requiring separate interventions, a claim no prior study had operationalized within a single design.
The reviewed literature leaves three convergent gaps that motivate this study. Conceptually, existing frameworks treat LLM engagement as a frequency construct rather than a typological one, thereby generating outcome measures that capture AI throughput rather than writing quality, and systematically misrepresenting students exercising the most cognitive agency as underperformers. Institutionally, the literature is disproportionately non-American and non-MSI; the equity dynamics of AI literacy and access remain under-examined in contexts where preparation heterogeneity and first-generation enrollment are defining rather than marginal features. Methodologically, large-N surveys establish prevalence without explaining the mechanism, while small qualitative studies illuminate meaning-making without establishing generalizability. No study had used a sequential explanatory mixed-methods design to measure reliance patterns at scale and then purposely investigate the meaning-making and motivational dynamics that explain those patterns in the same population. The present study addresses all three.
This study employed a sequential explanatory mixed-methods design [32] in which a quantitative survey phase (Phase 1) both addressed the primary research questions and informed the purposive sampling of qualitative participants (Phase 2). The design rested on a pragmatist epistemological stance in which research questions, rather than predetermined philosophical commitments, governed methodological choices, and quantitative and qualitative evidence was treated as complementary tools selected for their fit to the problem [33]. A purely quantitative design could establish prevalence and test hypothesized associations but could not explain why students adopted particular reliance patterns or how they reasoned through ethical tensions; a purely qualitative design could surface meaning-making but lacked the statistical power to test the theorized predictor structure in a representative sample. The sequential explanatory architecture resolved both constraints: Phase 1 established patterns at scale, and Phase 2 explained them in depth.
The qualitative phase comprised three complementary data strands selected to triangulate the typology against multiple sources of evidence. The primary strand consisted of 14 semi-structured interviews purposively sampled to maximize variation in reliance type, AI literacy level, first-generation status, and discipline. A second strand drew on open-ended responses from 35 participants who opted into a dedicated qualitative instrument embedded in the Phase 1 survey (hereafter SUR35). A third strand comprised a brief open-ended item completed by 361 of 382 survey respondents (94.5%; hereafter SUR361), which constituted the study’s largest qualitative source and enabled population-level qualitative analysis unavailable from purposive subsampling alone. Together, the three strands produced 1,435 coded instances across 47 codes and seven themes.
Participants were undergraduate students at a public R1 minority-serving institution (MSI) and Asian American and Native American Pacific Islander-Serving Institution (AANAPISI) in the mid-Atlantic, at which over half of undergraduates identify with one or more minoritized racial or ethnic groups [34]. The institution was selected purposively: MSIs concentrate precisely the conditions, preparation heterogeneity, high first-generation enrollment, and broad reliance on federal financial aid, under which differential AI reliance is theorized to carry the greatest equity consequences.
Recruitment proceeded through three channels over a six-week period: (1) email invitations distributed via departmental listservs coordinated through faculty liaisons across more than 30 academic departments representing STEM, social sciences, humanities, and professional programs; (2) in-class announcements delivered by cooperating instructors, supplementing email outreach, particularly in large-enrollment courses; and (3) a university-wide undergraduate email list. To maximize participation rates, 50 randomly selected respondents received $10 in monetary compensation; all 14 interview participants received $25. A staged reminder protocol targeting non-respondents was employed across the collection period. Recruitment materials emphasized voluntary participation, confidentiality, and institutional IRB approval. Following systematic data screening, the removal of one Qualtrics label row, two preview submissions, and one case with more than 20% item non-response, the final analytic sample comprised \(N = 382\) participants.
Sample size was justified through a priori power analysis using G*Power 3.1.9.7 [35]. For hierarchical regression with eight predictors, \(n = 382\) exceeded the minimum of \(n = 236\) required to detect small-to-medium effects (\(f^2 = 0.10\)) at power \(= .95\). For one-way ANOVA comparing four reliance types, \(n = 382\) substantially exceeded the \(n = 280\) needed for power \(= .95\) at medium effect sizes. The moderation analysis target was \(n = 395\) for small interaction effects (\(f^2 = 0.02\)); the achieved \(n = 382\) fell marginally short, meaning non-significant moderation findings should be interpreted as inconclusive rather than as evidence of true null effects.
The LLM Reliance construct is defined operationally as the characteristic manner in which a writer engages large language models during academic writing, distinguished by two dimensions: the locus of cognitive control (the degree to which the student retains authorial agency over ideational content) and the degree of critical evaluation applied to AI-generated outputs before incorporation. This definition separates reliance type, a qualitative orientation, from reliance intensity, the breadth and frequency of engagement, which are governed by distinct predictor systems and require separate measurement. Four dimensions were specified from writing process theory [12], [13] and preliminary pilot data: Strategic, Instrumental, Dialogic, and Dependent. Full conceptual definitions and theoretical grounding for each dimension are provided in the theoretical framework section and Table 1.
Items were generated from two sources: deductive sampling against the construct blueprint, with items anchored to each dimension’s theoretical definition, and inductive refinement through preliminary pilot interview data in which participants described their LLM engagement practices in their own words. Strategic reliance items were adapted from Schraw and Dennison’s [36] metacognitive awareness scales and Long and Magerko’s [7] AI literacy framework; Dependent reliance items were adapted from Bandura’s [37] self-efficacy and learned helplessness scales; Dialogic reliance items drew from Vygotsky’s [38] sociocultural theory and collaborative learning frameworks; and Instrumental reliance items were developed specifically for this study against the blueprint. All items were drafted as first-person behavioral statements rated on a 7-point Likert scale (1 \(=\) Strongly Disagree, 7 \(=\) Strongly Agree) and reviewed for behavioral specificity, temporal specificity, and construct fidelity, that is, each item was written to describe an observable action, anchored to current academic writing practice, and targeted to exactly one dimension of one construct [39].
Prior to pilot administration, item content was reviewed by a small group of colleagues with expertise in educational measurement and AI in education, who provided qualitative feedback on item clarity, construct coverage, and wording precision. Their feedback informed revisions to five items identified in the pilot as producing inconsistent response patterns. Formal computation of item-level content validity indices (I-CVI) was not conducted; this represents a limitation of the current validation sequence, and future refinement of the LLM Reliance Scale should include a structured expert panel using the [40] I-CVI procedure with a minimum of five raters and a retention threshold of I-CVI \(\geq .78\).
The LLM Reliance Scale and its companion measures were developed through a structured process: (1) construct definition and domain specification against a theoretical blueprint, (2) item generation drawing on pilot interviews and existing frameworks, (3) content review, (4) pilot administration with item analysis (\(n = 143\)), and (5) main-study psychometric evaluation. Pilot administration informed three critical refinements: open-ended survey questions were reduced from six to three due to low completion rates (38% vs.% for Likert items); wording was clarified on five items where pilot respondents showed inconsistent response patterns; and interview protocol sequencing was validated through three pilot interviews. Pilot internal consistency was acceptable across all subscales (\(\alpha = .79\)–\(.88\)). Table 1 presents the construct blueprint governing item development; psychometric properties of the final main-study instrument are reported in Table ¿tbl:tab:psychometrics? and the Psychometric Evaluation section.
| Construct | Dimension | Conceptual indicator | Final items | Source |
|---|---|---|---|---|
| LLM Reliance | Strategic | Deliberate, goal-directed engagement; active verification and critical evaluation of AI outputs; full authorial control maintained throughout | 8 | New |
| Instrumental | Bounded use for mechanical tasks (outlining, sentence revision, summarizing) without delegation of ideational content | 4 | New | |
| Dialogic | Iterative co-construction with AI as thinking partner; authorial agency preserved throughout the exchange | 4 | New | |
| Dependent | Cognitive abdication; wholesale outsourcing of writing tasks with minimal critical oversight | 4 | New | |
| AI Literacy | Core knowledge | Understanding of LLM functioning, output evaluation, error detection, citation knowledge, and ethical self-regulation | 6 | Allen & Kendeou (2024) |
| Prior Exposure | Breadth and depth of prior LLM experience across academic contexts | 7 | New |
Note. All items used 7-point Likert scales (1 \(=\) Strongly Disagree, 7 \(=\) Strongly Agree). “New” \(=\) items developed specifically for this study against the construct definitions above. Item counts reflect the final retained instrument after pilot testing (\(n = 143\)). CE \(=\) Critical Evaluation sub-facet; DEL \(=\)
Deliberate Use sub-facet of the Strategic subscale.
The LLM Reliance Scale comprised 20 items across four subscales. The Strategic subscale (8 items) integrated two sub-facets: Deliberate Use (4 items measuring goal-directed, purposeful engagement) and Critical Evaluation (4 items measuring active oversight and verification of AI-generated outputs). The Instrumental subscale (4 items) assessed bounded, task-directed AI use for specific mechanical functions, such as outlining, sentence revision, and summarizing, without delegation of ideational content. The Dialogic subscale (4 items) captured iterative co-construction in which AI functions as a conversational thinking partner while the student retains authorial agency. The Dependent subscale (4 items) measured cognitive abdication: wholesale outsourcing of writing tasks with minimal critical oversight. All items used 7-point Likert scales (1 \(=\) Strongly Disagree, 7 \(=\) Strongly Agree).
AI literacy was measured with 13 items across two subscales. Core AI Literacy (6 items adapted from [10]) measured understanding of LLM processing, output evaluation, error detection, citation knowledge, and ethical self-regulation. Prior Exposure (7 items) assessed the breadth and depth of prior LLM experience across academic contexts. Expectancy-Value beliefs were measured with 5 items adapted from [8], covering expectancy for success, efficiency value, intrinsic enjoyment, attainment value, and utility value. Writing process engagement was assessed with 12 items across four stages (Planning, Drafting, Revising, Editing; 3 items per stage), measuring AI-supported engagement at each stage. Writing outcomes were assessed using 6 single-item indicators (Quality, Self-Efficacy, Clarity, Grammar/Style, Originality, and Critical Thinking) adapted from [41], [42], and [43], which capture perceived AI-mediated writing attainment.
@p4.0cmccccp6.4cm@ Subscale & k & \(\boldsymbol{\alpha}\) & \(\boldsymbol{\omega}\) &
AVE & Representative item
Strategic & 8 & .786 & .789 & .331 & “I verify facts in AI-generated text against authoritative sources.”
Critical Evaluation facet & (above) & — & — & — & “I evaluate the logical consistency of AI output before incorporating it.”
Instrumental & 4 & .877 & .878 & .642 & “I use generative AI tools to rephrase complex sentences in my drafts.”
Dialogic & 4 & .859 & .860 & .607 & “I engage in multiple prompt rounds to co-develop text with AI.”
Dependent & 4 & .879 & .880 & .647 & “I incorporate AI suggestions into drafts with minimal revision.”
AI Literacy—Core & 6 & .775 & .795 & .402a & “I can detect errors and bias in AI-generated content.”
AI Literacy—Exposure & 7 & .664b & .653b & .240a & “I have used generative AI for discipline-specific writing tasks.”
Expectancy-Value & 5 & .923 & .924 & .709 & “I am confident AI tools will enhance my academic performance.”
Planning & 3 & .894 & .895 & .741 & “I use AI to develop detailed outlines before drafting.”
Drafting & 3 & .869 & .875 & .701 & “I use AI to overcome writer’s block more quickly.”
Revising & 3 & .938 & .939 & .837 & “I use AI to revise my draft more thoroughly.”
Editing & 3 & .900 & .901 & .752 & “I use AI to catch grammar and style errors.”
Note. k \(=\) number of items; \(\alpha\) \(=\) Cronbach’s alpha (main study); \(\omega\) \(=\) McDonald’s omega; AVE \(=\) average variance extracted from a single-factor solution. All items used 7-point Likert scales (1 \(=\) Strongly Disagree, 7 \(=\) Strongly Agree). The Critical Evaluation facet and Deliberate Use facet together constitute the 8-item Strategic subscale; both representative items are listed to illustrate the two sub-facets. Writing-process items are shown without the shared preamble “I use generative AI tools to…” for brevity. a AVE below the conventional \(\geq .50\) threshold. b \(\alpha\) and \(\omega\) below the conventional \(\geq .70\)
threshold; acknowledged as a scale limitation.
Subscale scores were computed as the mean of their constituent items, preserving the 1–7 metric so that scores remain interpretable relative to the scale midpoint (4.0) and are comparable across subscales of unequal length. A high subscale score indicates stronger endorsement of that reliance orientation; it does not indicate greater overall frequency of AI use. Reliance intensity, the continuous outcome in RQ2, was operationalized as the mean across all 20 reliance items.
Each participant was assigned a dominant reliance type corresponding to their highest subscale mean, provided that the mean reached a minimum endorsement threshold of 5.0 (above the scale midpoint), indicating a substantively adopted pattern rather than a marginal preference. This classification rule yielded the following: Strategic (\(n = 131\), 34.3%), Instrumental (\(n = 118\), 30.9%), Dialogic (\(n = 116\), 30.4%), and Dependent (\(n = 17\), 4.5%). All 382 participants met the threshold on at least one subscale, indicating that the abstention-mode population identified in the qualitative phase (13.0% of SUR361 respondents reporting principled non-use) is not detectable through scale scores alone; they registered low-intensity strategic profiles quantitatively while articulating categorical refusal qualitatively. This is the basis for the Two-Model Architecture introduced in the Discussion.
EFA was conducted on all 20 reliance items using principal-axis factoring with Promax rotation. Sampling adequacy was meritorious: KMO \(= .926\), Bartlett’s \(\chi^2(190) = 4{,}539.14\), \(p < .001\). Parallel analysis and the scree plot both supported retention of four factors, accounting for 57.2% of common variance (eigenvalues: 8.10, 3.24, 1.25, 1.11). The pattern matrix (Table 2) revealed three cleanly defined factors and one important cross-loading: F1 loaded all four Instrumental items (loadings .69–.84) and all four Dialogic items (loadings .59–.83); F2 loaded all four Dependent items (.69–.86); F3 loaded the four Critical Evaluation items of the Strategic subscale (.55–.88); and F4 loaded the four Deliberate Use items of the Strategic subscale (.45–.67). The Instrumental–Dialogic co-loading on F1 indicates that, at the indicator level, these two constructs share substantial empirical variance. This structural overlap is acknowledged as a scale limitation; the theoretical and qualitative distinction between them, Instrumental AI as a mechanical assistant versus Dialogic AI as a thinking partner, is supported by Theme 2 and Theme 6 of the qualitative analysis but requires indicator-level refinement in future scale development.
| Item | Empirical factor | F1 | F2 | F3 | F4 |
|---|---|---|---|---|---|
| I1 (Instrumental) | F1: Engagement | +.69* | — | — | — |
| I2 | +.84* | — | — | — | |
| I3 | +.83* | — | — | — | |
| I4 | +.79* | — | — | — | |
| DL1 (Dialogic) | F1: Engagement | +.64* | — | — | — |
| DL2 | +.71* | — | — | — | |
| DL3 | +.59* | — | — | — | |
| DL4 | +.83* | — | — | — | |
| DP1 (Dependent) | F2: Dependent | — | +.69* | — | — |
| DP2 | — | +.81* | — | — | |
| DP3 | — | +.86* | — | — | |
| DP4 | — | +.83* | — | — | |
| S5 (CE—Strategic) | F3: Crit. Eval. | — | — | +.82* | — |
| S6 | — | — | +.88* | — | |
| S7 | — | — | +.72* | — | |
| S8 | +.39 | — | +.55* | — | |
| S1 (DEL—Strategic) | F4: Delib. Use | — | — | — | +.58* |
| S2 | — | — | — | +.67* | |
| S3 | — | — | — | +.45* | |
| S4 | — | +.39 | — | +.48* |
Note. F1 \(=\) Engagement factor (Instrumental and Dialogic items co-load); F2 \(=\) Dependent; F3 \(=\) Critical Evaluation facet of Strategic; F4 \(=\) Deliberate Use facet of Strategic. Pattern coefficients below .30 are suppressed and replaced with an em dash. Bold type and an asterisk (*) indicate a primary loading (\(\lambda \geq .40\)). Unbolded coefficients without asterisks are cross-loadings retained for transparency. CE \(=\) Critical Evaluation sub-facet; DEL \(=\)
Deliberate Use sub-facet of the Strategic subscale.
Main-study internal consistency was indexed by both Cronbach’s \(\alpha\) and McDonald’s \(\omega\) (Table ¿tbl:tab:psychometrics?). Acceptable-to-excellent reliability was achieved for Instrumental (\(\alpha = .877\), \(\omega = .878\)), Dependent (\(\alpha = .879\), \(\omega = .880\)), Dialogic (\(\alpha = .859\), \(\omega = .860\)), Expectancy-Value (\(\alpha = .923\), \(\omega = .924\)), and all writing-process subscales (\(\alpha = .869\)–\(.938\)). The Strategic subscale produced lower reliability (\(\alpha = .786\), \(\omega = .789\)), consistent with the EFA finding that its two facets (Deliberate Use and Critical Evaluation) load on separate factors; the 8-item composite indexes a broad orientation rather than a unidimensional latent trait. The AI Literacy, Exposure subscale fell below the conventional threshold (\(\alpha = .664\), \(\omega = .653\)), reflecting heterogeneous item content spanning different types of AI experience; this subscale is used as a predictor of reliance intensity rather than as a stand-alone construct score, and its sub-threshold consistency is acknowledged as a limitation.
Convergent validity was supported by a theoretically predicted positive correlation between Strategic reliance and the AI Literacy composite (\(r = .610\)) and by average variance extracted (AVE) values that were acceptable for three of four reliance subscales (Instrumental \(= .642\), Dialogic \(= .607\), Dependent \(= .647\)). The Strategic subscale’s AVE (\(.331\)) fell below the conventional .50 threshold, attributable to its two-facet structure identified in EFA; facet-level AVEs are acceptable.
Discriminant validity was evaluated through three converging criteria. First, the complete inter-factor correlation matrix (Table 3, lower triangle) showed theoretically expected patterns: the Strategic–Dependent correlation was lowest (\(r = .187\)), consistent with their conceptual opposition as restrained versus wholesale AI engagement, while Instrumental, Dialogic, and Dependent intercorrelations were higher (\(r = .530\)–\(.774\)), reflecting their shared characteristic of high AI engagement volume. Second, the Fornell–Larcker criterion held for all pairs except marginally for Instrumental–Dialogic (\(\sqrt{\text{AVE}}_{\text{Instr}} = .801\) vs. \(r = .774\); \(\sqrt{\text{AVE}}_{\text{Dial}} = .779\) vs.\(r = .774\)). Third, the heterotrait–monotrait ratio (HTMT; [44]) fell below the conservative .85 threshold for all pairs except Instrumental–Dialogic (HTMT \(= .891\)). These results confirm that Dependent, Strategic, and the other pairings are empirically distinct but reveal that Instrumental and Dialogic share a level of empirical variance that challenges their separation as independent constructs at the indicator level, an expected consequence of the EFA’s F1 co-loading and a priority for future scale refinement.
| 1. Strat. | 2. Instr. | 3. Dial. | 4. Dep. | |
|---|---|---|---|---|
| 1. Strategic | .575 | .647 | .679 | .462 |
| 2. Instrumental | .548 | .801 | .891a | .604 |
| 3. Dialogic | .561 | .774 | .779 | .638 |
| 4. Dependent | .187 | .530 | .553 | .804 |
Note. Lower triangle \(=\) Pearson inter-factor correlations. Bold diagonal \(=\) square root of average variance extracted (\(\sqrt{\text{AVE}}\)). Upper triangle \(=\) heterotrait–monotrait ratio of correlations (HTMT; [44]). Discriminant validity is supported when (a) \(\sqrt{\text{AVE}}\) in each row exceeds the off-diagonal correlations in that row (Fornell–Larcker criterion) and (b) all HTMT ratios fall below the conservative threshold of .85. a HTMT \(= .891\)
for the Instrumental–Dialogic pair, exceeding the .85 threshold; discriminant validity for this pair is not supported.
Criterion validity was supported by significant one-way ANOVA effects of reliance type across all ten writing process and outcome variables (all \(p < .001\); \(\eta^2\) range \(= .166\)–\(.329\)), confirming that type membership systematically predicts theoretically expected behavioral differences. Known-groups validity was additionally supported by the multinomial logistic regression finding that AI literacy was the strongest predictor of type membership: each unit increase in AI literacy was associated with approximately a tenfold reduction in the odds of Dependent versus Strategic classification (\(OR = 0.11\), \(p < .001\)), confirming that high-AI-literacy students disproportionately exhibit Strategic reliance.
Common method variance (CMV) was assessed using Harman’s single-factor test, which produced a first unrotated factor accounting for 42.9% of total variance, below the 50% threshold, suggesting that common method variance was not a serious concern. Missing data were screened using Little’s [45] MCAR test, \(\chi^2(16) = 18.43\), \(p = .301\), confirming that data were missing completely at random; the overall missing rate was 0.26%, and listwise deletion was employed for all multivariate analyses.
The equity argument this study advances, that AI reliance patterns carry differential consequences for first-generation, multilingual, and minoritized students at MSIs, presupposes that the LLM Reliance Scale measures the same constructs in the same metric across demographic groups. Measurement invariance across gender, first-generation status, and race/ethnicity was not tested in this study, representing a limitation of the current validation evidence. Future work should implement a configural–metric–scalar invariance sequence using multigroup CFA, with invariance inferred from \(\Delta\)CFI \(\leq .010\) and \(\Delta\)RMSEA \(\leq .015\) [46], [47]. Scalar invariance is a necessary condition for the meaningful comparison of latent means across subgroups. Item-level differential item functioning (DIF) analysis should also be conducted to identify any items whose response patterns differ across groups in ways not explained by group differences on the underlying construct [39]. Until invariance is established, between-group comparisons of reliance type prevalence and intensity should be treated as exploratory rather than confirmatory.
Following best practice for applied educational instruments, the intended use of the LLM Reliance Scale is stated explicitly [39]. The scale is designed to classify undergraduate writers according to qualitatively distinct patterns of AI-supported writing behavior for the purposes of research, AI literacy assessment, and educational intervention design. Appropriate uses. The instrument is appropriate for: (1) research examining the relationship between reliance patterns and learning outcomes; (2) AI literacy program evaluation, in which pre- and post-instruction reliance profiles assess whether instruction shifts students from uncritical to strategic engagement; and (3) educational intervention design, in which student reliance profiles inform differentiated instructional responses. Inappropriate uses. The instrument is expressly not designed for: (1) grading students or evaluating individual academic performance, as subscale scores reflect orientation toward AI engagement rather than writing quality or academic achievement; (2) academic misconduct detection, as Strategic users, those exercising the greatest cognitive restraint, produce the lowest scores by design, rendering low scores an unreliable indicator of problematic use; and (3) evaluating writing quality or cognitive ability, as the measurement artifact finding in this study demonstrates that AI-reliance outcome instruments capture AI throughput rather than independent writing competence. Population boundaries. The scale was developed and validated with undergraduate students at a public R1 minority-serving institution in the United States. Transfer to graduate students, K–12 populations, faculty, or professional contexts should be preceded by revalidation, as AI engagement dynamics and construct relevance may differ meaningfully across these groups. Transfer to non-English-language contexts requires translation validation, including back-translation, bilingual expert review, and independent psychometric evaluation in the target population.
Following Phase 1 data analysis, maximum variation sampling identified 14 interview participants representing diversity across reliance types, AI literacy levels, first-generation status, and discipline. Selected participants received personalized invitations describing the study purpose, estimated duration (approximately 60 minutes), and compensation ($25). Interviews were conducted synchronously via Webex video conference, with video disabled to reduce self-consciousness, or asynchronously via structured written response for students facing scheduling constraints; pilot testing confirmed comparable data richness across modalities. Recordings were transcribed verbatim using Webex automated transcription, followed by manual verification of technical terminology. All participants received transcripts via email with a 72-hour member-checking review period [48]. The interview protocol comprised 13 core questions across four sections: LLM usage contexts, reliance-type self-perception, writing development impact, and ethical reasoning.
Quantitative analyses followed [49], with \(\alpha = .05\) throughout and distributional assumptions (normality, homoscedasticity, multicollinearity via VIF) checked before each model. RQ1 was addressed through one-way ANOVA, Welch’s F, where Levene’s test indicated heterogeneous variances, with Tukey HSD or Games–Howell post-hoc comparisons and Cohen’s d effect sizes, complemented by multiple regression entering reliance intensity and dummy-coded reliance type as predictors of each outcome. RQ2 was addressed through four-block hierarchical multiple regression predicting reliance intensity (demographics \(\rightarrow\) AI Literacy composite \(\rightarrow\) AI Literacy–Exposure \(\rightarrow\) Expectancy-Value), with each block’s incremental contribution tested via \(\Delta R^2\); reliance type membership was modeled through multinomial logistic regression with Strategic as the reference category. RQ3 was addressed using moderated multiple regression [50], with all continuous predictors mean-centered prior to interaction-term computation; significant interactions were probed using simple slopes analysis at \(\pm 1\) SD of each moderator.
The three qualitative strands were analyzed using Braun and Clarke’s [51] reflexive thematic analysis in NVivo 15, proceeding through six recursive phases: familiarization, initial coding, theme development, theme review, theme definition, and reporting. Analysis was primarily inductive but theoretically informed by the three guiding frameworks. Rigor was established through four trustworthiness strategies: (1) researcher reflexivity through reflexive journaling maintained throughout all analytic phases; (2) thick description enabling transferability assessment; (3) member checking with four interview participants who confirmed the representational accuracy of their accounts; and (4) cross-strand corroboration, in which patterns were not reported as themes until confirmed independently across at least two of the three qualitative strands. Abstention-mode participants, principled non-users whose near-zero reliance scores reflect values-based categorical refusal rather than low AI literacy or low exposure, were identified through spontaneous, values-grounded statements of refusal in SUR361, not through low scale scores, which is what permits them to be distinguished analytically from low-intensity strategic restrainers.
Integration of the two phases followed a connecting-and-merging logic [52]: quantitative classification structured the purposive sampling for Phase 2, and qualitative meta-inferences were developed by arraying the seven themes against the quantitative findings they explained, confirmed, or complicated. The measurement-artifact finding, that Strategic users score lowest on every outcome measure, emerged from this integration: the quantitative pattern was identified first, and qualitative evidence then provided the interpretive mechanism (Theme 2, Theme 5) explaining why outcome instruments capture AI throughput rather than writing quality. Researcher positionality, as writing instructors and AI education researchers at the study institution, was monitored through reflexive memoing and treated as an analytic resource informing rather than distorting interpretation.
Dominant reliance type was determined by the highest subscale mean score with a minimum threshold of 5.0 indicating meaningful pattern adoption. The resulting distribution was: Strategic (\(n = 131\), 34.3%), Instrumental (\(n = 118\), 30.9%), Dialogic (\(n = 116\), 30.4%), and Dependent (\(n = 17\), 4.5%). At the sample level, Strategic reliance recorded the highest mean (\(M = 4.33\), \(SD = 1.08\)), reflecting that deliberate, regulated AI engagement was normative in this population, while Dependent reliance was notably low (\(M = 2.64\), \(SD = 1.50\)), indicating that uncritical wholesale outsourcing was not. Among predictors, AI Literacy Composite scored at \(M = 4.48\) (\(SD = 0.91\)) and Expectancy-Value beliefs at \(M = 4.05\) (\(SD = 1.68\)), both near the scale midpoint and exhibiting substantial variance. Table ¿tbl:tab:descriptives? presents full descriptive statistics for all study variables.
@lrrrrrrr@ Variable & N & M & SD & Min & Max & Skew & Kurt
Instrumental & 382 & 3.92 & 1.80 & 1.00 & 7.00 & \(-0.27\) & \(-1.08\)
Strategic & 382 & 4.33 & 1.08 & 1.00 & 7.00 & \(-0.32\) & 0.40
Dependent & 382 & 2.64 & 1.50 & 1.00 & 7.00 & 0.68 & \(-0.58\)
Dialogic & 382 & 4.29 & 1.70 & 1.00 & 7.00 & \(-0.51\) & \(-0.67\)
AI Literacy–Exposure & 382 & 4.03 & 1.00 & 1.00 & 7.00 & 0.05 & \(-0.16\)
AI Literacy–Core & 382 & 4.92 & 1.09 & 1.67 & 7.00 & \(-0.27\) & \(-0.07\)
AI Literacy Composite & 382 & 4.48 & 0.91 & 1.93 & 7.00 & \(-0.02\) & \(-0.23\)
Expectancy-Value & 382 & 4.05 & 1.68 & 1.00 & 7.00 & \(-0.35\) & \(-0.89\)
Planning & 379 & 3.69 & 1.82 & 1.00 & 7.00 & \(-0.15\) & \(-1.16\)
Drafting & 380 & 3.51 & 1.78 & 1.00 & 7.00 & \(-0.03\) & \(-1.15\)
Revising & 382 & 3.96 & 1.97 & 1.00 & 7.00 & \(-0.33\) & \(-1.27\)
Editing & 382 & 4.18 & 1.99 & 1.00 & 7.00 & \(-0.39\) & \(-1.16\)
Quality & 380 & 4.02 & 2.09 & 1.00 & 7.00 & \(-0.27\) & \(-1.38\)
Self-Efficacy & 382 & 3.65 & 2.06 & 1.00 & 7.00 & 0.03 & \(-1.39\)
Clarity & 381 & 3.93 & 2.05 & 1.00 & 7.00 & \(-0.27\) & \(-1.39\)
Grammar/Style & 379 & 4.39 & 2.28 & 1.00 & 7.00 & \(-0.41\) & \(-1.41\)
Originality & 380 & 3.03 & 1.86 & 1.00 & 7.00 & 0.37 & \(-1.20\)
Critical Thinking & 380 & 3.47 & 1.95 & 1.00 & 7.00 & 0.08 & \(-1.27\)
Note. All items measured on 7-point Likert scales (1 \(=\) Strongly Disagree, 7 \(=\) Strongly Agree). Scale midpoint \(= 4.0\)
. Writing outcome variables are single-item measures.
One-way ANOVA produced statistically significant differences across all ten dependent variables (all \(p < .001\)) with large effect sizes throughout (\(\eta^2\) range: .166–.329). The pattern was consistent and theoretically consequential: Strategic users reported the lowest engagement and outcome scores across every variable, while Dependent users reported the highest.
For Writing Process Engagement, Planning showed the most pronounced differentiation (\(F = 55.21\), \(p < .001\), \(\eta^2 = .306\)). Strategic users scored nearly two standard deviations below Instrumental users on Planning (\(M = 2.32\) vs.\(4.62\); \(d = 1.49\), \(p < .001\)), with Dependent users reporting the most extensive AI-assisted planning (\(M = 4.82\)). This pattern replicated across Drafting (\(\eta^2 = .324\)), Revising (\(\eta^2 = .297\)), and Editing (\(\eta^2 = .254\)).
For Writing Outcomes, Quality (\(F = 61.32\), \(p < .001\), \(\eta^2 = .329\)) showed Strategic users at \(M = 2.37\) (\(SD = 1.77\)) and Dependent users at \(M = 5.24\) (\(SD = 1.15\)). Self-Efficacy (\(F = 42.70\), \(p < .001\), \(\eta^2 = .253\)) showed Strategic users at \(M = 2.24\) (\(SD = 1.71\)) and Dependent users at \(M = 5.12\) (\(SD = 1.22\)). Critical Thinking (\(F = 34.61\), \(p < .001\), \(\eta^2 = .216\)) showed Strategic users at \(M = 2.28\) (\(SD = 1.79\)) and Dependent users at \(M = 4.18\) (\(SD = 1.47\)). Originality produced the most theoretically significant pattern, with Strategic users at \(M = 1.85\) (\(SD = 1.44\)), substantially below the scale midpoint, and Dependent users at \(M = 4.35\) (\(SD = 1.32\)). Table ¿tbl:tab:anova? presents all group means, standard deviations, F statistics, and effect sizes.
@lccccrc@ Variable & & & & & F & \(\boldsymbol{\eta^2}\)
Planning & 4.62 (1.46) & 2.32 (1.61) & 4.82 (0.99) & 4.11 (1.55) & 55.21*** & 0.306
Drafting & 4.31 (1.42) & 2.11 (1.39) & 4.47 (1.34) & 4.11 (1.62) & 60.12*** & 0.324
Revising & 4.76 (1.48) & 2.48 (1.78) & 4.84 (1.30) & 4.68 (1.72) & 53.24*** & 0.297
Editing & 4.90 (1.49) & 2.79 (1.98) & 5.06 (1.14) & 4.88 (1.71) & 42.80*** & 0.254
Quality & 4.97 (1.64) & 2.37 (1.77) & 5.24 (1.15) & 4.73 (1.79) & 61.32*** & 0.329
Self-Efficacy & 4.50 (1.77) & 2.24 (1.71) & 5.12 (1.22) & 4.15 (1.95) & 42.70*** & 0.253
Clarity & 4.91 (1.65) & 2.47 (1.87) & 4.62 (1.36) & 4.47 (1.77) & 46.74*** & 0.271
Grammar/Style & 5.00 (1.86) & 3.12 (2.40) & 5.47 (1.50) & 5.05 (2.00) & 24.81*** & 0.166
Originality & 3.99 (1.76) & 1.85 (1.44) & 4.35 (1.32) & 3.21 (1.72) & 40.88*** & 0.246
Critical Thinking & 4.44 (1.61) & 2.28 (1.79) & 4.18 (1.47) & 3.73 (1.82) & 34.61*** & 0.216
Note. All variables measured on 7-point scales (1 \(=\) Strongly Disagree, 7 \(=\) Strongly Agree). Welch’s F reported; degrees of freedom adjusted. \(\eta^2\) interpreted per [53]: \(\geq .14 =\) large. *** \(p < .001\). Group \(n\)
s vary slightly across outcomes due to listwise deletion.
Post-hoc pairwise comparisons confirmed that the Strategic–Dependent contrast produced the largest effect sizes in the study (Planning: \(d = 1.61\); Drafting: \(d = 1.76\); Quality: \(d = 1.73\)), establishing these as the most divergent behavioral profiles. Instrumental, Dependent, and Dialogic users did not significantly differ from one another on most variables, with notable exceptions for Originality (Instrumental \(>\) Dialogic: \(d = 0.45\); \(p = .010\)) and Critical Thinking (Instrumental \(>\) Dialogic: \(d = 0.41\); \(p = .017\)). Multiple regression confirmed that Reliance Intensity was the strongest and most consistent predictor across all outcome models (\(\beta\) range: .598–.811, all \(p < .001\); \(R^2\) range: .403–.737), with the Strategic type dummy independently predicting lower AI-assisted attainment on Planning (\(\beta = -.173\)), Drafting (\(\beta = -.192\)), Quality (\(\beta = -.256\)), Self-Efficacy (\(\beta = -.255\)), and Originality (\(\beta = -.259\)) beyond overall intensity. Figures 2 and 3 display these group means across the four writing-process stages and six writing-outcome measures, respectively.
The pattern of Strategic users scoring lowest on every outcome variable requires direct interpretive intervention. A naive reading would suggest that strategic restraint underperforms relative to other reliance modes, a finding that would be theoretically incoherent given the construct’s definition. The correct interpretation emerges from close examination of item wording. Every outcome item in this study, and in the published literature, employs the structure “I use generative AI tools to [achieve X].” These items capture perceived AI-mediated attainment, not independent writing quality or cognitive growth. Students who deliberately refrain from using AI for most writing functions will score low on these items by design. Strategic users’ suppressed scores reflect their principled restraint, not their writing capacity.
Qualitative evidence confirmed this reinterpretation with precision. Tyler, a Mechanical Engineering junior representing the study’s most pronounced strategic profile, described approaching AI verification as “trying to disprove it almost. So, I’ll go back to the textbook.” His near-zero Originality score captured how little AI contributed to his writing, not how little originality his writing contained. This finding has field-wide implications: any study using analogous items to compare users and non-users, or frequent and infrequent users, is not measuring writing quality but AI throughput. The construct validity crisis this represents is not marginal. It is architectural.
The four-block hierarchical regression produced a well-specified model with \(R^2 = .722\) (adj.\(R^2 = .716\)). Block 1 demographics explained \(R^2 = .061\) (\(p < .001\)), with First-Generation status (\(\beta = .153\), \(p = .008\)) and STEM major (\(\beta = .127\), \(p = .013\)) as the only individually significant predictors. Block 2 added AI Literacy composite (\(\Delta R^2 = .264\), \(F(1, 368) = 143.10\), \(p < .001\); \(\beta = .519\), \(p < .001\)), representing the largest single-block increment. Block 3 added AI Literacy-Exposure (\(\Delta R^2 = .141\), \(F(1, 367) = 96.76\), \(p < .001\); \(\beta = .755\), \(p < .001\)), with its entry reducing AI Literacy composite to non-significance (\(\beta = -.143\), \(p = .141\)), indicating that prior exposure is the more proximal predictor of reliance breadth. Block 4 added Expectancy-Value beliefs, producing the strongest increment (\(\Delta R^2 = .256\), \(F(1, 366) = 335.92\), \(p < .001\); \(\beta = .630\), \(p < .001\)) and the largest unique contribution. In the full model, only AI Literacy-Exposure (\(\beta = .308\), 95% CI [0.189, 0.427]) and Expectancy-Value (\(\beta = .637\), 95% CI [0.568, 0.705]) retained significance. Table ¿tbl:tab:hierarchical? presents the full hierarchical regression model.
@lcccccc@ Predictor & & & & & & VIF
Gender (woman) & .049 & .051 & .032 & .029 & .311 & 1.11
First-generation & .153** & .130** & .102* & .040 & .212 & 1.08
STEM major & .127** & .089* & .111** & .027 & .348 & 1.10
Lower SES & .092 & .099* & .064 & .048 & .126 & 1.11
CGPA & .047 & .010 & \(-.006\) & \(-.013\) & .660 & 1.05
AI Literacy composite & — & .519*** & \(-.143\) & \(-.028\) & .506 & 1.30
AI Literacy–Exposure & — & — & .755*** & .308*** & \(<.001\) & 4.87
Expectancy-Value & — & — & — & .630*** & \(<.001\) & 4.87
\(R^2\) & .061 & .330 & .465 & .722 & &
Adjusted \(R^2\) & .053 & .319 & .455 & .716 & &
\(\Delta R^2\) & .061 & .264 & .145 & .256 & &
F for \(\Delta R^2\) & 4.80*** & 143.10*** & 96.76*** & 335.92*** & &
Note. \(\beta =\) standardized regression coefficient. \(\Delta R^2 =\) increment in \(R^2\) from previous block. Reference categories: Gender \(=\) non-woman; First-Gen \(=\) continuing-generation; STEM \(=\) non-STEM; SES \(=\) middle/upper. * \(p < .05\). ** \(p < .01\). *** \(p < .001\)
.
The multinomial logistic regression model was statistically significant (LR \(\chi^2(21, N = 382) = 164.52\), \(p < .001\); McFadden pseudo-\(R^2 = .175\); AIC \(= 823.44\)). AI Literacy was a significant negative predictor across all three contrasts: Instrumental vs.Strategic (\(OR = 0.37\), \(p = .006\)), Dependent vs.Strategic (\(OR = 0.11\), \(p < .001\)), and Dialogic vs.Strategic (\(OR = 0.25\), \(p < .001\)). The Dependent contrast produced the largest effect: each unit increase in AI literacy was associated with approximately a tenfold reduction in the odds of Dependent versus Strategic classification. AI Literacy-Exposure showed the inverse pattern (ORs: 2.89–6.83), with greater prior exposure positively predicting all non-Strategic types. Expectancy-Value beliefs were strong positive predictors of all three non-Strategic types (ORs: 2.07–2.17), with near-identical odds ratios across contrasts. Demographic predictors were non-significant in all contrasts. Table 4 presents all multinomial logistic regression results.
| Instrumental vs.Strategic | Dependent vs.Strategic | Dialogic vs.Strategic | |||||||
|---|---|---|---|---|---|---|---|---|---|
| 2-4(lr)5-7(lr)8-10 Predictor | OR | 95% CI | p | OR | 95% CI | p | OR | 95% CI | p |
| Gender (woman) | 0.99 | [0.54, 1.84] | .986 | 0.87 | [0.29, 2.59] | .807 | 0.72 | [0.40, 1.32] | .288 |
| First-generation | 2.16* | [1.04, 4.46] | .038 | 2.32 | [0.64, 8.34] | .198 | 1.61 | [0.77, 3.35] | .205 |
| STEM major | 0.92 | [0.45, 1.86] | .812 | 0.65 | [0.20, 2.13] | .477 | 1.14 | [0.56, 2.32] | .713 |
| Lower SES | 1.72 | [0.85, 3.47] | .129 | 0.51 | [0.13, 2.03] | .337 | 1.47 | [0.73, 2.96] | .276 |
| AI Literacy composite | 0.37** | [0.18, 0.75] | .006 | 0.11*** | [0.03, 0.37] | \(<.001\) | 0.25*** | [0.12, 0.50] | \(<.001\) |
| AI Literacy–Exposure | 2.89** | [1.45, 5.73] | .002 | 6.83** | [1.96, 23.83] | .003 | 3.29*** | [1.67, 6.49] | \(<.001\) |
| Expectancy-Value | 2.17*** | [1.70, 2.77] | \(<.001\) | 2.07*** | [1.35, 3.17] | \(<.001\) | 2.11*** | [1.66, 2.68] | \(<.001\) |
Note. OR \(=\) odds ratio. 95% CIs computed via Wald method. Reference category \(=\) Strategic reliance. Dependent group \(n = 17\); interpret with caution. * \(p < .05\). ** \(p < .01\). *** \(p < .001\)
.
The key interpretive insight is that AI literacy and Expectancy-Value beliefs operate at distinct levels: EVT predicts how intensely students rely on LLMs overall, while AI literacy determines which type of reliance they adopt. These are separate predictive systems requiring separate interventions.
Thirteen moderation models were estimated. Six significant interactions emerged. For H3a (AI Literacy as Moderator), two significant interactions were identified: AI Literacy \(\times\) Reliance Intensity predicting Originality (\(\Delta R^2 = .008\), \(F(1, 377) = 6.43\), \(p = .012\), \(\beta = +.092\)), where higher literacy amplified rather than constrained perceived creative benefits, with simple slopes \(b = 1.28\) (\(+1\) SD) and \(b = 1.02\) (\(-1\) SD) (both \(p < .001\)); and Dependent Reliance \(\times\) AI Literacy predicting Self-Efficacy (\(\Delta R^2 = .007\), \(F(1, 378) = 4.94\), \(p = .027\), \(\beta = -.085\)), where higher literacy buffered the self-efficacy inflation associated with dependent use (slopes: \(b = 0.71\) at \(+1\) SD vs.\(b = 0.92\) at \(-1\) SD). For H3b (EVT as Moderator), EVT significantly moderated Reliance Intensity predicting Grammar/Style (\(\Delta R^2 = .014\), \(F(1, 376) = 10.16\), \(p = .002\), \(\beta = -.126\)), consistent with a saturation effect (high-EVT slope: \(b = 0.30\); low-EVT slope: \(b = 0.74\)), and Reliance Intensity predicting Originality (\(\Delta R^2 = .010\), \(F(1, 377) = 7.55\), \(p = .006\), \(\beta = +.103\)). For H3c (Prior Exposure as Moderator), Dependent Reliance \(\times\) Exposure predicting Self-Efficacy was significant (\(\Delta R^2 = .012\), \(F(1, 379) = 8.70\), \(p = .003\), \(\beta = -.135\)), consistent with a developmental recalibration effect. Figures 4 and 5 plot the simple slopes for the significant AI Literacy and Expectancy-Value interactions, respectively. Table 5 presents the complete moderation summary.
| H | Interaction | Outcome | \(\boldsymbol{\Delta R^2}\) | F | p | \(\boldsymbol{\beta}\) | Interpretation |
|---|---|---|---|---|---|---|---|
| H3a | Dep \(\times\) AI Literacy | Self-Efficacy | .007 | 4.94 | .027 | \(-.085\) | Buffering: supported |
| H3a | RI \(\times\) AI Literacy | Originality | .008 | 6.43 | .012 | \(+.092\) | Amplification (unexpected) |
| H3a | RI \(\times\) AI Literacy | Quality | .000 | 0.12 | .725 | — | Not significant |
| H3a | RI \(\times\) AI Literacy | Self-Efficacy | .001 | 0.61 | .437 | — | Not significant |
| H3a | RI \(\times\) AI Literacy | Critical Thinking | .004 | 2.85 | .093 | — | Not significant |
| H3b | RI \(\times\) EVT | Grammar/Style | .014 | 10.16 | .002 | \(-.126\) | Attenuation: supported |
| H3b | RI \(\times\) EVT | Originality | .010 | 7.55 | .006 | \(+.103\) | Amplification (emergent) |
| H3b | Instr \(\times\) EVT | Grammar/Style | .009 | 6.54 | .011 | \(-.103\) | Attenuation replicated |
| H3b | RI \(\times\) EVT | Quality | .004 | 3.66 | .057 | — | Not significant |
| H3b | RI \(\times\) EVT | Critical Thinking | .001 | 0.64 | .424 | — | Not significant |
| H3c | RI \(\times\) Exposure | Self-Efficacy | \(<.001\) | 0.07 | .797 | — | Not supported |
| H3c | RI \(\times\) Exposure | Originality | .010 | 7.67 | .006 | \(+.102\) | Amplification (emergent) |
| H3c | Dep \(\times\) Exposure | Self-Efficacy | .012 | 8.70 | .003 | \(-.135\) | Recalibration effect |
Note. RI \(=\) overall reliance intensity; Dep \(=\) Dependent reliance subscale; Instr \(=\) Instrumental reliance subscale; EVT \(=\) Expectancy-Value Theory composite; Exp \(=\) AI Literacy-Exposure. \(\Delta R^2\) represents the incremental variance attributable to the interaction term. \(\beta =\) standardized regression coefficient. — \(=\)
interaction not significant.
Reflexive thematic analysis across three qualitative strands produced seven themes with 1,435 coded instances. Theme 1 (Functional Utility; 176 instances) established that LLM use is stage-specific and functionally differentiated, with grammar and polishing most demographically distributed. Theme 2 (Reliance Typologies in Practice; 120 instances) confirmed that the quantitative typology captures genuine variation in reliance orientation but revealed a developmental trajectory absent from cross-sectional data: participants narrated movement from dependent to strategic reliance across semesters. Destiny’s account was prototypical: “I realized my grades were going up…but I wasn’t actually learning anything.” Theme 3 (Epistemic Risk and Verification; 93 instances) grounded the AI literacy moderation finding in a near-universal hallucination experience: all 14 interview participants and 22 of 35 SUR35 respondents described encountering fabricated citations, with verification practices emerging through self-directed trial and error rather than formal instruction.
Theme 4 (Motivational Calculus; 136 instances) provided the lived account underlying EVT dominance (\(\Delta R^2 = .256\)), identifying time pressure and task efficiency as the primary motivational drivers of LLM engagement across all three strands (11/14 INT; 24/35 SUR35; 26/361 SUR361). Theme 5 (Ethical Navigation; 140 instances) revealed that professor policy functioned as an externalized moral compass for most students, with 11 of 14 interview participants and 18 of 35 SUR35 respondents reporting instructor AI policy as the primary determinant of their use decisions in any given course.
Theme 7 (Critical AI Ethics, Environmental Concerns, and Principled Non-Use; 188 instances across strands) introduced the study’s most theoretically significant revision. In the full SUR361 strand, 45 respondents (12.5%) spontaneously cited environmental concerns, 47 (13.0%) described principled non-use grounded in personal ethical values, and 12 (3.3%) invoked cognitive-offload concerns. This finding demands a Two-Model Architecture: the four-type reliance taxonomy and its EVT-based predictor model describe negotiation-mode students who balance expected value against cognitive cost in arriving at context-sensitive reliance decisions. Abstention-mode students, principled non-users for whom prior ethical commitments function as categorical constraints, are not represented by any existing reliance framework. Their near-zero scores on reliance scales result not from low AI literacy or weak EVT beliefs but from values-based categorical refusal, a motivational architecture entirely outside the EVT framework’s scope.
This study’s findings cohere around a central argument that extends and substantially corrects the existing literature on LLM reliance: the field has been measuring the wrong thing. By operationalizing LLM outcomes through perceived AI-mediated attainment items, existing studies, including the present study before qualitative reinterpretation, systematically misrepresent their most sophisticated students as their least effective. Strategic users, who exercise the most developed AI literacy, produce the most principled writing practices, and engage AI with genuine metacognitive oversight, score lowest on every outcome measure because those measures capture how much AI contributed to their writing, and their principled answer is: not much. This finding is not marginal. The accumulating evidence base about “outcomes” in LLM reliance research is, at best, evidence about AI throughput and, at worst, evidence that rewards cognitive dependency over intellectual authorship.
The hierarchical regression and multinomial logistic regression analyses revealed two empirically separable predictor systems. Expectancy–Value beliefs were the dominant predictor of reliance intensity (\(\beta = .630\), \(\Delta R^2 = .256\), total \(R^2 = .722\)), accounting for more variance than demographics, prior exposure, and AI literacy combined. AI literacy, by contrast, predicted reliance type, the qualitative character of engagement, rather than its volume (\(OR = 0.11\), Dependent vs.Strategic, \(p < .001\)). These two systems are not redundant: a student with high EVT beliefs may adopt any of the four reliance types, and a student with high AI literacy may still rely intensively, but will do so strategically.
The interpretation is practically direct: AI literacy interventions and motivational interventions are not interchangeable levers. Programs that develop students’ knowledge of LLM capabilities and limitations, their skill in evaluating AI-generated outputs, and their ethical awareness of AI use will shift students from uncritical to critical engagement without necessarily reducing the volume of engagement. Interventions that address knowledge and skills without simultaneously examining the motivational beliefs, perceived utility, efficiency value, and expectancy for success that drive reliance intensity will move one lever while the other remains untouched.
This two-lever structure is consistent with recent empirical work. [54] found that perceived value was far more strongly associated with student intention to use generative AI than perceived cost (\(r = .60\) vs.\(-.30\)) in an EVT-based instrument, indicating that motivational architecture governs the breadth of engagement. [55] found that challenge appraisals positively predicted students’ intention to use generative AI, extending UTAUT’s cognitive predictors with an affective dimension that motivational accounts of reliance would anticipate. Both findings converge on the same implication: AI literacy curriculum design must treat motivational beliefs as a co-equal instructional target alongside knowledge and skills, not as a background condition.
The most theoretically consequential quantitative finding is that Strategic users, the most deliberate, metacognitively regulated, and academically principled reliers in the sample, scored lowest on every one of the ten writing-process engagement and outcome variables (all ANOVA \(p\)s \(< .001\); \(\eta^2\) range \(= .166\)–\(.329\)). The Strategic–Dependent contrast produced the largest effect sizes in the study (Planning: \(d = 1.61\); Quality: \(d = 1.73\)), meaning that the most autonomous writers and the most dependent writers differed by nearly two standard deviations on outcomes that, on their face, should favor the autonomous group.
Qualitative evidence reinterpreted this pattern as a measurement artifact rather than a true performance difference. Every outcome item in this study, and in every published LLM reliance instrument reviewed in its preparation, employs the structural form “I use generative AI tools to achieve X.” These items do not measure writing quality, cognitive growth, or academic achievement; they measure AI throughput, the degree to which students attribute their writing outcomes to AI engagement. Students who deliberately limit AI’s role in their work score low on these items by design, not because their writing is weak, but because their principled restraint produces low AI-attributed attainment. Tyler, the Strategic user whose profile most clearly exemplified this pattern, described his approach as “trying to disprove it almost, so I’ll go back to the textbook.” His near-zero Originality score captured how little AI contributed to his writing, not how little originality his writing contained.
This finding has field-wide implications that extend beyond the present study. Lintner’s [56] systematic review found that only a few of the 16 validated AI literacy scales reviewed had been tested for responsiveness, the ability to detect genuine change in the construct over time, and none had been tested for cross-cultural validity or measurement error, raising a parallel validity concern: the field lacks not only outcome instruments capable of distinguishing AI-assisted from AI-independent performance, but also predictor instruments validated against observable behavioral change. Until both gaps are addressed, the evidence base about what produces good outcomes among LLM users remains systematically confounded by the measurement framework used to generate it.
Reflexive thematic analysis of the SUR361 strand identified a population that the four-type quantitative typology cannot accommodate: 13.0% of open-ended survey respondents described principled non-use grounded in ethical, environmental, or epistemic values, using language that indicated categorical refusal rather than selective restraint. This population produces near-zero scores on all four reliance subscales, scores that, quantitatively, resemble those of low-intensity Strategic users. The qualitative data, however, reveal that the two profiles are motivated by structurally different architectures.
What distinguishes a principled non-user from a highly Strategic user? Three diagnostic dimensions separate them. First, directionality of reasoning: Strategic users engage in an ongoing negotiation, deciding when and how to use AI based on task demands and authorial goals, so their low scores reflect selective, bounded engagement; principled non-users have decided not to use AI, and their near-zero scores reflect categorical exclusion from a decision they consider already settled. Second, context-sensitivity: Strategic users’ AI engagement varies across task types, assignment demands, and instructor policies, whereas principled non-users’ abstention is context-independent, applying the same categorical refusal regardless of task, policy, or time pressure. Third, motivational architecture: Strategic users operate within EVT’s negotiation framework, weighing expected value against perceived cost before each engagement decision, whereas principled non-users have resolved the engagement question before any cost–benefit calculation through prior ethical commitments that function as categorical constraints dominating all other inputs. In short, Strategic users choose not to rely heavily on AI in a given instance; principled non-users have chosen not to use AI at all.
Conflating these two populations in quantitative instruments produces heterogeneity in the predictor structure that regression models cannot disentangle. When abstention-mode students, whose near-zero scores derive from ethical commitments, are pooled with Strategic users, whose restrained scores reflect metacognitive governance, the predictor structure becomes internally inconsistent. The EVT-based negotiation model predicts engagement under varying motivational conditions with substantial explanatory power (total \(R^2 = .722\)), but it lacks a mechanism to predict outcomes determined by categorical prior commitments rather than contextual value calculations. This structural gap motivates the Two-Model Architecture: a negotiation model governing the four-type taxonomy for students for whom engagement is a live behavioral option, and a separate abstention model for students whose prior ethical commitments have already resolved the engagement question categorically. Future scale development must distinguish these populations before incorporating near-zero reliance scores as outcome data in regression models.
The institutional context of this study is not a background condition; it is central to interpreting the findings. First-generation status was the only individually significant demographic predictor of reliance intensity in Block 1 (\(\beta = .153\), \(p = .008\)), a result that carries particular interpretive weight given the institutional setting. At this institution, roughly a third of undergraduates are first-generation college students, more than a quarter identify as Black or African American, and over half receive Pell Grant support [34]. These are not incidental features of the sample; they are defining structural conditions of the MSI context in which AI reliance operates.
The interpretation is consistent with a compensation dynamic rather than a literacy deficit. Roughly a third of MSI students are first-generation, compared with about one in five at non-MSIs [57], meaning that the pattern identified here, first-generation students reporting higher reliance intensity as a proxy for academic support unavailable through other channels, is a structural feature of the MSI context rather than an individual anomaly. First-generation students may not rely more intensively because they have weaker AI literacy; they may do so because AI offers cost-effective access to support that their continuing-generation peers obtain through family cultural capital, writing-center fluency, and prior academic preparation.
The present study situates this compensation mechanism explicitly within an EVT framework: first-generation students’ higher reliance intensity may reflect not weaker AI literacy but stronger perceived utility value in a context where alternative academic support is structurally less available. That unequal access to AI can itself widen inequities in assessment was flagged in early ChatGPT scholarship [2]; the present study extends that concern by identifying which students turn to AI to compensate, and why. The equity implication is pointed. If AI literacy programs succeed in redirecting first-generation students from uncritical to strategic reliance without simultaneously addressing the underlying resource gap that drives their higher intensity, those programs will have changed how students engage without changing the structural conditions that produced the engagement.
Three forces converge at MSIs to produce a specific and correctable form of educational harm: AI literacy disparities compound existing heterogeneity in preparation; institutional policy frameworks respond with detection rather than development; and outcome measures misrepresent the most sophisticated students as the weakest performers. The Reliance Negotiation Framework, the Two-Model Architecture, and the measurement-artifact finding introduced in this study provide the theoretical vocabulary needed to address that harm. Whether institutions develop the will to act on that vocabulary is a question this study cannot answer but compels.
Several limitations bound the interpretive scope of these findings. The cross-sectional design precludes causal inference; the directionality of the EVT-reliance relationship and, particularly, the AI literacy-type membership association, cannot be established without longitudinal data. The developmental trajectories narrated qualitatively, students describing movement from dependent toward strategic reliance across semesters, are theoretically compelling but empirically unconfirmed within the present study’s design.
The study’s reliance on self-report data introduces social desirability concerns that are particularly acute in this domain: the near-universal attribution of dependent use to past selves or other students in interviews suggests systematic underreporting of habituated reliance, a pattern documented in analogous research on academically sensitive self-disclosure. In addition, formal item-level content validity indices were not computed; future studies should use the [40] procedure with at least five expert raters and an I-CVI threshold of .78.
At the same time, the single-institution design limits generalizability; the institution’s specific demographic composition, R1 status, and institutional AI policies may not characterize other MSI contexts, and findings should be interpreted with reference to the particular structural conditions of a public R1 AANAPISI in an urban mid-Atlantic setting. The moderation power analysis indicated \(n \geq 395\) for small interaction effects (\(f^2 = .02\)), and the achieved \(n = 382\) fell marginally short; non-significant moderation findings should be interpreted as inconclusive rather than as evidence of true null effects.
Finally, the measurement-construct gap identified in this study, the inability of existing outcome instruments to distinguish AI-assisted attainment from independent writing quality, is not resolved here; it defines the most urgent instrumentation challenge for the next generation of LLM reliance research, and its resolution will require not only new item structures but new theoretical clarity about what AI outcome measures are intended to capture.
Four implications follow directly from these findings. First, AI literacy curriculum design must be decoupled from reliance intensity reduction as a pedagogical goal. AI literacy predicts type of reliance, not intensity; programs that develop students’ evaluative and ethical competencies will move students from uncritical to critical engagement without necessarily reducing engagement volume. This is the correct pedagogical destination, but it requires explicit acknowledgment that literacy development is not synonymous with use reduction and that interventions calibrated to reduce AI use frequency may inadvertently suppress the very engagement that, when regulated, supports rather than undermines intellectual development.
Second, EVT beliefs must be incorporated as a target in AI literacy curricula, not merely as a background variable. The dominant predictor of reliance intensity is motivational, not competency-based, and interventions that address knowledge and skills without engaging students’ beliefs about AI’s utility and cost, about whether AI use is consistent with who they are and who they want to become as scholars, will fail to shift the most powerful behavioral driver identified in this study.
Third, outcome measurement instruments must be redesigned from the ground up. The field needs instruments that independently assess AI-free writing competence following AI-assisted practice, portfolios of unaided performance, pre-post transfer assessments, and longitudinal writing development measures, rather than perceived AI-mediated attainment captured at a single time point. The systematic review evidence confirms that validated instruments specifically measuring students’ beliefs, perceptions, and practices in AI-assisted academic writing remain underdeveloped across the field frontiers, with most studies relying on adapted instruments not designed for the generative AI context.
Fourth, future research must develop and validate abstention measures that distinguish the grounds for near-zero reliance scores, strategic metacognitive restraint, categorical ethical refusal, low-exposure non-use, and disability-related accommodations, before incorporating these populations into regression predictor models. Pooling abstention-mode students with strategic restraints conflates motivational architectures that are theoretically distinct and practically require different pedagogical responses.
This study established that LLM reliance in undergraduate academic writing is not unitary but structured into four qualitatively distinct types: Strategic, Instrumental, Dialogic, and Dependent, each defined by a characteristic relationship to AI’s affordances, each producing a distinctive profile of writing process engagement, and each requiring a different pedagogical response. The prediction architecture governing these types is two-tiered and empirically separable: EVT beliefs drive the intensity with which students rely on LLMs; AI literacy drives the direction, the type of reliance students adopt. These are not redundant or interchangeable predictors. They are separate systems governing separate dimensions of LLM engagement, and interventions that address only one while leaving the other untouched will produce partial and predictably insufficient results.
The most theoretically consequential finding in this study, that Strategic users score lowest on every outcome measure, is not an empirical discovery about writing quality. It is a measurement problem that, once identified, calls into question the construct validity of an entire literature on LLM outcomes. The outcome instruments the field has been using, structured around perceived AI-mediated attainment, are not measuring what researchers have been claiming they measure. They are measuring AI throughput. Until the field achieves new instrumentation clarity about what AI outcome measures actually assess, its evidence base about the consequences of LLM engagement for student learning will remain systematically distorted in the direction of rewarding dependence and penalizing the principled restraint that genuine intellectual development requires.
The MSI context does not merely add demographic texture to these findings; it specifies their equity stakes. At institutions where a substantial share of enrolled students are first-generation and the majority receive federal financial aid, the structural conditions that drive compensatory AI reliance are not incidental to the research problem. Three forces converge at MSIs to produce a specific and correctable form of educational harm. AI literacy disparities compound existing heterogeneity in preparation. Institutional policy frameworks respond with detection rather than development, and outcome measures misrepresent the most sophisticated students as the weakest performers. The Reliance Negotiation Framework, the Two-Model Architecture, and the Three-Tier Ethical Reasoning Model introduced in this study provide the theoretical vocabulary needed to address that harm. Whether institutions develop the will to act on that vocabulary is a question this study cannot answer but is determined to compel.
The author occasionally used Claude (Anthropic, models Sonnet 4.6 and Opus 4.8) for grammar checking of passages, table formatting, and copy editing of the manuscript. All content was reviewed and verified by the author.