July 16, 2026
The broad adoption of Artificial Intelligence (AI), especially Generative AI, raises pressing questions about how users interact with these systems to produce new content. In this paper, we introduce the concept of authorship calibration, defined as users’ awareness of their actual authorship when interacting with AI. Using the CoAuthor dataset, we empirically examine how authorship calibration varies across users and how it relates to their frequency of AI use. Our results reveal high variability: users relying heavily on AI tend to misjudge their authorship, whereas those using AI less frequently exhibit more accurate authorship calibration. These findings suggest that AI can obscure users’ perception of their own authorship. In learning contexts, miscalibration can affect metacognitive monitoring and learning strategies, ultimately impacting learning outcomes. Fostering authorship calibration then appears essential for promoting responsible and educationally meaningful AI integration.
The advent of sophisticated Artificial Intelligence (AI) algorithms, such as Large Language Models (LLMs), has significantly impacted our daily habits. Among the various tasks for which such generative AI models can be employed, writing stands out as one of the most prominent, as LLMs can produce human‑like content rapidly and efficiently [1]. AI can then support writers by stimulating creativity, improving grammatical and orthographic accuracy, and enhancing the overall coherence and flow of a text [2]. However, these capabilities also raise fundamental questions regarding ownership and authorship [3], since parts of the output may be generated entirely by the AI without meaningful user contribution. More specifically, depending on how users interact with AI, the system can assume a range of roles, from simple spelling and grammar corrector to functioning as an undeclared ghostwriter [4].
The education field emerges as a particularly significant area of adoption of generative AI [5], with learners now relying on such models to support them in achieving educational tasks. While providing them with permanent, personalized, and on‑demand assistance [6], AI integration into educational contexts also introduces important challenges, especially concerning its effects on metacognitive processes [7] that are essential for effective and autonomous learning [8]. Among metacognitive processes, calibration, or metacognitive calibration, plays an important role: it reflects the degree to which individuals’ judgments about their understanding, capability, competence, or preparedness align with their actual demonstrated performance [9]. Given that AI usage in educational contexts may affect such metacognitive processes, it also can also potentially impact associated learning outcomes.
In this paper, we aim to address both the challenges related to users’ authorship in AI‑assisted writing tasks and the associated impact on metacognitive processes, especially in the educational context. To guide our research work, we formulate the following Research Questions (RQs):
RQ1: How accurately do users evaluate their authorship when engaging in AI‑assisted tasks?
RQ2: Does the frequency of AI use influence authorship perception, and in what ways?
To investigate these questions, and building on the established notion of calibration [10], we introduce the concept of authorship calibration, defined as users’ accurate evaluation of their own contribution during an AI‑assisted task. High authorship calibration indicates that users are aware of how AI is used and integrated into the final output, while accurately evaluating their own authorship. In contrast, miscalibration arises when users either overestimate or underestimate their actual contribution. In an educational context, as with traditional performance‑related miscalibration, poor authorship calibration may hinder learning, potentially resulting in dissatisfaction, overconfidence, or reduced motivation.
We study authorship calibration in the context of writing tasks leveraging the CoAuthor dataset [11], and situates our findings within educational settings to discuss the possible implications for metacognitive processes and learning outcomes. This paper advances the field by (1) introducing the concept of authorship calibration in GAI environments, (2) offering insights into how GAI usage shapes this calibration, and (3) providing new evidence that accurate calibration can benefit educational contexts.
The advent of AI and its spread to the general public has highly transformed our daily lives, and the education sector is highly impacted [5]. Built upon AI algorithms that leverage existing content (text, audio, images, etc.), generative AI systems are able to automatically generate new content [12]. Their integration in educational contexts presents promising opportunities for learners to access permanent, personalized, and on-demand feedback [6], thereby potentially enhancing their overall learning experience [13]. AI is sometimes described as a transformative force capable of empowering learners and revolutionizing pedagogical practices [14]. However, beside these potential benefits, AI in education also raises challenges, especially related to the development of cognitive processes underlying learning [7], [15]. Of particular concern is the risk of reduced cognitive effort and learner engagement when AI is used for educational tasks [16] as learners can potentially become passive participants rather than active actors of their learning process.
To better understand how AI can impact learning, researchers have devoted considerable attention to identifying and describing user-AI interaction patterns. Gomez et al. [17] propose a taxonomy of seven human-AI collaboration patterns, differentiated by the temporal evolution of interactions and by which agent (human or AI) initiates the interaction. Complementing this framework, vanBerkel et al. [18] define three types of human-AI interaction based on the trigger mechanism for AI input: intermittent (explicit user request), continuous (implicit ongoing integration), and proactive (condition-dependent AI initiation). Within educational contexts, Memarian et al. [19] described a multidimensional taxonomy of learner-AI interaction centered on learning alignment and compatibility alignment, referring to the coherence among learning activities, learning experiences, and learning outcomes.
Beyond conceptualizing learner-AI interactions, a growing body of research examines how these interaction patterns impact learning performance. Specifically about AI-support for writing tasks, Nguyen et al. [20] demonstrated that learners who engage more actively in the interaction process achieve higher performances. Similarly, Yang et al. [21] revealed that active engagement in higher-order cognitive processes, such as critical evaluation, synthesis, and revision, is consistently associated with improved essay quality. Addressing this concern about reduced learner engagement, Arnold et al. [22] demonstrated that by modifying the writer-AI interaction design, AI can remain supportive without replacing the writer. However, the role AI plays for writers can vary a lot: from editor supporting writers in the reviewing of their output, to co-author involving collaborative patterns, to ideas source providing inspiration and direction, to ghostwriter generating content that human authors may not transparently declare [4]. This raises fundamental questions about authorship and ownership in AI-assisted writing tasks [23].
The notion of authorship, traditionally referring to the state or fact of being the writer of a document or the creator of a work, becomes highly questionable when AI is integrated into the writing or creation process. As AI models now produce human-like text using powerful natural language models, considering AI as a co-author represents a legitimate conceptual question [24]. This raises complex issues regarding how to appropriately declare AI usage in a work, and what constitutes legitimate authorship in human-AI collaboration. Closely related is the concept of ownership, referring to the legal rights to control and profit from a work, which can be separate from its creator. Considering these related concepts of authorship and ownership, Draxler et al. [3] introduced the "AI Ghostwriter Effect", describing situations where AI users do not attribute authorship to the AI, although they attribute ownership to it. Declaration of authorship by AI users then become skewed, as they avoid acknowledging the AI’s contribution.
These authorship and ownership concepts are particularly interesting in educational settings, where academic tasks may be jointly completed by learners and AI. This makes it highly difficult for teachers to evaluate knowledge acquisition and competence development as traditional assessment methods focus primarily on final products (i.e. product-based assessment) rather than the processes through which they were created (i.e. process-based assessment) [25]. Cheng et al. [26] develop a process-based assessment approach, where the final product is not the unique object being evaluated; instead, the interaction between learners and AI becomes equally essential. Complementing this growing assessment process, technical approaches relying on natural language processing models are able to automatically classify the work as belonging to the learner or the AI [27], [28], then informing about authorship and contribution in AI-assisted works.
Finally, an alternative approach focuses on making interaction patterns and text provenance visible [29]–[31]. In educational contexts, this represents the opportunity to raise learners’ awareness about their engagement in the learning process. For teachers, this gives important insights that can inform evaluation process or adaptations of teaching practices. This awareness-oriented approach connects authorship concerns directly to metacognitive processes, suggesting that learners understanding their own contribution can then adapt their cognitive mechanisms to foster a more effective learning.
In the educational field, metacognition refers to learners’ awareness and regulation of their own cognitive processes, encompassing the planning, monitoring, and evaluating of their learning activities [32]. Metacognition plays a crucial role by enabling learners to become more strategic, reflective, and autonomous [8]. Closely aligned with metacognition, Self-Regulated Learning (SRL) constitutes a well-established theoretical framework describing how learners actively control their own learning processes through a cyclical model involving planning, monitoring, and self-reflection phases [33], [34]. SRL has been extensively studied in educational research, with empirical evidence consistently demonstrating that it serves as a robust predictor of academic success [35]. Recently, Xu et al. [36] specifically examined the relationship between SRL and metacognition in AI environments, revealing that metacognitive support enhances SRL abilities, thereby improving the overall learning experience.
Within the broader metacognitive and SRL frameworks, the concept of calibration, or metacognitive calibration, represents an interesting construct [10]. Calibration reflects the degree of correspondence between individuals’ subjective judgments about their understanding, capability, competence, or preparedness and their objectively measured performance across these dimensions [9]. Well-calibrated learners demonstrate accurate awareness of what they know and what they do not know, enabling effective allocation of study resources and appropriate help-seeking behaviors. Conversely, miscalibrated learners exhibit poor metacognitive awareness and may be either overconfident, overestimating their competencies, or underconfident, underestimating their competencies. Calibration is related to SRL, with calibration accuracy influencing the efficiency of self-regulatory behaviors [37]. Well-calibrated learners can accurately adapt their learning strategies throughout the SRL cycle, thereby enhancing learning outcomes, whereas miscalibrated learners may implement ineffective strategies, resulting in suboptimal performance [10].
In today’s rapidly evolving educational landscape, ensuring that learners remain aware of their own abilities and of the role AI plays in their learning is a major challenge. Yet, to the best of our knowledge, metacognitive calibration remains poorly explored within the context of AI usage in education. In particular, studies exploring how calibration relates to authorship in AI‑assisted work are lacking. We therefore propose to investigate how learner–AI interactions influence what we refer to as authorship calibration, meaning a user’s awareness of his own contribution during an AI‑assisted task.
In this work, we conceptualize calibration through the lens of authorship, introducing the concept of authorship calibration. This authorship calibration reflects the extent to which an individual’s judgment about his contribution in an AI-assisted task aligns with his actual contribution in the final output. Authorship calibration is operationalized using Equation 1 .
\[\text{Authorship calibration} = \text{Declared Authorship} - \text{Actual Authorship} \label{eq:calibration}\tag{1}\]
Declared and actual authorship represent, respectively, the proportion of the final output that the user believes he has authored and the proportion he has actually authored. These proportion are expressed with continuous values ranging in \([0;1]\) (e.g., if the user authored half of the output (\(50\%\)), the corresponding value is \(0.5\)). The resulting authorship calibration score ranges in \([-1; 1]\), with a score of \(0\) indicating a perfect calibration: the user’s perceived contribution matches his actual contribution. Negative scores reflect underestimation, meaning the user contributed more than declared; positive scores reflect overestimation, meaning the user contributed less than declared. The absolute magnitude value of the calibration score (\(|\text{Authorship calibration}|\)) indicates the degree of miscalibration, with larger deviations from zero corresponding to poorer calibration. Figure 1 illustrates how authorship calibration scores vary as a function of declared and actual authorship.
This study relies on the CoAuthor Dataset1 [11], which includes data about 1,445 AI-assisted writing sessions from 60 authors recruited on Amazon Mechanical Turk. The dataset encompasses two writing genres: 830 creative and 615 argumentative writing sessions. Each writing session began with an initial prompt providing an overview of the assigned topic. Users could then choose to write independently (without using GPT) or request AI assistance through five GPT-generated suggestions, which they could ignore, modify, or incorporate without modifications. The dataset captures comprehensive interaction data, including GPT calls and keystroke events recording text insertions and deletions, cursor movements, and interactions with GPT suggestions.
Beyond interaction logs, the dataset provides additional resources including metadata about system configuration, and user survey responses. The metadata includes GPT parameters (temperature and frequency penalty settings) and some quantitative metrics characterizing GPT usage (i.e. number of queries, number of accepted suggestions, proportion of final text written by user, etc.). The survey includes five sections assessing writer demographic and background information, perceived benefits of collaborative writing, user perceptions of GPT’s capabilities and limitations, and overall writing experience. The CoAuthor Dataset is particularly well-suited to our research objectives, as it combines interaction data, essential for analyzing learner-AI collaboration patterns, with survey responses that enable assessment of authorship calibration.
To address our second research question (RQ2), we split users into two groups based on their AI usage. Following the methodology of Shibani et al. [29], we classify users into Low- and High‑AI usage groups based on the median number of AI calls. This metric offers a comprehensive indicator of interaction intensity, capturing the full range of possible engagement with AI (e.g., requesting ideation suggestions, integrating suggestions with or without modification, etc.). Users whose number of AI calls exceeds the median are assigned to the High‑AI usage group, while those below the median are assigned to the Low‑AI usage group. The full code is accessible on our github repository2.
To operationalize our introduced concept of authorship calibration, we leverage survey responses provided in the CoAuthor dataset. Importantly, participants were asked to answer the following question: "—% of the essay/story is written by me (and the rest is written by taking the suggestions)". Users then estimated the percentage (\([0-100\%]\)) of their personal contribution in the final submitted text, as distinct from content originating from GPT suggestions. The actual percentage of text written by the human author is computed by comparing the number of user‑written sentences with the number of sentences originating from GPT suggestions. This percentage is provided in the CoAuthor metadata. We then compute authorship calibration using Equation 1 , which indicates the extent to which users accurately assess their authorship.
From the original 1,445 writing sessions of CoAuthor, we kept only those with both interaction data and complete survey responses. Filtering sessions with missing data, 1,252 sessions remained for analysis (754 creative writing sessions and 508 argumentative writing sessions). For each session, authorship calibration is assessed by comparing declared authorship against actual authorship (See Section 3).
The distribution of authorship calibration scores shows a roughly bell‑shaped pattern centered around zero (mean value \(= -0.003\)), indicating accurate calibration for most writing sessions (See Figure 2). However, the Standard Deviation (SD) is high (\(SD=0.138\)), with scores ranging from \(-0.53\) to \(0.49\), indicating the presence of both strongly under‑ and over‑calibrated users. A limited skewness toward negative values is observed, suggesting a small tendency toward underestimation. Overall, these results highlight considerable variability, motivating further analysis of differences across Low- and High-AI usage groups.
To deepen our analysis, we examine whether and how the frequency of AI use impact authorship calibration. As detailed in Section 4.2, we differentiate between Low-AI usage and High-AI usage comparing the number of calls made to AI (i.e. requests for AI-generated suggestions). In the filtered dataset of 1,251 writing sessions, the mean number of AI calls is \(12.98\) and varies greatly between users (\(SD=9.6\)), ranging from \(0\) calls to \(65\) calls. Relying on the median value of \(11\) AI calls, we classified 647 sessions as High-AI usage and 605 sessions as Low-AI usage.
To compare calibration patterns across user groups, we plot calibration curves (Figure 3), where each point represents a user’s declared authorship positioned against the corresponding actual authorship, and the distribution of corresponding authorship calibration scores is presented in Figure 4. To further analyze how users are distributed within the calibration space, we also provide density heatmaps in Figure 5, offering a more aggregated view of local concentrations and global patterns.
Figure 3: Calibration curves.. a — Low-AI usage, b — High-AI usage
Figure 4: Distribution of Authorship Calibration Scores in Users Groups.. a — Low-AI usage, b — High-AI usage
First, users from the the Low-AI usage group are predominantly concentrated in the upper-right quadrant, with the majority of both declared and real authorship values ranging between \(70\%-100\%\) (See Figure 3 (a)). Low-GPT usage users are relatively evenly distributed on both sides of the ideal calibration line, though a slight underestimation tendency is observable (more user above the ideal calibration line). This is confirmed by the distribution of authorship calibration score that are skewed towards negative values (See Figure 4 (a)), with a mean authorship calibration score of \(-0.004\) (\(SD=0.112\)), ranging from \(-0.53\) to \(0.26\). This negative bias indicates that users with limited AI interaction during writing sessions tend to slightly underestimate their authorship, declaring authorship lower than the actual one. The corresponding heatmap confirms this pattern with the highest density in the \(90\%-100\%\) range for both declared and real authorship (See Figure 3 (a)). This tight clustering around maximal authorship values demonstrates that writers with less frequent AI usage maintain high authorship, while mostly accurately evaluating it.
Second, users from the High-AI usage group are broadly distributed in the calibration space, with both declared and actual authorship spanning from \(10\%\) to almost \(100\%\) (See Figure 3 (b)). Specifically, there is considerable spread below the ideal calibration line (overestimation), particularly visible in the \(30\%-90\%\) declared authorship range. Corresponding authorship calibration values also show greater variability compared to the Low-GAI usage group, with a mean authorship calibration value of \(0.003\), a higher standard deviation (\(SD=0.146\)) and a wider range of \([-0.39,0.49]\) (See Figure 4 (b)). This high variability suggests that users experience an inconsistent awareness of their actual authorship, while having a slight tendency to overestimate it. The heatmap distribution reinforces these observations, with the high densities (\(20 - 50\) users) occurring across the \(50\%-90\%\) range for both declared and real authorship (See Figure 5 (b)). Importantly, a higher density appears below the ideal calibration line, particularly in regions where declared authorship of \(60\%-80\%\) corresponds to real authorship of only \(40\%-70\%\). Users relying more frequently on AI during the writing process then tend to overestimate their authorship. Finally, a statistical comparison further confirms that the distributions of authorship calibration scores differ significantly between Low‑AI and High‑AI users (Mann–Whitney U test, \(p<0.05\)).
Figure 5: Heatmap of calibration curves.. a — Low-AI usage, b — High-AI usage
Our experimental study reveals that while numerous users have accurate authorship calibration scores, an important proportion still show poor calibration, either overestimating or underestimating their authorship. This confirms that authorship calibration is not a trivial construct in AI‑assisted contexts, and that important differences persist across users, as not all of them are able to accurately assess it (RQ1). Besides, our results confirms that the accuracy of authorship calibration vary depending on the frequency of AI use during the writing task (RQ2). Specifically, users in the Low‑AI usage group show a more accurate authorship calibration, but a slight tendency to underestimate it. This may reflect stronger metacognitive awareness, as these users remain attentive to even minimal AI assistance and consequently minimize their own contribution. In contrast, users in the High‑AI usage group show poorer authorship calibration, with a tendency to overestimate it. Extensive AI interactions then appears to blur the distinction between human‑generated and AI‑generated content. This may be the results of an “effort blend” phenomenon: the cognitive work involved in prompting, selecting, and integrating AI suggestions is subjectively experienced as equivalent to generating original text, resulting in a higher sense of authorship. These observations raise important concerns about authorship in AI‑assisted learning environments. Given the conceptual links between authorship calibration and higher‑order metacognitive processes, improving authorship calibration may represent a promising pathway to enhancing learning outcomes in AI‑assisted contexts.
Despite introducing the concept of authorship calibration and offering promising insights, this study has several limitations. First, authorship is examined only within writing tasks rather than in authentic educational settings. Exploring a wider range of educational tasks would allow for richer experiments, ultimately providing a more comprehensive understanding of learner–AI collaboration. Second, the distinction between Low‑ and High‑AI usage groups is based on a simple yet effective methodology. Applying more fine‑grained Learning Analytics pipelines to educational datasets would allow for richer classifications of learner–AI interaction patterns, ultimately providing richer insights into how specific interaction patterns influence authorship calibration.
Finally, this study opens promising directions for future work. A key avenue lies in designing and evaluating methods that support authorship calibration, e.g. through adapted user interfaces or personalized feedback. By increasing learners’ awareness about their actual contribution and fostering more accurate authorship calibration, such interventions may promote more appropriate metacognitive engagement, ultimately improving learning processes and outcomes. In addition, further research is needed to examine how authorship calibration relates to broader metacognitive processes, particularly within AI‑assisted learning environments.
This paper introduces authorship calibration, the extent to which users accurately perceive their own authorship during AI-assisted tasks. Our results show clear variability across users, with authorship calibration being influenced by the level of AI usage. While some users maintain an accurate sense of authorship, others substantially misjudge their contribution, revealing how easily AI can blur boundaries between human- and AI-produced content. Importantly, authorship calibration offers a powerful lens for understanding learning in AI‑assisted environments. Accurate authorship calibration may supports metacognitive monitoring and informed engagement, whereas miscalibration risks unproductive behaviors and weaker learning outcomes. As AI becomes broadly embedded in education settings, helping learners stay aware of their actual authorship appears essential. Designing tools and pedagogical strategies that strengthen authorship calibration will be key to fostering responsible, transparent, and educationally meaningful integration of AI in educational settings.
https://github.com/celinatreuillier/Authorship_Calibration_CoAuthor (will be publicly available upon acceptance).↩︎