January 09, 2026
Educational innovation is often constrained by the cost, time, and ethical complexity of human-subject studies. Simulated students offer a complementary methodology: learner models that support instructional testing, data generation, teacher training, and social learning [@vanlehn1998applications; @koedinger2015methods]. Large language models (LLMs) substantially expand the expressive power of such simulators, but also introduce a validity risk: fluent interaction can obscure unrealistic error patterns and learning dynamics. We argue that this competence paradox arises when a broadly capable model is asked to behave as a partially knowledgeable learner. We reframe the problem as one of epistemic specification: educational validity depends on explicitly defining what the simulated learner can access at a given moment (knowledge, strategies, representations, and resources), how errors are structured, and whether and how the state evolves with instruction. We formalize LLM-based student simulation as constrained generation under such specifications and propose a goal-by-environment framework that situates simulators by pedagogical goals and deployment contexts. We conclude with open challenges for building valid, reusable, and safe LLM-based student simulators.
Educational innovation is often constrained by the logistical, ethical, and financial challenges of human-subject research. Studying and deploying new educational interventions requires navigating informed-consent obligations, mitigating privacy risks in digital settings, and managing non-random attrition, all while bearing the costs of implementation across diverse sites [@pardo2014ethical; @fixsen2005implementation].
To complement these constraints, researchers have long studied simulated students: computational models of learners that can be instantiated in diverse pedagogical roles. Classic work emphasizes applications including teacher training, learning-by-teaching and collaboration, and formative use in instructional development [@vanlehn1998applications].
While student simulation has a long history within intelligent tutoring systems, recent work has explored the use of Large Language Models (LLMs) as a flexible substrate for simulated students. By supporting open-ended natural-language interaction, LLM-based systems lower the barrier to constructing interactive learner models in less structured domains and interaction settings [@chu-etal-2025-llm]. This flexibility has motivated investigations into richer forms of simulated behavior, including extended interaction over time and participation in multi-agent learning environments [@zhang-etal-2025-simulating].
At the same time, increased fluency and generative capacity introduce new threats to educational validity. Convincing language can obscure unrealistic learning dynamics, compress learner heterogeneity into stereotypes, or introduce artifacts that mislead interpretation [@li2025llm]. A key failure mode is the competence paradox: models trained to be broadly capable are tasked with exhibiting partial knowledge and misconception-driven behavior. We argue that this tension reflects a mismatch between model capability and the learner state the simulator is intended to represent. Because LLMs do not inherently bind generation to an explicit learner state, their errors can resemble superficial deviations from expert reasoning rather than stable, diagnosis-relevant misconceptions. As a result, educational validity depends on explicitly specifying what the simulated learner can and cannot access, and how its behavior is expected to change under instruction.
Student simulation refers to generative systems that produce learner-like behaviors (e.g., answers, questions, errors, or dialogue moves). We distinguish student simulation from student modeling, which focuses on inferring or predicting learner states without enacting interactive behavior. We also note that contemporary LLM-based student simulators are rarely used to test pedagogical theories; accordingly, we focus on LLM-based simulated students in current practice.
To make educational claims comparable and evaluations aligned, we call for an explicit Epistemic State Specification as a required system declaration. This specification states what concepts, strategies, representations, and resources are accessible to the simulated learner at a given moment, what gaps or misconceptions structure its behavior, and how this access evolves over time. We situate this requirement within a context-aware framework that organizes simulated student systems along two dimensions: behavioral goals (performance, learning, and human aspects) and the environment (subject domain, learner population, and interaction modality). Treated as a cross-cutting requirement, the Epistemic State Specification anchors evaluation, prevents overclaiming, and supports meaningful comparison beyond fluency-based proxies.
Our contributions are as follows.
We position simulated students as a complementary methodology for educational research and system development, spanning roles such as teacher training, data generation, social learning and content evaluation.
We analyze the competence paradox as a mismatch between model capability and intended learner-state access, and formalize student simulation as constrained generation under an explicit Epistemic State Specification.
We propose a context-aware framework that categorizes simulated student systems by behavioral goals and environment to support evaluation alignment and comparison.
We identify open challenges for cumulative progress, including validity threats, harm risks, and the need for reusable baselines and standardized evaluations.
Simulated students have long served as a bridge between learning theory and the design and evaluation of educational systems, supporting applications such as teacher training, learning-by-teaching, instructional evaluation, and content authoring [@vanlehn1998applications; @koedinger2015methods].
Early work was dominated by rule-based cognitive models that adopt a white-box view of learning and error, where behavior arises from explicit symbolic mechanisms. Errors were modeled as systematic procedural bugs [@brown1982toward], and later formalized within cognitive architectures such as ACT-R [@anderson1995cognitive; @taatgen2010past]. Systems like SimStudent and the Apprentice Learner extended this paradigm by learning production rules from interaction while retaining explicit representations that support diagnosis and pedagogical intervention [@matsuda2007simstudent; @maclellan2016apprentice]. Although theoretically grounded and interpretable, these approaches require substantial manual structure and scale poorly to open-ended domains.
Recent work leverages large language models (LLMs) as simulated students, exploiting their low authoring cost and flexible dialogue. In learning-by-teaching settings, prior systems constrain LLM agents with explicit or evolving learner states to regulate competence and promote productive interaction, demonstrating that educational usefulness depends on epistemic control rather than fluency alone [@jin2024teach; @rogers2025playing].
A parallel line of work scales from individual learners to classroom-level simulations, using role grounding, planning, and memory to sustain multi-agent educational interactions [@yue2024mathvc; @zhang-etal-2025-simulating]. Related efforts seek to elicit structured reasoning traces and misconceptions from LLM-based learners to support interpretability and formative use [@sonkar2024malalgopy]. Together, these studies highlight both the promise and the risk of LLM-based simulators: while expressive and scalable, their validity hinges on explicit specification of learner state and learning dynamics.
Large language models (LLMs) have enabled a recent shift in educational simulation research from modeling expert tutors toward modeling learners themselves. Rather than focusing on the optimization of instructional policies alone, LLM-based systems increasingly aim to reproduce learner behavior [@QI2026130753], reasoning trajectories [@genstu], and interactional dynamics [@zhou2025socialworldmodels]. This shift is not a superficial role-play change. A simulated student must be intentionally bounded: unlike a general-purpose assistant optimized for helpfulness and correctness, a high-fidelity simulated student must reproduce the characteristic imperfections of human learning, including systematic misconceptions, non-optimal reasoning paths, affect-driven behaviors (e.g., anxiety, disengagement), and learning trajectories that evolve over time.
A core challenge is the competence paradox, which we argue is rooted in an irreducible prior-knowledge entanglement problem. Mainstream LLMs are trained to be broadly capable, self-correcting, and prosocial. Unlike human learners, they cannot genuinely “unknow” solution schemas or expert heuristics once internalized, so even when prompted to act like a novice, their latent reasoning trajectories remain shaped by expert priors that the target student has not acquired.
We therefore define LLM-based student simulation as a constrained generation task whose goal is not to maximize correctness, but to generate responses that remain within an explicit epistemic boundary, namely a specification of what the simulated learner can legitimately access at a given moment (concepts, strategies, representations, and resources), together with the misconceptions or gaps that structure their errors.
Concretely, a simulated student must satisfy three coupled requirements:
Fidelity of Error: given a target misconception (or limitation), the simulator should apply it consistently on the problem types where it is relevant;
Epistemic Consistency: generated mistakes and explanations should be causally attributable to the stated epistemic boundary, and remain stable across isomorphic items and over multi-turn interaction, rather than appearing as one-off surface deviations from expert reasoning; and
Boundary of Competence: outside those regions, the simulator should still behave according to its assumed ability level, avoiding both expert shortcuts that leak inaccessible knowledge and degeneration into random noise.
This framing motivates explicit control mechanisms (prompting, decoding constraints, state variables, external controllers, or structured knowledge representations) that counter the model’s default tendency to self-correct and to converge on globally optimal reasoning, while allowing the epistemic boundary to be dynamically updated as the simulated learner progresses.
To make student simulators comparable and to align evaluation with what is actually being claimed, we organize the design space along two fundamental dimensions: Behavioral Goals and Environment. Behavioral goals specify which aspects of learner behavior a simulator aims to reproduce (e.g., response distributions, learning trajectories, or socio-affective behaviors). Environment specifies where these behaviors are expressed and constrained, including the subject domain, the target learner population, and the interaction modality. The key implication is that two systems can both be called “simulated students” yet be incommensurate: a simulator designed to match error distributions in short-answer math under objective grading is not directly comparable to one designed to model long-horizon learning and help-seeking in open-ended dialogue. Making these two dimensions explicit prevents overclaiming and enables evaluations that test the right notion of fidelity for a given setting.
Behavioral goals specify which facets of learner behavior the simulator must replicate. We distinguish three goals that are often conflated in prior work.
The simulator reproduces a student’s observable outputs at a time point, including success rates and characteristic error patterns. This goal is central for question difficulty estimation, distractor generation, and synthetic data augmentation where realistic response distributions matter more than modeling how the student acquired or failed to acquire the underlying knowledge.
The simulator models the process of learning as a trajectory. Unlike static performance snapshots, learning simulation requires state evolution across interactions, including skill acquisition, forgetting, and differential sensitivity to interventions (e.g., scaffolding, feedback timing, peer effects). This goal is essential for teacher training (observing pedagogical consequences) and for training adaptive tutors, where policies depend on how the student changes rather than what the student answers once.
The simulator reproduces non-cognitive attributes that shape learning interactions, such as personality, motivation, emotion, and socio-linguistic style. These factors determine whether classroom and peer-learning interactions feel behaviorally plausible. Affective realism also requires modeling help-seeking behaviors (instrumental, executive, avoidant) that vary with confidence, anxiety, and perceived social risk, rather than defaulting to uniformly cooperative “good student” behavior.
Environment specifies external constraints that bound what realism means in practice, and what inputs and modalities are available.
Different disciplines require different knowledge representations and error models. Structured domains (e.g., math, programming) afford objective correctness and well-defined misconception patterns; open-ended domains (e.g., writing, discussion) require modeling subjective reasoning, argumentation strategies, and rubric-dependent evaluation.
A simulator should match the target population (age, proficiency, language background, cultural context, neurodiversity) because these factors shape both misconceptions and interaction style. Persona is not only a prompt attribute. It should function as a consistency constraint across turns and tasks, preventing implausible ability drift (e.g., claiming low attention while exhibiting sustained high-focus reasoning).
Interaction setting constrains observable behavior. Text-only dialogue differs from classrooms or VR settings where speech, gaze, and embodied actions matter. Even within text, the availability of artifacts (equations, diagrams, code editors) changes what errors look like. Environment modeling should also account for instructional structure, such as scaffolding and fading: a realistic student may rely on prompts and deteriorate when supports are removed, rather than instantly generalizing.
In addition to the two preceding dimensions, we require each simulated-student system to explicitly declare its epistemic state specification (ESS). ESS defines what the learner knows and can access at a given moment, how errors are generated, and whether and how that state changes over time. Concretely, an ESS specifies (i) the representations, knowledge elements, strategies, and resources available to the learner at time \(t\); (ii) the sources of systematic error, such as misconceptions or incomplete procedures; and (iii) the update mechanism, if any, governing transitions between states.
ESS is intended to clarify the degree to which a system simulates performance at a fixed competence level versus learning as a stateful process. We operationalize ESS as a lightweight reporting label with five levels:
E0: Unspecified. The learner’s internal knowledge, error sources, and state transitions are not defined. Outputs are generated without an explicit epistemic constraint.
E1: Static bounded. The learner operates with a fixed, pre-specified set of knowledge elements, skills, or error templates that do not change during interaction. Behavior reflects performance given an initial competence level, without learning.
E2: Curriculum-indexed. The learner’s accessible knowledge or error patterns are updated according to an external progression signal, such as curriculum position, problem index, or mastery variable, without an explicit model of misconceptions or strategy change.
E3: Misconception-structured. The learner is governed by an explicit, stable model of misconceptions, strategies, or partial procedures that causally determine behavior. Errors arise from identifiable epistemic structures rather than surface randomness.
E4: Calibrated or learned. The learner’s state representation and transition dynamics are learned from or calibrated against human interaction data, such that both performance and state evolution are empirically grounded in observed learner trajectories.
This label functions as a cross-cutting declaration intended to prevent overclaiming, enable meaningful comparison across systems, and align evaluation protocols with the simulator’s stated epistemic constraints. Illustrative examples of several ESS level are provided in Appendix 9.
This section outlines where LLM-based simulated students are most promising and why. We organize the opportunities into four directions: Teacher Training, Social Learning (including Learning-by-Teaching and Collaborative Learning), Data Generation, and Content Evaluation. For each direction, we highlight the educational motivation and distinctive capabilities that LLMs bring. Table 1 summarizes these four directions alongside three unifying benefits that LLM-based simulated students provide across all of them: Scalability (deployment beyond human availability), Safety (risk-free experimentation that reduces privacy, liability, and ethical concerns), and Versatility (behavioral modeling of diverse learner characteristics that are difficult or impossible to reproduce reliably with real students).
| Applied Context | Associated Behavioral Goals | Environmental Considerations |
|---|---|---|
| Teacher Training | High fidelity in Simulating Learning (trajectories) and Human Aspects (affect, personality). | Requires high student group heterogeneity and classroom-like interaction settings, potentially multimodal. |
| Social Learning | Focus on Human Aspects (socio-linguistics) and Simulating Learning (to induce protégé effect). | Dialogue-based interfaces; collaborative or 1-on-1 peer environments. |
| Data Generation | Emphasis on Simulating Performance (error patterns) and Learning (longitudinal logs). | Often structured subjects (STEM, programming); digital interaction logs and standardized tasks. |
| Content Evaluation | Primary focus on Simulating Performance (static success/failure distributions). | High-throughput platform integration; objective correctness criteria and scalable test suites. |
High-quality instruction is a primary driver of student achievement, whereas inadequate teacher preparation leads to detrimental short- and long-term academic outcomes [@markel-etal-2023-gpteach]. However, providing rigorous training has become increasingly challenging due to the rapid expansion of online education and the reliance on large cohorts of novice tutors and teaching assistants [@pan2025tutorup; @markel-etal-2023-gpteach]. Opportunities for actual teaching are often scarce and constrained by factors outside the training program’s control [@zheng2023educasim]. Other traditional training solutions, ranging from webinars to human-authored visual or VR simulations, are resource-intensive and time-consuming to maintain. Consequently, these methods are difficult to deploy at the scale required to standardize training for the growing workforce of pre-service and part-time educators [@lee2023generativeagent; @christensen2011simschool].
Beyond scalability, traditional field experiences are fraught with inherent structural limitations. Training in real classrooms introduces significant accountability, liability, and privacy risks [@christensen2011simschool]. Crucially, limited contact hours restrict the exposure of a novice to student heterogeneity; real-world placements cannot guaranty the variety of personalities, learning styles, and aptitudes necessary to master adaptive instruction [@sanyal2025pedagogical; @knezek2024specialneeds]. This lack of diverse exposure hinders the development of complex skills, such as managing whole-class dynamics or facilitating discussion-based on argumentation [@nazaretsky2023ai].
LLM-based simulated students offer a solution to these challenges by providing frequent, structured, and risk-free rehearsal opportunities [@judge2013visual]. Unlike static environments, LLMs can model diverse, personality-aligned behaviors, creating a rich testbed for practicing adaptive strategies [@ma2025soei; @pentangelo2025senemai]. Research indicates that these simulated environments alleviate the immediate time pressure of real classrooms, allowing teachers to draft more thoughtful, inclusive responses and strategize around learning goals [@markel-etal-2023-gpteach]. By embedding these agents in realistic settings, systems like TutorUp enable novices to practice navigating disengagement and confusion in a safe, controllable environment, bridging the gap between theory and practice [@pan2025tutorup; @lee2023generativeagent].
Simulated students can further facilitate education by functioning as social learning partners. Learning by Teaching (LbT) is a robust pedagogical model where students achieve deeper conceptual understanding through the acts of structuring knowledge for explanation, taking responsibility for a tutee, and reflecting on their own learning processes [@biswas2005learning]. Despite its efficacy across age groups and domains, designing effective LbT environments presents challenges. A critical limitation occurs when neither the tutor nor the tutee is an expert; without guidance, participants often reinforce shared misconceptions or focus on trivial details rather than core concepts [@Debbané:2023:LBT:CSCW]. While human experts can mitigate this by feigning ignorance to elicit explanations, such role-playing is labor-intensive and inherently unscalable [@duran2017learning].
LLM-based simulated students offer a scalable alternative to this dynamic. These agents can simulate a novice learner to induce the protégé effect, in which students make greater effort to learn material they expect to teach, while simultaneously leveraging their underlying expert knowledge to subtly guide the student toward correct understanding [@rogers2025playing]. Furthermore, teaching an artificial agent creates a psychologically safe environment that mitigates the fear of judgment and the social pressure often associated with real-time human interaction [@Chase:2009:TAP:Journal; @Debbané:2023:LBT:CSCW].
Beyond the tutor–tutee dynamic, simulated students can also serve as peers within collaborative learning frameworks. Positive peer interactions are known to enhance academic outcomes and foster essential soft skills, such as communication and group coordination [@Veldman:2020:YCW:LI; @Rohrbeck:2003:PAL:JEP]. Conversely, educational settings that lack real-time feedback and interaction frequently yield suboptimal learning outcomes. In the current digital landscape, which is characterized by pre-recorded lectures, remote instruction, and asynchronous schedules, students often face logistical and social barriers to finding study partners [@Wang:2025:GCL:GROUP]. Simulated students bridge this gap by serving as always-available virtual peers or learning companions that restore the benefits of cooperative interaction to otherwise isolated learning environments.
The efficacy of modern intelligent education platforms (e.g., Coursera) and LLM-based educational tools (e.g., Google LearnLM) relies heavily on the availability of high-quality training data to support personalization and algorithm optimization [@Gao:2025:A4E:arXiv; @Song2024LearnLM]. However, the acquisition of high-fidelity educational datasets, such as student responses and conversation logs, is severely constrained by strict privacy regulations and the high costs associated with manual annotation. Furthermore, real-world educational data is often sparse or noisy, while achieving strong algorithmic performance requires comprehensive and well-structured datasets [@Chen:2025:SIM:arXiv; @Gao:2025:A4E:arXiv].
Simulated students address this scarcity by serving as scalable engines for synthetic data generation [@tutorial]. By modeling diverse learner profiles, these agents can produce large-scale, privacy-compliant datasets that mimic real-world distributions. Recent literature demonstrates the utility of this approach in two key areas: generating realistic static student responses [@Miroyan:2025:PGS:arXiv; @Benedetto:2024:ULMS; @Wu:2025:EISS; @Ross:2025:LMM], and synthesizing complex, longitudinal task-solving trajectories [@Ross:2025:MSL; @Sharma:2024:DSS:EDM].
The rapid expansion of educational technology platforms and pedagogical conversational agents (PCAs), such as Khan Academy’s Khanmigo, Squirrel AI, and Duolingo, presents both opportunities for personalized instruction and challenges in ensuring content quality and pedagogical efficacy [@Gonnermann-Muller:2025:FCE:arXiv; @Jin:2024:TTR:TR]. Evaluating these systems faces critical bottlenecks. Expert teacher assessment is prohibitively expensive, while pilot testing with real students introduces privacy risks and the potential exposure to harmful content [@Ezzaki:2024:EEC:LLM; @Jin:2024:TTR:TR]. Simulated students offer a compelling alternative by providing a high-throughput and risk-free environment for both system-level validation and content-level assessment, particularly for item difficulty modeling [@hambleton1991fundamentals; @hsu2018automated; @alkhuzaey2021systematic; @li2025item; @peters2025text]. Although accurate item difficulty modeling is essential for assessment validity and adaptive learning algorithms [@Duenas:2024:UPN:BEA], traditional approaches, including manual expert labeling and item response theory, suffer from subjectivity, limited scalability, or reliance on extensive historical student data. These limitations make simulated students an increasingly attractive solution.
Simulated students address these limitations by acting as synthetic test-takers. By modeling learners with varying ability levels, these agents can generate performance data to predict item difficulty without the need for large-scale human pilots. Recent frameworks such as UPN-ICC, QG-SMS, and SMART demonstrate that simulated students, whether used independently or to augment IRT models, can achieve high accuracy in calibrating educational content [@Duenas:2024:UPN:BEA; @Nguyen:2025:QGS:arXiv; @Scarlatos:2025:SMA:arXiv]. At the same time, recent large-scale evidence suggests this setting is not solved by simply substituting an LLM for human pilots: even with difficulty labels grounded in real student field testing, off-the-shelf LLMs can be systematically misaligned with human difficulty judgments, and they struggle to reliably simulate lower-proficiency cognitive states [@Li2025DifficultyAlignment]. This underscores the need for calibrated, goal-conditioned simulated students whose proficiency controls and validation protocols are explicitly specified.
High-fidelity student simulation depends on granular, real-world traces that capture learner errors, feedback, and instructional context. In practice, such data are scarce: collecting them is costly (often requiring expert annotation) and sustained access to classrooms or proprietary platforms is difficult to obtain. Missing instructional materials and intervention context can also create a persistent contextual void, limiting a simulator’s ability to represent how pedagogy shapes learning outcomes [@Zhai:2020:AIReview; @Baker:2014:EDM; @Xu:2024:EGSA].
Privacy constraints further restrict both what can be collected and what can be shared. Educational traces frequently contain rich demographic and behavioral identifiers, so many datasets cannot be redistributed, undermining reproducibility and slowing cumulative progress [@Hoel:2018:Privacy; @Wu:2025:EISS]. For LLM-based simulation, the risk compounds because models trained on educational traces may memorize and reproduce private information, or exhibit privacy-related biases that distort pedagogical interactions [@shvartzshnaider2025privacybiaslanguagemodels].
As simulated students are increasingly deployed in long-form inquiry, open-ended reasoning, and authentic classroom dialogue, evaluation becomes difficult because the target behavior is rarely verifiable by binary correctness. Recent evidence suggests a broader non-verifiability crisis: standard automated metrics and even LLM-based judges can misalign with expert human preferences, with reported preference-matching accuracy on the order of \(\sim\)65% in some scientific reasoning settings [@Cohan2025SciArena]. This limitation is amplified by a style-substance mismatch, where models produce responses that are polished and confident yet ungrounded or logically inconsistent. Consequently, automated evaluators may fail to distinguish a student who is productively struggling from a simulator that is convincingly hallucinating [@Zhang2025SimulRAG].
These measurement challenges directly motivate the need for standardized evaluation frameworks that are explicitly aligned with a simulator’s intended behavioral goal and deployment environment. In particular, assessing learning dynamics requires different instruments than those used for static correctness. Yet the absence of shared benchmarks and reporting conventions has pushed current work toward ad-hoc expert ratings that are costly to scale, difficult to reproduce, and hard to compare across studies. This ambiguity makes cross-paper comparisons fragile, since the appropriate notion of fidelity depends on the simulator’s behavioral goal and deployment environment, and remains under-specified without shared, goal-conditioned benchmarks.
Epistemic State Specification (ESS; E0–E4) should be treated as a required reporting artifact for simulated student systems because the competence paradox is ultimately an epistemic mismatch: LLMs cannot truly “unknow,” so they may leak expert knowledge while producing novice-like errors. System descriptions should therefore explicitly state what knowledge, strategies, and resources are accessible at each turn and how this access evolves over time; moving from unspecified (E0) to misconception-structured (E3) or calibrated (E4) specifications makes claims falsifiable (for example, stable misconception behavior under paraphrase and across isomorphic items), turns “student level” into an auditable design choice, and enables meaningful cross-paper comparison.
Evaluation should prioritize goal-aligned fidelity rather than surface realism. The fidelity-evaluation gap arises when fluent dialogue and generic “humanness” scores are taken as evidence of educational validity, even though they can mask implausible learning dynamics. Moreover, simulators built for different purposes (for example data generation versus teacher training) require fundamentally different success criteria.
Metrics should be derived from the simulator’s behavioral goal and environment. Performance and data-generation settings call for controlled error distributions and resistance to drifting into expert reasoning. Learning-oriented settings require trajectory metrics such as gradual improvement and appropriate sensitivity to feedback. Human-aspects and social learning settings require interactional metrics that match the application, such as persistence, help-seeking, or affective coherence when relevant.
Longitudinal simulation benefits from explicit learning mechanisms rather than context-window prompting alone. Prompt-only approaches often yield unrealistic dynamics, including abrupt novice-to-expert jumps, unstable competence across turns, and brittle dependence on phrasing, which undermines the validity of using simulators to study learning or stress-test instructional interventions.
A practical direction is hybrid architectures that pair an LLM with an explicit learner-state representation and a defined transition rule, such as knowledge tracing, proficiency variables, misconception graphs, or cognitive models. This makes learning trajectories interpretable and calibratable, and it makes failures diagnosable as state, transition, or interface errors rather than opaque model variance.
Shared benchmarks are needed for non-verifiable educational behaviors, especially misconception consistency and longitudinal socio-affective trajectories. Many properties central to educational validity cannot be assessed by correctness alone, and automated judges can be misled by fluent but incoherent behavior, leaving the field reliant on expensive, ad hoc expert ratings and fragile comparisons.
Benchmark suites should emphasize consistency under controlled variation. Misconception tests can use isomorphic items and paraphrases to assess stability of error signatures aligned with declared epistemic states. Learning tests can use multi-turn curricula to assess gradual, selective revision under feedback. Socio-affective tests can use scenarios that elicit frustration or re-engagement to assess coherent trajectories over time. Open, standardized suites would directly improve reproducibility and comparability.
In conclusion, to advance LLM-based student simulation from exploratory prototypes to rigorous scientific instruments, the field must resolve the “competence paradox” by prioritizing epistemic fidelity over surface realism. We formalized this shift through the Goal-by-Environment framework and the Epistemic State Specification (ESS), which ensure that simulated behaviors are causally attributable to defined learning states rather than stochastic hallucinations. By adopting these standards, researchers can transform simulated students into trustworthy, reproducible testbeds for educational innovation, bridging the gap between conversational fluency and valid pedagogical modeling.
This study is primarily theoretical and synthesizes insights from previous literature in educational psychology, learning sciences, and natural language processing. We did not conduct empirical evaluations, classroom deployments, or large-scale psychometric validations of the proposed Epistemic State Specification. As such, our claims regarding the resolution of the “competence paradox” are not validated through direct pedagogical interaction or longitudinal system testing. Although we propose a distinct taxonomy for epistemic states (E0–E4), real-world model behaviors often exhibit fluctuating or indeterminate knowledge boundaries. Our framework idealizes these dimensions for analytical clarity, which may limit its robustness when applied to stochastic, non-deterministic foundation models.
Moreover, the proposed solution assumes that developers can enforce strict knowledge constraints via architectural design, which may be undermined by the inherent fragility of prompt engineering and the opacity of commercial black-box models. Future work should focus on validating the framework through controlled empirical studies that measure whether ESS-compliant simulators effectively prevent knowledge leakage and by developing automated metrics to quantify epistemic fidelity across diverse learning scenarios.
The deployment of LLM-based student simulators entails significant pedagogical risks if the generated behaviors are not rigorously grounded in well-defined epistemic states. A primary concern is the potential for negative training transfer in teacher education, in which novice instructors may develop maladaptive teaching strategies by practicing with simulators that exhibit realistic fluency but unrealistic learning dynamics. For example, if a simulator responds to a teaching intervention with plausible yet causally disconnected improvements, such as appearing to master a complex concept immediately after a simple hint, it reinforces superficial instructional moves rather than deep pedagogical reasoning [@dai2025embracingcontradictiontheoreticalinconsistency]. This misalignment creates a hazard in which the simulator functions as a “pedagogical placebo,” offering the illusion of effective practice while failing to reflect the cognitive resistance and gradual skill acquisition observed in real classrooms.
We thank Prof.Ken Koedinger and Prof.Carolyn Rosé for their valuable feedback and suggestions which we heavily incoporated into this paper. We also thank Gati Aher and Aylin ÖZTÜRK for providing helpful suggestions during the early-stage ideation of this project.
Furthermore, LLM-based student profiles can establish normative or cultural biases if the behavioral and linguistic cues associated with “struggle” or “misconception” are not localized or participatory in design [@xiao-etal-2025-humanizing]. This marginalizes underrepresented cultures [@alkhamissi2025hireanthropologistrethinkingculture] or reinforces dominant stereotypes regarding academic ability. Although our Goal-by-Environment framework advocates for context-sensitive profile calibration, more empirical research is needed to verify the effectiveness of such strategies across diverse educational settings.
We also recognize the potential for dual use of our framework. The taxonomy of “human aspects” could be used to inform more persuasive or emotionally manipulative systems, especially in commercial tutoring or surveillance contexts. We encourage future work to develop mitigation strategies, such as interpretability indicators (e.g., exposing the active ESS state to the user), constrained anthropomorphic profiles, or gated release mechanisms, to help monitor and control simulated student behavior. We also stress the importance of interdisciplinary collaboration with learning scientists, ethicists, and affected communities during system development.
By articulating both the functional benefits and the possible harms of epistemic simulation in LLMs, our goal is to support transparent, socially aligned, and user-aware design practices. We strongly encourage future research to empirically validate and refine this framework, particularly through participatory codesign and cross-cultural evaluation.
To identify relevant literature on LLM-based student agents, we conducted a comprehensive search across six major academic repositories: Semantic Scholar, LearnTechLib, arXiv, the ACM Digital Library, SpringerLink, and IEEE Xplore. Following the methodological approach of [@kaser-alexandron-2024-simulated], we employed a Boolean search string targeting variations of simulation terminology:
“simulated students” OR “student simulation” OR “simulated learners” OR simstudent OR simstudents OR “simulated student” OR “simulated learner”
The initial search yielded a total of 973 records.
We removed repeated records and manually screened the retrieved papers to ensure alignment with the specific scope of this survey. A paper was included only if it met the following criteria:
LLM-Based Architecture: The system utilizes Large Language Models as the primary mechanism for generating learner behavior, effectively excluding pre-LLM, rule-based, or hand-crafted simulation systems.
Generative Nature: We strictly distinguished between student modeling and student simulation. In line with our scope, we included systems that take a generative perspective (producing learner-like data or behavior) and excluded pure student modeling systems focused solely on inference, prediction, or analytics without a simulation component.
Explicit Simulation Focus: We excluded general-purpose chatbots or superficial implementations that mimic learner behavior only incidentally. We prioritized studies that explicitly aim to advance the methodology, architecture, or fidelity of student simulation, filtering out works that employ generic personas solely as a utility without technical innovation.
To ensure comprehensive coverage, we performed backward and forward snowballing, manually tracing citations and references of the selected papers to identify high-impact works that may have been missed in the keyword search.
The screening and snowballing process resulted in a final corpus of 73 papers. Based on their primary research objectives, we categorized these works into four distinct clusters: Data Generation (\(n=23\)), Social Learning (\(n=19\)), Teacher Training (\(n=17\)), and Content Evaluation (\(n=14\)).
To illustrate the utility of the proposed framework, we analyze five recent simulated student systems: GPTeach, MATHVC, HypoCompass, Agent4Edu, and Generative Students. We classify each system according to its Dimensional Characterization and Epistemic State Specification (ESS), highlighting how specific design choices mitigate the “competence paradox,” defined here as the tendency of LLMs to default to expert performance, and how these choices necessitate distinct evaluation strategies.
GPTeach [@markel-etal-2023-gpteach] acts as a flight simulator for pedagogy, allowing novice teaching assistants (TAs) in computer science to practice high-stakes office hour interactions in a low-risk environment. Dimensional Characterization:
Behavioral Goals: Simulating Human Aspects. The primary utility lies in reproducing the affective and social dynamics of teaching (e.g., handling frustration, impostor syndrome) rather than modeling precise long-term cognitive decay.
Environment: Context: Intro CS office hours; User: University-level TAs; Modality: Text-based chat.
Epistemic State Specification (ESS): Level E1 (Static Bounded). GPTeach operates at E1 by defining student states through static “persona” templates (e.g., “confused about variable scope”) and scenario-specific constraints. The agent’s ignorance is bound by the prompt description for that specific session, without a persistent memory that evolves across sessions.
Implementation Design: The system is built on a modular “teaching scenario” architecture. Each scenario initializes a student profile containing three distinct layers: (1) Background (name, major, programming experience), (2) Pedagogical State (specific misconception, e.g., confusion between recursion and iteration), and (3) Affective State (e.g., defensive, anxious). To bridge the gap between abstract profiles and realistic dialogue, the system utilizes few-shot prompting populated with excerpts from authentic student-TA transcripts. This retrieval-augmented approach grounds the LLM’s tone, ensuring it mimics the hesitant, often imprecise language of a struggling novice rather than the polished prose of a textbook.
Evaluation Methodology: Evaluation prioritized utility and affective realism over cognitive precision. The authors conducted a mixed-methods study involving TAs who engaged with the tool. Quantitative metrics included self-efficacy ratings and perceived realism scores. Qualitatively, the evaluation focused on the “safe space” affordance: did the simulation allow TAs to experiment with different pedagogical moves (e.g., Socratic questioning vs. direct explanation) without the fear of harming a real student? Results indicated that while the “student” occasionally hallucinated competence (solving the code too quickly), the primary value was the social pressure simulation.
MATHVC [@yue2024mathvc] simulates a collaborative learning environment where a human student solves mathematical modeling problems alongside three LLM-simulated peers, each exhibiting distinct skill levels and social roles.
Dimensional Characterization:
Behavioral Goals: Simulating Learning and Human Aspects. The system models how students with varying aptitudes contribute to and evolve during collaborative problem-solving (CPS).
Environment: Context: Middle school mathematics; Modality: Multi-party dialogue grounded in symbolic tasks.
Epistemic State Specification (ESS): Level E3 (Misconception-Structured). MATHVC exemplifies E3 by implementing a “Symbolic Character Schema.” Instead of relying solely on natural language prompts, the system maintains a structured representation of the specific variables and sub-tasks the agent currently “knows” or “misunderstands.”
Implementation Design: The architecture employs a dual-schema approach. First, a Task Schema defines the ground truth of the math problem. Second, specific errors (e.g., variable omission, incorrect relationship mapping) are injected into this schema to generate a restricted Character Schema for each agent. A specialized Dialogue Controller mediates all interaction: before generating a response, the LLM must first “ground” its intended dialogue act in its restricted schema. If the LLM attempts to reference a variable not present in its character schema (a symptom of the competence paradox), the controller filters or regenerates the response.
Evaluation Methodology: The evaluation framework introduces two novel metrics: Characteristics Alignment and Procedural Alignment. The former verifies that the agent’s errors are causally attributable to its assigned schema (e.g., verifying that a wrong answer stems from the specific variable the agent was programmed to misunderstand). The latter assesses whether the group dynamic follows the natural phases of collaborative problem solving, moving beyond simple “Turing Test” realism to verify structural fidelity.
HypoCompass [@ma-etal-2024-teach] inverts the standard tutoring paradigm: human students act as TAs to assist an LLM-simulated agent that has written buggy code, thereby practicing their own debugging skills.
Dimensional Characterization:
Behavioral Goals: Simulating Performance (Error Production). The goal is to generate high-fidelity artifacts (buggy code) that serve as robust training material.
Environment: Context: Introductory programming; User: Novice programmers; Modality: Code-centric dialogue.
Epistemic State Specification (ESS): Level E1 (Static Bounded). The system uses an E1 specification where the epistemic boundary is materialized as the specific bug injected into the code. The agent’s “knowledge state” is effectively frozen at the moment of error injection.
Implementation Design: To ensure the “student” agent is helpfully incompetent, HypoCompass employs an Over-Generate-Then-Select pipeline. The system first prompts an LLM to generate multiple buggy variations of a correct reference code, filtering for those that are syntactically parsable but logically incorrect. In the interaction phase, the system enforces a strict Boundary of Competence: the LLM agent is permitted to perform code completion (the “hands”) but is explicitly barred from performing hypothesis generation (the “brain”), forcing the human user to articulate the why behind the bug.
Evaluation Methodology: Evaluation focused on the Fidelity of Error and learning outcomes. The authors first validated the quality of the generated bugs to ensure they were non-trivial. Subsequently, a controlled study measured the learning gains of human students. The key finding was that students who taught the HypoCompass agent showed a 12% improvement in debugging post-tests compared to controls, validating the system by its effectiveness as a pedagogical tool.
Agent4Edu [@Gao:2025:A4E:arXiv] is a comprehensive simulation framework designed to generate synthetic learner response data for training Computerized Adaptive Testing (CAT) algorithms.
Dimensional Characterization:
Behavioral Goals: Simulating Performance and Learning. The system targets the production of response patterns that reflect both static ability levels and the temporal evolution of proficiency (learning and forgetting).
Environment: Context: Intelligent Tutoring Systems; Modality: Interaction with algorithmic recommenders.
Epistemic State Specification (ESS): Level E2 (Curriculum-Indexed). Agent4Edu represents an advanced E2 system. It explicitly decouples the learner’s cognitive state from the LLM’s text generation parameters, tracking proficiency as a dynamic variable updated by interaction history.
Implementation Design: The architecture mimics human cognition through three coupled modules: (1) A Learner Profile initialized with real-world data (via Item Response Theory parameters) to capture baseline ability; (2) A Memory Module that integrates an Ebbinghaus forgetting curve, ensuring that the agent’s recall of concepts decays realistically over time; and (3) An Action Module where the LLM generates responses conditioned on the retrieved memory state. This separation allows the system to simulate complex behaviors like the “practice effect” and “slip” driven by state variables rather than random hallucination.
Evaluation Methodology: The evaluation centered on Predictive Validity and Distributional Fidelity. Rather than qualitative chats, the authors assessed whether the synthetic data could train a CAT model as effectively as real data. Metrics included the alignment of psychometric curves (IRT) between simulated and real students, and the performance of downstream recommendation algorithms trained on the synthetic corpus. Results demonstrated that Agent4Edu faithfully reproduced the statistical properties of human learning trajectories.
Generative Students [@genstu] utilizes LLMs to simulate diverse student response patterns to Multiple Choice Questions (MCQs) in Human-Computer Interaction, specifically to identify poorly designed assessment items.
Dimensional Characterization:
Behavioral Goals: Simulating Performance (Misconception Patterns). The objective is to replicate specific confusion patterns so that if a question is ambiguous, the “confused” student will plausibly select a distractor.
Environment: Context: HCI Heuristic Evaluation; Modality: Multiple Choice Questions.
Epistemic State Specification (ESS): Level E3 (Misconception-Structured). This system provides a textbook example of E3 by defining the learner’s state through explicit Knowledge Component (KC) triples: Mastered, Unknown, and Confusion.
Implementation Design: The core innovation is the structured Confusion Tuple (e.g., confusion(KC_A, KC_B)). The prompting strategy explicitly instructs the model that the student cannot distinguish between these
two specific concepts. This constraint forces the LLM to generate errors that are not random guesses, but structural failures consistent with that specific misconception. By systematically permuting these profiles, the system generates a distribution of
responses that reveals which distractors are attractive to which types of learners.
Evaluation Methodology: The system was validated through Item Difficulty Correlation. The authors compared the difficulty ranking of questions derived from the simulated students against historical data from real students. A high correlation indicated that the simulator correctly identified “hard” questions for the right reasons. Qualitative analysis confirmed that the simulated students selected distractors that aligned with their assigned confusion tuples, confirming the system successfully bounded the LLM’s expert priors.
To reduce ambiguity and prevent overclaiming, this reporting card operationalizes ESS as a set of falsifiable requirements. Authors should complete this card for each simulated-student system. A system may only claim level \(E_k\) if it satisfies all criteria for levels \(E_1\) through \(E_k\).
Equal contribution.↩︎