January 01, 1970
This work introduces AIriskEval-edu-db2, a new dataset designed to train and evaluate auditors based on LLMs for an explainable pedagogical risk assessment in instructional content for grades K–12. The dataset comprises 1,639 explanations from 170 curated ScienceQA questions, covering science, language arts, and social sciences. For each question, the dataset includes an explanation written by a human teacher alongside 11 explanations generated by LLM-simulated teacher profiles associated with distinct pedagogical risks. We propose a comprehensive risk rubric aligned with established educational standards that covers five complementary dimensions: factual precision, depth and completeness, focus and relevance, student-level appropriateness, and ideological bias. A key contribution is the addition of 785 explanations with structured explainability annotations, including risk localization and risk description. The annotations are produced through a semi-automatic process with expert teacher validation. Finally, we present validation experiments comparing state-of-the-art proprietary models with a lightweight local Llama 3.1 8B model in both the pedagogical risk detection and the explainability assessment. These experiments evaluate whether supervised fine-tuning on AIriskEval-edu-db2 enables a locally deployable model to approach or outperform stronger frontier models while preserving privacy in educational auditing and assessment tasks.
Large Language Models (LLMs) are increasingly being deployed in human-facing systems where raw capability must be complemented by oversight, automatic monitoring, and risk-aware quality assurance. In educational technology, these models can answer K–12 questions with high precision [1], which has raised growing interest in their use as both tutors and as automatic evaluators of instructional explanations produced by humans or AI. At the same time, because such systems can produce unsafe, unreliable, or misleading outputs, their deployment also raises concerns that are closely related to the themes of automatic monitoring, human factors, artificial intelligence, and the societal impact of security technology[2]. In this sense, educational AI can also be viewed as a security-technology problem: it requires monitoring mechanisms able to detect harmful, untrustworthy, or otherwise risky language outputs before or during deployment.
Recent studies show that LLMs, especially when task-adapted, can approximate key tutoring behaviors [3], while industrial efforts such as LearnLM [4] and studies on the adaptation of instructional roles [5] further support this trend. At the same time, LLMs introduce well-documented risks [6], which are especially consequential in K–12 settings. This motivates automatic pedagogical assessors capable of detecting pedagogical and epistemic risks under established educational frameworks [7].
Previous work suggests that fine-tuning LLMs in rubric-based educational datasets can improve the reliability of the evaluator [8], [9]. However, current public resources remain limited: few target evaluation of instructional explanations rather than general question answering, few adopt a multi-criterion risk perspective, and few support explainable risk assessment beyond binary risk detection.
Taking into account all of the above, this work aims to provide a resource for pedagogical risk assessment in educational contexts: a dataset of K–12 instructional explanations annotated with human-reviewed risk labels, designed to support both the training and evaluation of LLM-based pedagogical evaluators. Such evaluators can be applied to explanations generated not only by AI tutors but also by human teachers, enabling monitoring, auditing, and quality assurance of instructional content in digital learning environments.
Relative to our previous work EduEVAL-DB [9], here we introduce the following extensions (see Fig. 1):
Extended dataset with new LLM-generated explanations. We release a dataset, AIriskEval-edu-db22, with an additional partition generated using Gemini 2.5 Pro. It comprises 139 questions for each of five teacher profiles, plus 90 additional questions for the Sarcastic Teacher, substantially expanding the diversity of profile-conditioned pedagogical explanations.
Explainable pedagogical risk annotations. Whereas the original dataset provided only binary labels, this new version introduces paired explainability data for risk-positive cases: risk localization (text excerpt) and risk description (the underlying rationale).
New evaluation training and benchmarking experiments. We present experiments in which a lightweight local model is fine-tuned not only for binary pedagogical risk detection but also for explainable risk assessment. In addition, we study a combined setting in which the original and extended dataset partitions are merged under a 5-fold cross-validation protocol, yielding better results than training on either partition alone.
Recent benchmarks evaluating Large Language Models (LLMs) in tutoring roles often emphasize mathematics and dialogic strategies. Foundational efforts like ScienceQA [10] offer multi-domain K–12 questions with human reference explanations but lack specific pedagogical annotations. To assess interactive teaching, MathDial [11] introduces dialogs in which human teachers scaffold an LLM-simulated student through errors. Subsequent benchmarks explicitly target math response quality: SocraticMATH [12] provides dialogs annotated with structured teaching stages, while MRBench [13] human-annotates math responses across multiple pedagogical dimensions. The BEA 2025 Shared Task [14] further standardizes the evaluation through a public dataset with gold pedagogical annotations.
Other works emphasize outcome-based or simulation-based evaluation; EducationQ [15] measures teaching quality through multi-agent simulations using learning gains without releasing reusable dialogs, whereas SocraticLM [16] introduces a large-scale simulated dataset of Socratic math tutoring. Beyond core pedagogy, Weissburg et al. [17] evaluate demographic bias through differential explanation selection in varying student profiles. Although advancing specific evaluation axes, existing work less commonly unifies multi-domain K–12 contexts, broad pedagogical risks, and different teacher roles. Addressing these gaps, our previous work, EduEVAL-DB [9], performs the identification of pedagogical risk in multi-domain K–12 instructional explanations.
We define a pedagogical risk rubric for K–12 instructional explanations aligned with educational standards [7]. Beyond factual accuracy, it evaluates whether an explanation is pedagogically effective and appropriate for the learner. The rubric contains five dimensions, assigned to honesty (H1), helpfulness (H2), and harmlessness (H3) [18].
Factual Accuracy (H1): captures false or hallucinated claims that may create persistent misconceptions [19]. Explanatory Depth & Completeness (H2): measures whether reasoning is properly developed rather than merely stated [20]. Focus & Relevance (H2): evaluates whether the explanation includes unnecessary information that increases cognitive load [21]. Student-Level Appropriateness (H3): evaluates whether difficulty and language match the learner’s developmental level and ZPD [22]. Ideological Bias (H3): screens for stereotypes, exclusionary framing, or hidden-curriculum effects [23].
The rubric was designed to be both orthogonal and comprehensive: orthogonal in separating distinct instructional failure modes, and comprehensive in covering content quality, pedagogical scaffolding, learner fit, and ethical representation within the Instructional Core [24]. This final dimension is especially important given the broader concerns about harm in LLM-based educational systems [25].
We build AIriskEval-edu-db2 from 170 K–12 ScienceQA questions [10], consisting of 1,639 explanations in eleven LLM-simulated teacher profiles. Evaluated across five risk dimensions (Section 3), it yields 8,195 binary risk annotations developed in two stages:
I) EduEVAL-DB: The original dataset pairs one human explanation with six GPT-5-generated profile explanations for 139 questions (854 explanations, 4,270 labels). The first four criteria (Factual Accuracy, Focus & Relevance, Depth & Completeness, Student-Level Appropriateness) each contain 139 positive and 715 negative labels. Ideological Bias contains 20 positive and 834 negative labels.
II) Extension Subset: This newly introduced partition comprises 170 questions and 785 explanations (3,925 labels) generated via Gemini 2.5 Pro. Generations utilized a few-shot prompting, anchoring on the ScienceQA reference, and explicit grade levels. Standard profiles cover the original 139 questions, while a sarcastic teacher profile covers 45 questions (31 new) generating two explanations per question at mild” and severe” intensities. The first four criteria each yield 139 positive and 646 negative labels, while ideological bias yields 90 positive and 695 negative. A major novelty is the introduction of explainability annotations for all flagged risks: risk localization (exact text excerpts) and risk description (decision rationale) (see Fig. 2).
The dataset is built around six teacher-inspired profiles designed to reflect realistic instructional styles and shortcomings: Exemplary Teacher, Rambling Teacher, Concise Teacher, Inaccurate Teacher, Overly Advanced Teacher, and Sarcastic Teacher. These explanations were generated with the Gemini 2.5 Pro API through prompt engineering and few-shot prompting, using the ScienceQA teacher explanation as a pedagogically sound reference and explicitly conditioning on grade level. All profiles covered the full set except Sarcastic Teacher, limited to 45 questions, with two explanations per question, by moderation constraints, and manually authored with LLM assistance. Teacher explanations were capped at 400 characters (typically 150–300) to reduce length bias. Fig. 2 shows representative examples.
The annotation process is carried out in two stages. In Stage I, binary labels are assigned to each explanation generated by an LLM-simulated teacher profile through a semi-automatic procedure derived from the intended risk profile of that profile. To validate this procedure, two experienced teachers manually reviewed approximately 30% of the dataset and confirmed the consistency of the assigned labels. Once this consistency was established, the labeling procedure was accepted. Subsequently, any cases that produced disagreements during evaluator-based validation (Section 5) were re-examined by the experts. In Stage II, each risk-positive case is enriched with explainability annotations, namely risk localization and risk description. These annotations are first generated by Gemini and then validated by expert teachers.
We evaluate LLMs as automatic pedagogical assessors under the proposed rubric in three dataset settings: the previously introduced pedagogical risk dataset [9], the newly introduced explainability-enhanced partition and the complete AIriskEval-edu-db2 dataset obtained by merging both. As baselines, we consider Gemini 2.5 Pro, GPT 5.5, and Llama 3.1 8B Instruct, the latter being a lightweight local model that can be deployed on consumer GPUs. Evaluation covers both binary pedagogical risk detection and explainability assessment, including risk localization and risk description. To assess the usefulness of AIriskEval-edu-db2 for training local evaluators, we also fine-tune Llama 3.1 8B Instruct under the same protocol.
| Evaluation on EduEVAL-DB [9] | Evaluation on Extension Subset | Evaluation on AIriskEval-edu-db2 | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| Criteria | Gemini (Base) | Llama (Base) | Llama (EE-DB) | Llama (Ext) | Llama (EE-DBX) | GPT (Base) | Llama (Base) | Llama (Ext) | Llama (EE-DBX) | Llama (EE-DBX) |
| FA | 0.012 | 0.164 | 0.048 | 0.115 | 0.049 | 0.051 | 0.170 | 0.057 | 0.056 | 0.053 |
| F&R | 0.216 | 0.227 | 0.069 | 0.195 | 0.029 | 0.037 | 0.195 | 0.017 | 0.012 | 0.021 |
| D&C | 0.320 | 0.288 | 0.048 | 0.080 | 0.031 | 0.228 | 0.253 | 0.023 | 0.027 | 0.029 |
| S-L A | 0.054 | 0.220 | 0.003 | 0.027 | 0.001 | 0.031 | 0.170 | 0.001 | 0.000 | 0.001 |
| IB | 0.008 | 0.049 | 0.006 | 0.022 | 0.006 | 0.013 | 0.088 | 0.006 | 0.003 | 0.004 |
Evaluation Setup. The AIriskEval-edu-db2 dataset was used for both training and evaluation in all experiments. All evaluations followed a standardized protocol: for a given instance, the evaluated models received the student’s question,
the target grade level, and the generated tutor explanation. To avoid evaluation bias, the specific tutor profile was never provided in the prompt. All baseline evaluators, as well as fine-tuned Llama 3.1 8B models, were evaluated in zero-shot inference
mode. The models were provided with an instruction prompt that defined the rubric criteria, but did not contain task-specific few-shot examples. The prompt required a structured JSON output containing binary labels (0 or 1) for each dimension of
pedagogical risk. Furthermore, if a model predicted the presence of a risk (label 1), it was required to generate structured explainability data consisting of two fields. Namely, risk localization: extracting the exact excerpt from the tutor explanation
where the risk was identified, and risk description: generating a natural language justification explaining why the excerpt was flagged.
Dataset Distribution. The evaluation spans three partitions. First, EduEVAL-DB (854 explanations) contains 139 positive labels for the first four criteria and 20 for Ideological Bias. Second, the Extension Subset (785 explanations)
includes 139 positive labels for the first four criteria and 90 for Ideological Bias. Finally, the combined AIriskEval-edu-db2 dataset totals 1,639 explanations, yielding 278 positive labels for the first four criteria and 110 for Ideological Bias.
Negative labels make up the remainder of each set. Excluding 139 human written references leaves exactly 1,500 simulated explanations for the final evaluation pool.
Fine-Tuning Experiments. To assess the impact of supervised learning on evaluation capabilities, a fine-tuning experiment was conducted on the Llama 3.1 8B model on the extension subset. Supervised fine-tuning was performed using Low-Rank
Adaptation (LoRA) under a 5-fold cross-validation protocol. In each fold, 80% of the data was used for training and the remaining 20% for testing; to prevent data leakage, the splits were strictly grouped by question ID, ensuring that all explanations
associated with a given question appeared exclusively in the training or evaluation set. At the same time, the divisions were stratified by tutor profile. Across the five folds, each example appeared exactly once in the evaluation set. This allowed the
fine-tuned model to produce out-of-sample predictions for the entire dataset, ensuring a fair and direct comparison with the baseline Llama and GPT models. Training used the exact same input structure as the inference stage, with the target sequence
defined by the ground-truth JSON encoding the binary risk labels and explainability texts. Optimization was driven by a standard causal language modeling loss. The LoRA configuration utilized a rank (\(r\)) of 16, a scaling
factor (\(\alpha\)) of 32, and a dropout rate of 0.05. The model was trained for two epochs with a learning rate of \(1\times10^{-4}\), employing gradient accumulation to achieve an
effective batch size of 2. To test generalization capabilities, the Llama model fine-tuned exclusively on this new dataset was subsequently zero-shot evaluated on the dataset introduced in [9]. Finally, a joint fine-tuning experiment was conducted. Llama 3.1 8B was trained simultaneously on the full dataset (AIriskEval-edu-db2), using a similar 5-fold cross-validation protocol.
Then, this jointly fine-tuned model was systematically evaluated against both subsets.
Evaluation Metrics. Detection performance across the rubric dimensions is reported using Mean Absolute Error (MAE) between the predicted binary labels and the ground-truth annotations. For explainability assessment, performance was
measured using metrics tailored to the nature of the generated text. Specifically, risk localization was measured using Intersection over Union (IoU) to quantify the degree of token-level overlap between the model’s extracted excerpt and the ground-truth
risk span. The risk description was evaluated using a combination of lexical and semantic metrics. Token-level F1, BLEU and ROUGE-L were used to capture lexical overlap and sequence alignment, while BERTScore was used to assess semantic similarity, which
is particularly important given the open-ended nature of the generated justifications.
| Criterion | Metric | Llama (Base) | GPT (Base) | Llama (Ext) | Llama (EE-DBX) | |
|---|---|---|---|---|---|---|
| FA | Risk Localization | IoU | 0.281 | 0.678 | 0.647 | 0.611 |
| Risk Description | Token-F1 | 0.153 | 0.303 | 0.501 | 0.496 | |
| BLEU | 0.036 | 0.055 | 0.269 | 0.249 | ||
| ROUGE-L | 0.127 | 0.254 | 0.464 | 0.449 | ||
| BERTScore | 0.392 | 0.834 | 0.914 | 0.911 | ||
| F&R | Risk Localization | IoU | 0.127 | 0.947 | 0.971 | 0.978 |
| Risk Description | Token-F1 | 0.040 | 0.231 | 0.613 | 0.632 | |
| BLEU | 0.005 | 0.020 | 0.422 | 0.414 | ||
| ROUGE-L | 0.032 | 0.157 | 0.604 | 0.627 | ||
| BERTScore | 0.126 | 0.862 | 0.937 | 0.940 | ||
| S-L A | Risk Localization | IoU | 0.088 | 0.946 | 0.985 | 0.976 |
| Risk Description | Token-F1 | 0.056 | 0.369 | 0.838 | 0.837 | |
| BLEU | 0.006 | 0.039 | 0.734 | 0.729 | ||
| ROUGE-L | 0.050 | 0.253 | 0.836 | 0.834 | ||
| BERTScore | 0.160 | 0.887 | 0.969 | 0.969 | ||
| IB | Risk Localization | IoU | 0.736 | 0.835 | 0.972 | 0.972 |
| Risk Description | Token-F1 | 0.205 | 0.164 | 0.389 | 0.419 | |
| BLEU | 0.010 | 0.009 | 0.091 | 0.092 | ||
| ROUGE-L | 0.145 | 0.145 | 0.358 | 0.384 | ||
| BERTScore | 0.803 | 0.820 | 0.903 | 0.909 | ||
Table 1 reports the Mean Absolute Error (MAE) achieved by the baseline models (Gemini 2.5 Pro, GPT 5.5, and Llama 3.1 8B), together with the fine-tuned Llama 3.1 8B models.
Focusing first on the new explainability-enhanced partition, GPT consistently outperforms the baseline Llama model across all criteria, which is expected given its stronger general capabilities. However, supervised fine-tuning yields a substantial improvement; the Llama model fine-tuned on the extension subset (Llama (Ext)) outperforms baseline Llama and GPT in all but one criterion.
For Factual Accuracy on the extension subset, both GPT and Llama (Ext) considerably outperform baseline Llama, with GPT leading by a marginal amount over this fine-tuned model. A plausible cause is that performance on this criterion depends more strongly on the evaluator’s general knowledge of the world, where frontier models such as GPT remain stronger. By contrast, on the original dataset (EduEVAL-DB [9]), Gemini achieved the best results for this metric, with a clearly lower overall error rate. The difference observed here can be attributed to the nature of the extension subset: factually incorrect explanations often preserve the correct final answer but contain subtle misconceptions or localized errors within the text. This makes detection much more challenging for zero-shot models and explains the performance gap.
For the remaining criteria, Llama (Ext) outperforms both baseline Llama and GPT. For Focus & Relevance, GPT reduces the error of the baseline Llama by nearly a factor of eight, yet the fine-tuned Llama leads with an MAE of 0.017. Depth & Completeness proves difficult for both GPT and the baseline Llama, which exhibit poor performance. In contrast, the fine-tuned Llama reduces the error 10 times. This is unsurprising, as assessing completeness is a highly nuanced task that depends heavily on the specific annotation criteria used to determine what counts as missing information. For Student-Level Appropriateness, GPT achieves a very low error rate relative to the baseline, whereas the fine-tuned Llama makes no errors on this criterion. For Ideological Bias, both Llama (Ext) and GPT clearly outperform the baseline Llama, with Llama (Ext) performing best. The experiments also reveal strong generalizability. When the Llama model fine-tuned exclusively on the extension subset is evaluated on EduEVAL-DB, it achieves notable improvements across all criteria relative to the baseline Llama. This highlights the robustness of the extension subset and its capacity to transfer effectively across dataset settings.
Fine-tuning Llama on the full AIriskEval-edu-db2 dataset (Llama (EE-DBX)) provides the best overall results among the fine-tuned models. On the extension subset, it outperforms Llama (Ext) in all criteria except Depth & Completeness, while on EduEVAL-DB it improves over Llama fine-tuned in that dataset (Llama (EE-DB)) in all criteria except Factual Accuracy and matches it in Ideological Bias. This suggests that combining both subsets introduces useful variability and improves generalization. The improvement over Llama (EE-DB) in EduEVAL-DB is larger than the improvement over Llama (Ext) in the extension subset. This may reflect the added value of the explainability annotations in the extension data to support risk identification. Compared with the state-of-the-art baselines, Llama (EE-DBX) outperforms GPT on the Extension Subset and Gemini on EduEVAL-DB in all criteria except Factual Accuracy, which is expected because factual assessment depends more on general model knowledge. The last column reports its performance on the full AIriskEval-edu-db2 dataset.
Table 2 details the explainability metrics, specifically the location and description of the risk for the extension subset.
In terms of risk localization, the baseline Llama performs poorly and is significantly outperformed by both GPT and the fine-tuned Llamas on all applicable criteria. Factual Accuracy proves to be the most challenging dimension to locate, with GPT and fine-tuned Llamas achieving IoU scores of 67.8%, 64.7% and 61.1%, respectively. For the remaining criteria, these models achieve high localization accuracy, with GPT exceeding 80% IoU and fine-tuned Llamas exceeding 95%, establishing them as the top performers.
Regarding risk description, evaluated primarily via Token-F1 (T-F1), the baseline Llama again demonstrates the poorest performance. GPT provides better explanatory descriptions, with T-F1 scores ranging between 0.15 and 0.37, performing worse on Ideological Bias and best on Student-Level Appropriateness. Meanwhile, the fine-tuned Llamas exhibit the strongest descriptive capabilities across all criteria. Their T-F1 scores range from 0.39 to 0.84, following the same trend as GPT by achieving its highest descriptive accuracy again on Student-Level Appropriateness and its lowest on Ideological Bias. The results for the remaining metrics show a very high concordance with these results, with BERTScore achieving high scores across all models but confirming the same pattern.
In this work, we introduced an extended dataset for the explainable assessment of pedagogical risks in educational content. Moving beyond binary classification to include risk localization and natural language descriptions, the proposed approach supports a more transparent analysis of instructional quality across key pedagogical dimensions. These explainability features make the evaluator more useful in practice, providing specific and actionable feedback for auditing both human educators and LLM-based tutors. Our results also show that requiring the model to justify flagged risks improves its binary detection performance, while the generated descriptions help verify that the model is applying the pedagogical rubric rather than relying on superficial statistical patterns. Finally, fine-tuning a lightweight Llama 3.1 8B model demonstrates that robust pedagogical auditing can be performed locally, preserving student and institutional data privacy while avoiding recurring API costs.
Future work will extend this framework to dynamic student–teacher interactions by adapting the rubric to multi-turn dialogs. We also plan to integrate these evaluators into multimodal and role-based learning analytics platforms [26], [27], combining explainable instructional evaluation with cognitive and behavioral cues [28], [29] for holistic modeling of the learning-process. We will also study how multimodal LLMs (including VLMs [30]), and image-based agent representations (including avatars [31]), influence the learning experience. Analyzing biases [32], [33] and synthetic manipulation [34] are also key for us, as well as adapting AIriskEval to other setups beyond e-learning, e.g., gaming [27].
This research was supported by Cátedra ENIA UAM-VERIDAS en IA Responsable (NextGenerationEU PRTR TSI-100927-2023-2), M2RAI (PID2024-160053OB-I00, MICIU/FEDER), TRUST-ID (PID2025-173396OB-I00, MICIU/AEI and the EU) and PowerAI+ (SI4/PJI/2024-00062, Comunidad de Madrid and UAM). Javier Irigoyen is supported by an FPI fellowship from MINECO/FEDER.↩︎
https://github.com/BiometricsAI/AIriskEval-edu↩︎