Qualified Educational Capacity Planning under Heterogeneous Student Support Needs: A Synthetic Benchmark and Decision-Support Framework

Carlos Eduardo Sanoja
Quanta Labs, LLC
Professor, FCEA, Universidad Monteávila
Edificio Lomas del Sol, Calle Humboldt, Lomas del Sol, Caracas, Venezuela
csanoja@somosquanta.com
ORCID:
0009-0000-0339-7072

,

Oscar Enrique Moreno Mayz
Quanta Labs, LLC
omoreno@somosquanta.com


Abstract

In educational support services, the binding resource is often staff time that is both available and qualified for the task — and qualification is dynamic: preparation decays, new support needs arrive that nobody is yet prepared for, and training consumes the same staff hours that current students need. We introduce a benchmark specification and decision-support framework for qualified educational capacity planning. The model is a stylized single-institution service system with heterogeneous support-demand categories, backlog-only dynamics (services are not storable), continuous preparation states with hard threshold qualification and decay, and capacity-consuming training, hardened before implementation by an adversarial multi-agent design review. The benchmark provides six seed-controlled scenario families — announced and surprise new support categories, staff absences, and demand surges — with exact feasibility discipline, declared per-policy information sets, within-episode requalification and greenfield-qualification counters, access-dispersion metrics, replay checksums, and paired statistics. We compare service-only, reactive, static-insurance, water-filling, and rolling-horizon mixed-integer controllers, with an attribution chain separating service planning, qualification maintenance, and acquisition, and a perfect-foresight reference. Our central result is a regime map governed by one quantity: whether a newly required qualification can be acquired within the controller’s reaction reach (its planning horizon plus the window over which backlog stays recoverable). When it can — the regime that includes the core instance, and a frozen \(2{,}620\)-episode adversarial suite spanning a shock-focused static plan, disruption windows down to two periods, training rates over a \(4{\times}\) range, and low slack — the closed-loop controller dominates (\(158\) paired cells, no static plan wins), and the attribution chain locates the value in just-in-time qualification acquisition. When it cannot — a cold start with deep retraining, so the training lag exceeds the horizon — a \(4{,}420\)-episode boundary search shows lean static insurance winning by up to \(2\times\) on long windows; the win is structural, since the perfect-foresight peer ties the controller exactly and also loses. A reactive trainer that starts after onset wastes effort and is worst of all, and once demand structurally outruns qualified capacity no policy choice matters. Backlog perishability, tested both post-hoc and as genuine dynamics, shifts this boundary without erasing either regime. This contrasts with our companion manufacturing studies, where lean static insurance holds a regime near the capacity boundary. A transparent scenario-analysis interface, EduCapacity Studio, reproduces any exported scenario bit-for-bit. All evidence is stylized and synthetic; the framework makes no claims about real student outcomes, compliance, or individual placements.

1 Introduction↩︎

Educational institutions plan around a resource that standard operations models treat as fixed: qualified staff time. Yet the capacity that matters for student support is not headcount — it is staff time that is both available and qualified for the specific support task at hand, whether that is remediation, language support, individualized learning support, or assistive-technology assistance. Qualification is dynamic: preparation fades without practice, new support needs arrive that nobody on staff is yet prepared for, and the only way to create qualified capacity — training — consumes the very staff hours that current students need. The literature documents persistent shortages and turnover of qualified support personnel [1], [2], regulatory qualification requirements [3], and the workload pressure on the staff who remain [4][6]; practice tools for support-service scheduling [7][9] confirm the operational need. What is missing is a reproducible way to study the resulting decision problem.

The decision problem is genuinely dynamic and genuinely hard. Serving demand today and being qualified to serve demand tomorrow draw on the same staff hours, so upskilling is never free. Qualifications decay, making preparation a maintained asset rather than a one-off credential [10], [11]. Disruptions — a new support category, staff absences, a demand surge — can be announced in advance or arrive as surprises, and because educational services cannot be stored, capacity that becomes qualified too late cannot make up for service that was never delivered. Reacting after a shock can therefore be structurally too late, while insuring against every contingency in advance is expensive precisely because training consumes service time.

Existing education-operations research does not provide a testbed for this trade-off. Timetabling benchmarks [12][14] optimize courses and rooms with static staff qualifications; scheduling and assignment decision-support systems [15][17] treat qualification as input, not state; workforce-planning research outside education models skills and training [18][20] but not the education-specific combination of support-demand categories, non-storable service, and announced-versus- surprise information regimes; and learning-analytics dashboards [21], [22] visualize rather than plan. We position this paper in the gap: a benchmark specification, not a deployed scheduler.

We make four contributions. (1) A stylized dynamic model of qualified educational staff capacity: heterogeneous support-demand categories, backlog-only service dynamics (services are not storable), continuous preparation states with hard threshold qualification, decay, and training as a capacity-consuming action — locked before implementation through a multi-agent design review with an adversarial pass. (2) A reproducible synthetic benchmark: six seed-controlled scenario families spanning announced and surprise new-category, absence, and surge shocks, with exact feasibility discipline, declared per-policy information sets, replay checksums, and paired statistics. (3) A policy study comparing service-only, reactive, static-insurance, water-filling, and rolling-horizon MILP controllers — with an attribution chain separating service planning, qualification maintenance, and acquisition, and an oracle reference labeled as such. (4) EduCapacity Studio, a transparent scenario-analysis interface whose export/replay loop reproduces results bit-for-bit, making the interface a reproducibility instrument rather than a recommendation engine.

The empirical message is sharper than, and instructive against, our companion manufacturing studies. There, lean static insurance holds a regime: under surprise shocks near the capacity boundary, reacting after onset is structurally too late, so pre-bought cross-training wins. Here we map both sides of that boundary and report the map as the finding. The governing quantity is whether a newly required qualification can be acquired within the controller’s reaction reach — its planning horizon plus the window over which backlog stays recoverable. In the reaction-feasible regime, which includes the core instance, the closed-loop controller dominates, and we are careful to show this is legitimate rather than an artifact of weak comparators: the static baselines are lean, monotone, and exactly feasible, they pre-qualify the shock category, and a frozen adversarial suite of \(2{,}620\) episodes — a shock-focused static plan, disruption windows down to two periods, training rates across a \(4\times\) range, and low slack — finds no static win across \(158\) paired cells. But when we push the training lag past the controller’s horizon (a cold start with deep retraining), a \(4{,}420\)-episode boundary search locates the other side: lean static insurance wins by up to \(2\times\) on long windows, and the win is structural rather than a controller defect, because the perfect-foresight peer ties the controller exactly and also loses. Two further regimes complete the map — a reactive trainer that starts after onset wastes effort and is worst of all, and under structural capacity insufficiency no policy choice matters — and backlog perishability shifts the boundary without erasing it. We treat this reaction-versus-pre-positioning map, and the cross-domain contrast it draws with our manufacturing studies, as the finding. All results are stylized computational evidence on synthetic scenarios; the framework makes no claims about real student outcomes, legal compliance, or individual placements.

2 Related Work↩︎

2.0.0.1 Educational timetabling and scheduling benchmarks.

Education operations research has mature, reusable benchmarks for timetabling: the international timetabling competitions and their instance formats [12], [13], [23], [24], curriculum-based formulations [25][27], and school timetabling surveys [14], [28], [29]. These benchmarks optimize courses, rooms, sections, and assignments under rich constraints, and they define the template — instances, validators, baselines — that any new education-operations benchmark must follow. They do not model staff qualification as a dynamic state, nor training as a decision.

2.0.0.2 Decision support for educational planning.

Web-based scheduling and assignment systems are likewise established: udpSkeduler [15], group decision support for academic term preparation [30], optimization-based student-to-teacher assignment [16], and teacher–course assignment with preferences and workload [31], [32]; operational research in education is surveyed by [33]. The closest dynamic threat is predict-and-prescribe course scheduling [17], which scales course/section/location capacity from forecasts on real university data; its capacity object is sections and rooms, not qualified support-service staff, and upskilling never enters as a service-time opportunity cost.

2.0.0.3 Workforce scheduling and training.

Outside education, staff scheduling and multi-skilled workforce planning are mature [18], [19], [34], including joint timetabling with trainer rostering [20] and training/service delivery planning [35]. We adapt these ideas — and the qualified-capacity control formulation of our companion manufacturing studies — to an education-specific setting; the skills-as-state, training-as-control structure is not claimed as new.

2.0.0.4 Qualified educational support staff.

The practical relevance of qualification-constrained support capacity is documented in the special-education and student-services literature: personnel qualification requirements [3], shortages and attrition of qualified special educators [1], [2], certification breadth [36], [37], in-service training [10], [11], [38], paraprofessional employment [39], workload/caseload guidance [4], [5], and teacher time use [6]. Practice tools for support-service scheduling [7][9] prove the operational need while also limiting any product-novelty claim; none provides a peer-reviewed, reproducible policy benchmark.

2.0.0.5 Learning analytics dashboards.

Dashboard and advisor-analytics research is extensive [21], [40][44] and increasingly critical of impact claims [22], [45][47]. This blocks interface novelty: our EduCapacity Studio is deliberately a transparent wrapper around the benchmark — a scenario/replay layer with declared information sets — not an analytics dashboard with outcome claims.

2.0.0.6 Positioning.

Prior work covers timetabling benchmarks, educational scheduling DSSs, assignment models, workforce planning with skills, and dashboards. What we did not find documented is the full intersection this paper targets: a reproducible synthetic benchmark in which heterogeneous support-demand categories meet dynamic qualified-staff capacity, upskilling consumes the same staff hours as current service, demand and absence shocks come in announced and surprise information regimes, and policy classes — including a rolling-horizon controller — are compared under paired statistics with a replayable decision-support interface. The contribution is that testbed and its regime analysis, not a new timetabling solver, staffing product, or dashboard.

3 Problem Formulation↩︎

We model one educational institution (a school, campus, department, or support center) as a dynamic service-capacity system. The scope is deliberately stylized: planning happens at the level of support-demand categories, never individual students; no protected attributes, legal compliance logic, or learning-outcome models are included; and there is no master timetable, room, or transportation modeling. All instances are synthetic and seed-controlled.

3.0.0.1 Sets and parameters.

Periods \(t \in \{0,\dots,T-1\}\) (one period is one service block; \(T = 40\)), support-demand categories \(g \in \mathcal{G}\), staff members \(w \in \mathcal{W}\), and qualifications \(k \in \mathcal{K}\) with the identity map \(k(g) = g\) in the MVP (one required qualification per category). Static parameters: staff hours \(A_{w,t} \le H_w\) per period, qualification thresholds \(\theta_k\), training gain \(\alpha_k\) per hour, decay \(\delta_k\) per period, training-seat slots \(cap^{\mathrm{train}}_k\), and cost coefficients \(c^B_g\) per unmet hour-period and \(c^Y\) per training hour.

3.0.0.2 State.

At period \(t\): backlog of unmet support hours \(B_{g,t} \ge 0\); staff availability \(A_{w,t}\); continuous preparation levels \(S_{w,k,t} \in [0,1]\) with hard qualification \[Q_{w,k,t} = \mathbf{1}\!\left[S_{w,k,t} \ge \theta_k\right]; \label{eq:p3qual}\tag{1}\] and a demand-forecast window \(\hat{D}_{g,t:t+F}\) whose content depends on the scenario’s information regime (announced shocks appear in the window before onset; surprise shocks are hidden until they occur). There is no inventory state: educational services are not storable, so capacity unused today cannot serve tomorrow’s demand.

3.0.0.3 Actions and feasibility.

Each period the planner assigns service hours \(x^{\mathrm{service}}_{w,g,t} \ge 0\) and training hours \(x^{\mathrm{train}}_{w,k,t} \ge 0\) subject to \[\begin{align} \sum_g x^{\mathrm{service}}_{w,g,t} + \sum_k x^{\mathrm{train}}_{w,k,t} &\;\le\; A_{w,t} \quad \forall w, \tag{2}\\ x^{\mathrm{service}}_{w,g,t} \;\le\; A_{w,t}\, Q_{w,k(g),t}, \qquad \bigl|\{w : x^{\mathrm{train}}_{w,k,t} > 0\}\bigr| &\;\le\; cap^{\mathrm{train}}_k \quad \forall k. \tag{3} \end{align}\] Constraint 2 is the central mechanism: upskilling consumes the same scarce staff hours that direct service needs now. Eligibility is hard — only qualified staff can serve a category.

3.0.0.4 Dynamics.

Realized demand is the scenario mean with multiplicative truncated noise, gated to active categories. Served hours and the backlog-only queue evolve as \[\begin{align} y_{g,t} = \sum_w x^{\mathrm{service}}_{w,g,t}, \qquad B_{g,t+1} &= \max\!\bigl(0,\; B_{g,t} + D_{g,t} - y_{g,t}\bigr), \tag{4}\\ S_{w,k,t+1} = \Pi_{[0,1]}\!\bigl((1-\delta_k)\, S_{w,k,t} + \alpha_k\, x^{\mathrm{train}}_{w,k,t}\bigr), \qquad Q_{w,k,t+1} &= \mathbf{1}\!\left[S_{w,k,t+1} \ge \theta_k\right]. \tag{5} \end{align}\] Decay makes qualification a maintained asset: an unrefreshed qualification eventually lapses and can only be recovered by further training.

3.0.0.5 Objective and information regimes.

The per-period cost is \(c_t = c^B \sum_g B_{g,t+1} + c^Y \sum_{w,k} x^{\mathrm{train}}_{w,k,t}\), and a policy is a causal map from the observation to actions minimizing \(\sum_t c_t\) under the scenario’s disruption process (new support categories, staff absences, and demand surges, each in announced and surprise variants). Every policy declares its information set explicitly, and oracle references are labeled as such. The benchmark’s decision tension is intertemporal: serving demand today and being qualified to serve demand tomorrow draw on the same staff hours through 2 , while hard qualification 1 makes future capacity a discrete, training-lagged consequence of today’s allocation — and because services cannot be stored 4 , capacity that arrives too late cannot be made up in advance.

4 Benchmark Design↩︎

The benchmark is a reproducible, seed-controlled environment realizing the formulation of Section 3, hardened before implementation by a multi-agent design review (two specialist proposals reconciled by an adversarial pass that rejected, among others, a decay rate that degenerated the no-shock control and a category count that was infeasible by construction).

4.0.0.1 Frozen instance.

Eight staff members with \(8\) hours per period; \(G \in \{3, 6, 9\}\) support categories with \(G = 6\) as the default and the category-count sweep holding the slack ratio constant (\(\bar{D}_g = 64/(1.5\,G)\)) so that coverage differences come from qualification breadth rather than raw overload; \(\theta = 0.6\), \(\alpha = 0.08\) (one cold-start qualification costs \(7.5\) training hours, roughly one staff-period), \(\delta = 0.01\); per-category demand mean \(8.0\) hours (slack ratio \(1.33\)) with multiplicative truncated noise (\(\sigma = 0.2\)); three training seats per qualification per period. The deterministic initial qualification matrix is lean on purpose: two home categories per staff member at \(S_0 = 0.8\), everything else warm but unqualified (\(0.25\)), and the latent shock category unqualified for everyone (\(0.2\)) — so a new support category can never be served without prior training.

4.0.0.2 Scenario families and information regimes.

Six seed-controlled families: a stationary no-shock control; a new support category activating mid-episode (demand \(12\) h for \(20\) periods, onset randomized per seed) in announced (visible once the \(F{=}4\) forecast window reaches the onset) and surprise (hidden until onset) variants; staff absence (three of eight staff unavailable for eight periods) in announced and surprise variants; and a demand surge (\(\times 1.75\) for ten periods) that pushes the system transiently over capacity. Realized onsets are logged for replay.

4.0.0.3 Metrics.

Operational: total unmet support hours, mean coverage ratio (\(\min(1, y/(D{+}B))\) over active category-periods — the storable-goods notion of fill rate does not apply), peak and area-under backlog, staff utilization. Resilience: recovery rate and time (backlog within half a period of demand of its pre-shock level for two consecutive periods), unrecovered counts. Capability: within-episode requalifications and greenfield qualifications (first-ever threshold crossings), training hours by staff and qualification. Access dispersion: Jain indices over per-category coverage and per-staff training hours plus minimum floors — reported strictly as coverage/training dispersion; no demographic attributes are modeled and no fairness claims are made. Reproducibility: seeds, configs, solver status and runtime, fallback counts, and a replay checksum over the backlog/qualification trajectory.

4.0.0.4 Feasibility discipline.

The environment validates every action and repairs infeasibility deterministically while counting violations and qualification-zeroed hours; all policies in this paper emit exactly feasible actions, so these diagnostics are identically zero in every reported run.

4.0.0.5 Decision-support interface.

EduCapacity Studio, a research demonstrator, exposes the benchmark through seven screens (scenario builder; category, staffing, and training editors; policy comparison with each policy’s declared information set; a bottleneck explorer; and scenario replay/export). Exported scenarios store the configuration, seed, and replay checksums; re-importing re-runs the simulator and verifies identical results, making the interface itself a reproducibility check rather than a recommendation engine. It is not a scheduling, compliance, or student-placement tool.

5 Policies and the Rolling-Horizon Controller↩︎

5.0.0.1 Baseline policy classes (explicit information sets).

ServiceOnly sees the current backlog, nowcast demand, qualifications, and availability; it never trains and allocates service hours proportionally to outstanding demand (a water-filling split — the serve-scarcest-to-exhaustion rule was rejected because it oscillates near tight capacity). ReactiveGap additionally trains toward a category’s qualification only after its uncovered demand exceeds a threshold. StaticInsurance\(\{b\}\) executes a fixed lean cross-qualification plan from \(t=0\) using only the \(t=0\) configuration, spending a budget fraction \(b \in \{0.05, 0.10, 0.20\}\) of total staff hours spread over half the horizon (\(b = 0.20\) is a deliberately saturating endpoint whose per-period diversion exceeds typical slack, documented as such). WaterFillingTraining spreads a \(10\%\) training budget across observed qualification gaps weighted by visible demand. OracleUpperReference runs the same receding-horizon controller (\(H{=}6\)) as the primary but with surprise activations unmasked in its forecast — perfect demand foresight, privileged information, clearly labeled. It is a perfect-foresight peer reference, not a global upper bound (the name is historical): on announced shocks it is informationally identical to the primary and may tie or be edged by it within solver tolerance, and it never competes in the policy tables.

5.0.0.2 ForecastAwareMPC (primary).

At every period the controller observes the state, builds demand and availability forecasts from the observation window (surprise shocks are masked by the environment until onset; announced absences expose the true window, surprise absences only persistence), solves a finite-horizon mixed-integer program, applies only the first-period action, and replans. The prediction model uses hard observed qualification at \(h=0\) (so the executed action is exactly feasible), binary predicted qualification \(\theta\, c_{w,k,h} \le S_{w,k,h}\) for \(h \ge 1\), the shared time budget 2 , seat caps (hours relaxation inside the program, slot trim on the executed action), and the backlog queue encoded as \(B_{h+1} \ge B_h + \hat{D}_h - y_h\), \(B \ge 0\), which is tight at the optimum because backlog costs are strictly positive; realized backlog is recomputed by the environment’s true \(\max(0,\cdot)\) recursion. The terminal value prices qualification gaps left open at the horizon edge, \[gap_k = \max\!\bigl(0,\; \widehat{sd}_k - \textstyle\sum_w \hat{A}_{w,H-1}\, c^T_{w,k}\bigr), \qquad V_f = \lambda_{\mathrm{gap}} \sum_k gap_k , \label{eq:p3gap}\tag{6}\] with \(\widehat{sd}_k\) the per-qualification maximum of visible forecast demand from the horizon edge onward and binary terminal qualifications \(\theta\, c^T_{w,k} \le S_{w,k,H}\). The primary configuration, locked ex ante before any validation run, is \(H = 6\), \(\lambda_{\mathrm{gap}} = 1.5\); the \(\lambda \in \{0, 1.5, 5\}\) and horizon sweeps are sensitivity analyses, never best-of selection. Attribution ablations mirror the chain used in our companion controller study: ServiceOnlyMPC (receding-horizon service allocation without training variables) and MaintenanceMPC (training restricted to currently held qualifications, which therefore can maintain but never acquire or recover a qualification), so the full controller’s value decomposes into service planning, qualification maintenance, and acquisition.

5.0.0.3 Solver and diagnostics.

Each replanning step solves with scipy.optimize.milp (HiGHS branch and bound) under a one-second limit and a \(10^{-2}\) relative gap — replanning corrects residual suboptimality — with binaries reduced by fixing \(c = 1\) for qualification cells that cannot decay below threshold within the horizon. Every solve logs status and wall-clock time; a failed solve would fall back to WaterFillingTraining and be counted, and fallbacks are zero in all accepted runs.

6 Experimental Setup↩︎

6.0.0.1 Protocol.

All experiments are synthetic and seed-controlled on the frozen instance of Section 4; no real institutional or student data are used anywhere. The primary controller configuration (ForecastAwareMPC, \(H{=}6\), \(\lambda_{\mathrm{gap}}{=}1.5\)) was locked ex ante; horizon and \(\lambda\) variations are sensitivity analyses. Twenty paired seeds per validation cell (the oracle reference uses eight, each a perfect-foresight receding-horizon \(H{=}6\) solve, not a full-horizon program); a smoke suite at \(T{=}20\) gates the implementation with eight checks including exact feasibility, deterministic replay across all scenario–policy pairs, announced-vs-surprise distinguishability over multiple seeds, structural failure of ServiceOnly on the new-category scenarios, and a no-pre-onset-leakage regression for surprise shocks.

6.0.0.2 Experiment families.

(i) The six core scenarios against the seven-policy set; (ii) the attribution chain (ServiceOnlyMPC, MaintenanceMPC) on no-shock, announced and surprise new-category, and announced absence; (iii) the \(\lambda\) ablation on the new-category scenarios; (iv) a slack sweep (demand \(10.16/8.0/5.33\) hours per category, i.e.slack ratios \(1.05/1.33/2.0\)) on the surprise new category; (v) a training-speed sweep (\(\alpha \in \{0.04, 0.08, 0.16\}\)) on the announced new category, including short-horizon \(\lambda\) cells in the slow regime where the training lag approaches the horizon; (vi) the category-count sweep (\(G \in \{3, 6, 9\}\) at constant slack) probing where lean static insurance stops scaling; (vii) the oracle reference on three shocked scenarios; (viii) a frozen adversarial threat-closure suite (\(2{,}620\) episodes) that adds a shock-focused static comparator and grids training speed \(\alpha \in \{0.02, 0.04, 0.08, 0.16\}\) against disruption duration \(\in \{2, 4, 8, 12, 20\}\) under surprise, plus low-slack/slow-training cells; (ix) a post-hoc deadline/perishability re-scoring of the executed episodes at \(L \in \{1, 2, 4, \infty\}\); and (x) a frozen boundary search (\(4{,}420\) episodes) that locates where pre-positioning beats reaction, gridding training speed \(\alpha \in \{0.01, 0.02, 0.04, 0.08\}\) against disruption duration with cold and warm shock starts, a static backup-count sweep, the perfect-foresight Oracle peer in every cell, a genuine perishable-dynamics variant (perishable_env, deadline \(L\), lost penalty \(c_{\mathrm{lost}}\)), and a structural-insufficiency sweep (\(\rho \in \{1.33, 0.89, 0.67, 0.53\}\)). Families (viii)–(x) are reported separately under experiments/ (adversarial_validation_*, deadline_backlog_*, boundary_search_*) and are never pooled with the core tables.

6.0.0.3 Statistics and reproducibility.

Policy comparisons are paired by seed: per-seed win rates with exact two-sided sign tests, and separately a seeded paired bootstrap (10,000 resamples) confidence interval on the mean cost difference with relative effect sizes. The two criteria are not interchangeable; borderline results are described as mean-effect evidence rather than decisive win-rate evidence. Episodes replay deterministically: each run records a checksum over the backlog/qualification trajectory, and two independent executions of the full suite produce byte-identical artifacts once wall-clock solver columns are excluded. Every policy row reports solver status, mean solve time, fallback counts (zero throughout), and the environment’s repair/eligibility diagnostics (zero throughout).

7 Results↩︎

We report results by mechanism and regime, and the headline is a regime map rather than a “method wins” claim. In the reaction-feasible regime that includes the core instance — where a newly required qualification can be acquired within the controller’s planning horizon and recoverable window — the closed-loop controller dominates, and we establish this is a legitimate consequence of problem structure (the controller coincides with its perfect-foresight peer and the static baselines are lean and functional) rather than weak comparators. We then locate the boundary: when the training lag exceeds the controller’s reach, lean static insurance wins; a reactive trainer that starts too late wastes effort; and under structural capacity insufficiency no policy choice matters.

7.1 Implementation and reproducibility↩︎

The validation suite runs 1,544 episodes across 79 cells. Every policy is exactly feasible: zero environment repairs and zero qualification-zeroed service hours across all rows (including the oracle reference). The MILP controllers run a node-limited branch-and-bound made deterministic by single-threaded execution (set via OMP_NUM_THREADS=1; the HiGHS threads option is also forwarded, though the SciPy wrapper reports it as unrecognized); determinism is verified, not merely asserted, by a replay-checksum spot-check over the hardest MPC cells, which passes here and in both the adversarial and deadline suites. No NaNs or negative states occur.

7.2 A perfect-foresight reference and the value of foresight↩︎

Table 1 compares the primary controller against a perfect-foresight reference: the same receding-horizon controller with surprise activations unmasked in its forecast. We use it to isolate the value of foresight, not as a global optimum — it is a peer controller, not an upper bound, and on announced shocks it is informationally identical to the primary (which already sees the announced shock), so the two coincide to within solver tie-breaking tolerance (\(\pm1\)\(2\%\), including a case where the primary edges it). The informative cell is the surprise shock, where the primary is genuinely information-limited: foresight is worth a \(24\%\) cost reduction there. The legitimacy of the dominance results below therefore does not rest on this reference; it rests on the baseline-fairness and adversarial-robustness evidence of Sections 7.3 and 7.5.

Table 1: Primary vs.the perfect-foresight reference (mean cost; 8 seeds,the reference’s budget). On announced shocks the two coincide within solvertolerance (foresight adds nothing once the shock is announced); the surprisegap isolates the value of foresight.
Scenario Primary Perfect-foresight gap
announced new category 335 339 \(-1\%\)
surprise new category 419 339 \(+24\%\)
announced absence 670 657 \(+2\%\)

7.3 Core regimes↩︎

Table 2 reports mean cost across the six core scenarios for the lean, functional baselines and the primary controller. Three patterns hold. First, the service-only policy collapses wherever qualification binds: on the new-category scenarios no staff is initially qualified, so its shock coverage is zero and cost is an order of magnitude above the controller. Second, every training-capable policy improves on service-only, and the static-insurance plans are genuinely lean here — they pre-qualify backups for the latent shock category (shock coverage \(0.55\)\(0.56\), \(6\)\(11\) greenfield qualifications, \(53\)\(95\) training hours) rather than the self-defeating blanket over-training of a misconfigured plan. Third, the closed-loop controller nonetheless wins every regime, including surprise shocks, low slack, slow training, and the category-count sweep.

Table 2: Core scenarios (mean cost, 20 seeds). ST10 = StaticInsurance (10%backup target), WF = WaterFillingTraining. All comparisons Primary vs.eachbaseline are 20–0 on paired seeds (sign \(p = 1.9\times10^{-6}\)), withbootstrap CIs excluding zero.
Scenario ServiceOnly ReactiveGap ST10 WF Primary
no_shock 3,333 345 2,450 650 164
announced new category 9,319 1,635 3,602 1,741 335
surprise new category 9,319 1,635 3,602 1,845 419
announced absence 5,510 2,209 3,628 2,137 670
surprise absence 5,510 2,209 3,628 2,137 699
demand surge 8,600 7,240 6,599 5,678 3,035

7.4 Attribution: where the value comes from↩︎

The ablation chain (Table 3) decomposes the controller’s advantage on the announced new-category shock, holding the receding-horizon service allocation constant. Replacing the heuristic service split with the MILP allocation (ServiceOnly \(\to\) ServiceOnlyMPC) accounts for a \(27\%\) reduction; adding qualification maintenance against decay (ServiceOnlyMPC \(\to\) MaintenanceMPC) roughly halves cost again; and enabling acquisition of the new qualification just in time (MaintenanceMPC \(\to\) Primary) removes \(91\%\) of the remaining gap. The dominant mechanism is therefore targeted just-in-time qualification, not merely better allocation — which is precisely the capability static insurance approximates but pays for in advance.

Table 3: Attribution chain on the announced new-category shock (mean cost,20 seeds), holding service allocation constant across the MPC variants.
Policy cost mechanism added
ServiceOnly 9,319 — (heuristic service, no training)
ServiceOnlyMPC 6,793 MILP service allocation
MaintenanceMPC 3,770 + qualification maintenance
Primary 335 + just-in-time acquisition

7.5 A regime map: where reaction beats pre-positioning, and where it does not↩︎

The companion manufacturing studies find a regime split: lean static insurance wins under surprise shocks near the capacity boundary, where a reaction transient is structurally unrecoverable. The scientific question is not whether the controller always wins here, but where the boundary lies. The governing quantity is whether a newly required qualification can be acquired within the controller’s reaction reach — its planning horizon \(H{=}6\) plus the window over which backlog stays recoverable. We map both sides of that boundary and treat the map, not a dominance claim, as the finding.

7.5.0.1 Reaction-feasible regime (the controller wins).

In the core instance the disruption window is long (\(20\) periods) relative to a short training lag (\(1\)\(2\) periods), and non-storable backlog is recoverable over such a window, so the controller reacts within the disruption and the surprise/announced distinction that static insurance hedges largely dissolves. The sweeps confirm the gap narrows but does not close: at low slack the relative margin shrinks (Primary \(4{,}115\) vs.ST10 \(9{,}700\), a \(2.4\times\) gap versus \(8\times\) at high slack), and at the slow training rate the controller still wins \(20\)\(0\). To stress this side of the boundary we ran a dedicated, frozen adversarial threat-closure suite (adversarial_validation_*; \(2{,}620\) episodes, zero fallbacks, zero repairs, deterministic spot-check passing) attacking the result along the four axes a reviewer would press — short disruption windows, slow training, low slack, and a stronger static comparator. The comparator set adds ShockFocusedStatic, a legitimate exact-feasible plan that concentrates its entire slack-metered, seat-capped backup budget on the latent shock qualification from \(t=0\) (it pre-commits to the shock slot but never learns its timing) — one of several static variants spanned by a backup-count sweep. Across all \(158\) paired Primary-vs-comparator cells, no static plan beats the controller. Three representative cells: (i) the shock-focused plan under surprise \(+\) low slack \(+\) slow training (\(\alpha{=}0.04\)) loses \(20\)\(0\) (Primary \(4{,}159\) vs. \(14{,}254\) — concentrating the premium on the guessed slot starves current service); (ii) a two-period window with the training lag deliberately made longer than the window (\(\alpha{=}0.02\), lag \(\approx 3\) periods), the canonical case where pre-positioning should win, loses \(8\)\(0\) (Primary \(263\) vs.StaticInsurance10 \(2{,}737\) — the controller absorbs and serves down a short backlog rather than pre-training for a shock that barely materializes); and (iii) a grid over training speed \(\alpha \in \{0.02, 0.04, 0.08, 0.16\}\) and disruption duration \(\in \{2, 4, 8, 12, 20\}\) in which the controller wins all \(20\) cells. These cells share warm starts and reskilling rates at which the training lag fits inside the controller’s reach, so they all sit in the reaction-feasible regime.

7.5.0.2 Pre-positioning regime (static insurance wins).

To find the other side of the boundary we ran a second frozen suite (boundary_search_*; \(4{,}420\) episodes, zero fallbacks, zero repairs, determinism spot-check passing) that pushes the training lag past the controller’s horizon: a cold shock start (no prior preparation) and deep retraining (\(\alpha{=}0.01\), so a new qualification costs \(\approx\!60\) staff-hours — a lag of \(\approx\!7\)\(8\) periods), with ample slack so pre-positioning is affordable. Here a lean static plan (StaticInsurance\(10\)) wins once the window is long enough to accumulate costly backlog: it beats the controller \(0\)\(20\) (sign \(p{=}1.9\times10^{-6}\)) at every window \(\ge\!8\) periods, by up to \(2\times\) at duration \(20\) (Primary \(3{,}570\) vs. \(1{,}725\)). The win is structural, not a baseline or information artifact: the perfect-foresight Oracle — the same \(H{=}6\) controller with the shock unmasked — ties the primary exactly (\(0\)\(0\)\(20\), mean difference \(0.0\)) and loses to static identically, because no \(H{=}6\) controller can accumulate \(60\) training-hours within a six-period horizon. At \(\alpha\ge0.02\) (lag inside the horizon) the ordering flips back: foresight pays (the Oracle beats the primary) and the primary beats static. The boundary is therefore set by the training lag relative to the controller’s reach (Figure 1); a static commitment from \(t=0\) — which needs the shock slot but never its timing — is the only policy that pre-qualifies in time when that lag is long. The heavier static plans over-train and lose: lean pre-positioning is what wins.

Table 4: Regime map. Which wins — just-in-time reaction (the controller) oradvance pre-positioning (static insurance) — is governed by thenew-qualification training lag relative to the controller’s reaction reach(horizon \(H{=}6\) plus the recoverable/perishable window). All four regimes arereproduced in frozen suites.
Regime Condition Best policy
Reaction-feasible lag \(\le\) reach (warm start; reskilling \(\alpha\ge0.02\); or short window) controller
Pre-positioning lag \(>\) reach (cold start, deep retrain \(\alpha{=}0.01\)) and long window (\(\ge8\)) lean static insurance
Reaction-too-late reactive trainer starts after onset; cannot qualify in time none (wasted training)
Structurally insufficient demand \(>\) qualified capacity (\(\rho<1\)) none (all collapse)
Figure 1: The reaction-versus-pre-positioning boundary (boundary_search_*). Left: winner over training rate \alpha (slow at top) and disruption duration, backlog-only with a cold start; the static-favoring corner (slow training, long window) is where the qualification lag outruns the controller’s reach. Right: at \alpha{=}0.01 the controller and its perfect-foresight peer coincide exactly (the lag exceeds the H{=}6 horizon, so foresight cannot help), while lean static insurance, flat in duration, overtakes both once the window is long enough to amortize its premium.

7.5.0.3 Reaction-too-late and structurally-insufficient regimes.

Two further regimes complete the map. A purely reactive trainer (ReactiveGap), which trains only after the gap appears, is the worst policy in the pre-positioning regime: at \(\alpha{=}0.01\), duration \(20\) it burns \(1{,}100\) training-hours after onset, too late to qualify (shock coverage \(0.01\)), and underperforms even no-training (\(8{,}404\) vs.ServiceOnly \(6{,}877\)) — training that starts too late is worse than not training. And when demand structurally outruns total qualified capacity (\(\rho<1\)), every policy collapses to uniformly poor coverage and the choice stops mattering: as \(\rho\) falls from \(1.33\) to \(0.53\), mean coverage drops from \(0.88\) to \(0.12\) and the worst comparator’s relative gap over the controller shrinks from \(+2{,}118\%\) to \(+20\%\). The terminal qualification-gap penalty has a small effect at this fast training rate (Primary vs.\(\lambda{=}0\): \(15\)\(5\) on the announced shock, \(13\)\(7\) on surprise, mean effect \(14\%\) on announced-shock cost), consistent with just-in-time training being near-optimal when the lag is short; we report it as a sensitivity rather than a decisive lever and do not pool it across rates.

7.6 Backlog perishability moves the boundary↩︎

The backlog-only dynamics treat uncovered support as recoverable later, which a reviewer can fairly attack: perhaps the controller “wins” only by serving late what an institution cannot ethically or operationally defer. We test this with a frozen deadline re-scoring (deadline_backlog_*): each executed trajectory is re-evaluated under a deadline \(L\), counting support served more than \(L\) periods after it arrived as late/lost rather than recovered. This is an explicit post-hoc lens, not a re-optimization — policies are not re-planned against the deadline — so it isolates whether the advantage depends on late service; \(L{=}\infty\) reproduces the backlog-only result exactly. It does not depend on late service: the controller loses the fewest hours at every deadline \(L \in \{1, 2, 4, \infty\}\) across all four tested scenarios, \(20\)\(0\) paired in every cell. Even at the strictest deadline \(L{=}1\) on the announced new-category shock, the controller leaves \(20\) hours undelivered against \(540\)\(680\) for the static plans (on-time coverage \(0.99\) vs.\(0.69\)\(0.75\)). In the reaction-feasible regime the controller’s lead therefore does not rest on forgiving backlog recovery. But perishability is not neutral at the boundary. We also re-ran the boundary suite under genuine perishable dynamics (perishable_env: support unserved past a deadline \(L\) is permanently lost at penalty \(c_{\mathrm{lost}}{=}20\), \(L\in\{2,4\}\)). A hard deadline is double-edged: it penalizes the controller’s deferred shock service — which, in the pre-positioning regime, it never delivers — but also penalizes the static plan’s starvation of base categories while it pre-trains. The net effect moves the boundary without erasing either regime: at the cold, deep-retraining corner the static-win region persists but contracts to the longest windows — static wins at duration \(20\) (Primary \(5{,}330\) vs.\(4{,}436\) at \(L{=}2\); \(5{,}721\) vs.\(3{,}794\) at \(L{=}4\), both \(0\)\(20\)) while the controller reclaims the shorter eight-period windows. A deadline-aware controller is the natural response in that regime and is left to future work.

Figures 23 summarize the core costs, the category-count sweep, and the training-rate sweep.

Figure 2: Core scenarios: mean cost by policy (log scale, 20 seeds).

a

b

Figure 3: Category-count sweep (left) and training-rate sweep (right) on the surprise/announced new-category shocks..

8 Discussion and Limitations↩︎

8.0.0.1 Reading the results carefully.

The benchmark’s value is the regimes it makes measurable. A service-only policy collapses whenever qualification is the binding resource — structurally on new support categories (nobody is qualified; eligibility is hard) and gradually under decay, where unmaintained qualifications lapse mid-horizon. The notable finding is a regime map, not a dominance claim. In the reaction-feasible regime — where a new qualification can be acquired within the controller’s horizon and recoverable window, including the core instance — the closed-loop controller dominates, and we take pains to establish this is not a weak-baseline artifact: the static-insurance plans are lean, monotone in their budget, and exactly feasible, they pre-qualify the shock category, and a frozen adversarial suite (adversarial_validation_*) adding a shock-focused static plan and sweeping disruption windows down to two periods, training rates across a \(4\times\) range, and low slack finds no static win across \(158\) paired cells; a deadline re-scoring (deadline_backlog_*) shows that lead does not depend on backlog being recoverable. The perfect-foresight reference adds nothing over the primary on announced shocks — it is a peer, not an upper bound — and isolates a \(24\%\) value of foresight only under surprise. But we deliberately sought the other side of the boundary and found it: a frozen boundary search (boundary_search_*) pushing the training lag past the controller’s horizon (cold start, deep retraining) shows lean static insurance winning by up to \(2\times\) on long windows. That win is structural — the perfect-foresight peer (\(H{=}6\)) ties the primary exactly and loses identically, because no six-period controller can acquire a qualification that takes longer than six periods to learn. This is exactly the manufacturing-studies regime (pre-bought capacity wins when reaction is structurally too late), reached here through a long training lag rather than an unrecoverable capacity transient. Two further regimes complete the map: a reactive trainer that starts after onset wastes effort and is worst of all, and under structural capacity insufficiency (\(\rho<1\)) coverage collapses for every policy and the choice stops mattering. Forecast privilege is made explicit by construction: every policy declares its information set, surprise variants mask the forecast until onset, and the oracle is labeled and never competes.

8.0.0.2 What the attribution chain establishes.

The ServiceOnlyMPC \(\to\) MaintenanceMPC \(\to\) full-controller chain holds receding-horizon service allocation constant and varies only training eligibility, separating three mechanisms with direct metric support from the within-episode counters: qualification maintenance against decay, requalification of lapsed cells (invisible to terminal counts by construction), and greenfield acquisition for categories nobody initially serves. The same counters discipline the narrative: a policy that wins on cost while acquiring zero new qualifications is winning on allocation, not capability.

8.0.0.3 Access dispersion, carefully.

The Jain indices over per-category coverage and per-staff training hours, with their minimum floors, expose concentration: policies that achieve average coverage by systematically under-serving one category, or that concentrate training on few staff members. These are operational dispersion metrics over synthetic categories and staff groups. No demographic attributes exist anywhere in the model, and no fairness or equity claims are made or implied.

8.0.0.4 Limitations.

Everything here is stylized computational evidence. The skill model is a two-parameter abstraction (linear gain, geometric decay, hard threshold); demand categories are synthetic and student-free; the instance is small (eight staff, up to nine categories, one institution); absences and surges are simple windows; forecast modes are masks over scenario means rather than learned forecasters. Excluded by design: timetabling, rooms, transportation, student-level assignment, protected attributes, compliance logic, learning outcomes, hiring, and procurement. The MILP controller runs with a deterministic node budget; its solutions are good incumbents, not proven optima, and solver scalability beyond this instance is untested. The static-insurance plans are informed comparators (they know which qualification slots exist, though never whether demand will arrive), and the oracle reference is a perfect-foresight peer (a receding-horizon \(H{=}6\) incumbent within a time budget), not an upper bound. None of the policies constitutes deployment advice: the framework identifies which institutional data — real demand categories, training durations, decay rates, absence patterns — would be needed before any real-world use.

8.0.0.5 Outlook.

The natural next steps follow the benchmark discipline: richer qualification structures (many-to-many category–skill maps where blanket insurance cannot scale), stochastic or scenario-based controllers that price hidden shocks, calibration of demand shapes and training durations to public education statistics [48], and user studies of the Studio interface with educational operations leaders — each an extension of the testbed, not a claim the current paper makes.

9 Conclusion↩︎

We introduced a benchmark specification and decision-support framework for qualified educational capacity planning: a stylized, single-institution service system in which support demand is heterogeneous by category, services are not storable, staff qualification is a dynamic state with hard threshold eligibility and decay, and training is a control action that consumes the same staff hours current service needs. The benchmark ships six seed-controlled scenario families with announced and surprise information regimes, an exact-feasibility discipline with declared per-policy information sets, within-episode requalification and greenfield counters, access-dispersion metrics, replay checksums, and paired statistics; a policy suite spanning service-only, reactive, static- insurance, water-filling, and rolling-horizon MILP control with an attribution chain and a labeled oracle reference; and EduCapacity Studio, a scenario-analysis interface whose export/replay loop reproduces results exactly.

The empirical picture is a regime map, governed by whether a newly required qualification can be acquired within the controller’s reaction reach. In the reaction-feasible regime — which includes the core instance and a frozen \(2{,}620\)-episode adversarial suite (a shock-focused static plan, windows down to two periods, a \(4\times\) training-rate range, low slack) — the closed-loop controller dominates over all \(158\) paired cells, and we establish this is legitimate rather than a weak-baseline artifact. But when the training lag is pushed past the controller’s horizon (a cold start with deep retraining), a \(4{,}420\)-episode boundary search finds lean static insurance winning by up to \(2\times\) on long windows; the win is structural, since the perfect-foresight peer ties the controller exactly and loses identically. A reactive trainer that starts too late is worst of all, and under structural capacity insufficiency no policy choice matters; backlog perishability shifts the boundary without erasing either regime. This recovers — rather than contradicts — the manufacturing-studies regime split: pre-positioning wins where reaction is structurally too late, reached here through a long training lag relative to the controller’s horizon. The benchmark’s contribution is to make these regimes — and the data one would need to locate a real institution within them — measurable, reproducible, and inspectable, without claiming validated educational impact, compliance, or student-level recommendations. Extensions in qualification structure, uncertainty-aware control, public-data calibration, and interface evaluation are left as future work on top of the released specification.

Funding↩︎

This research received no external funding.

References↩︎

[1]
E. Bettini, T. D. Nguyen, A. F. Gilmour, and C. Redding, “Disparities in access to well-qualified, well-supported special educators across higher- versus lower-poverty schools over time,” Exceptional Children, vol. 88, no. 3, pp. 283–301, 2022, doi: 10.1177/00144029211024137.
[2]
B. Billingsley and E. Bettini, “Special education teacher attrition and retention: A review of the literature,” Review of Educational Research, vol. 89, no. 5, pp. 697–744, 2019, doi: 10.3102/0034654319862495.
[3]
U.S. Department of Education, IDEA regulations: Sec. 300.156 personnel qualifications.” 2026, [Online]. Available: https://sites.ed.gov/idea/regs/b/b/300.156.
[4]
American Speech-Language-Hearing Association, “A workload analysis approach for establishing speech-language caseload standards in the school: Position statement.” 2002, [Online]. Available: https://www.asha.org/policy/ps2002-00122/.
[5]
C. H. Carlin, “Workload versus caseload: An exploratory comparison study of individualized education program progress and other outcomes,” Language, Speech, and Hearing Services in Schools, vol. 55, no. 2, pp. 259–275, 2024, doi: 10.1044/2023_LSHSS-23-00075.
[6]
K. J. Vannest and S. Hagan-Burke, “Teacher time use in special education,” Remedial and Special Education, vol. 31, no. 2, pp. 126–142, 2010, doi: 10.1177/0741932508327459.
[7]
SEATS, “Special education automated teacher scheduling.” 2026, [Online]. Available: https://www.seatscheduling.com/about.
[8]
DMSchedules, “Special education scheduling software.” 2026, [Online]. Available: https://dmschedules.com/special-education-scheduling-software.
[9]
ParaFlowTool, SPED para scheduling software.” 2026, [Online]. Available: https://www.paraflowtool.com/.
[10]
L. Feng and T. R. Sass, “What makes special-education teachers special? Teacher training and achievement of students with disabilities,” Economics of Education Review, vol. 36, pp. 122–134, 2013, doi: 10.1016/j.econedurev.2013.06.006.
[11]
M. E. Brock and E. W. Carter, “Effects of a professional development package to prepare special education paraprofessionals to implement evidence-based practice,” The Journal of Special Education, vol. 49, no. 1, pp. 39–51, 2015, doi: 10.1177/0022466913501882.
[12]
B. McCollum et al., “Setting the research agenda in automated timetabling: The second international timetabling competition,” INFORMS Journal on Computing, vol. 22, no. 1, pp. 120–130, 2010, doi: 10.1287/ijoc.1090.0320.
[13]
G. Post et al., XHSTT: An XML archive for high school timetabling problems in different countries,” Annals of Operations Research, vol. 218, no. 1, pp. 295–301, 2014, doi: 10.1007/s10479-011-1012-2.
[14]
S. Ceschia, L. Di Gaspero, and A. Schaerf, “Educational timetabling: Problems, benchmarks, and state-of-the-art results,” European Journal of Operational Research, vol. 308, no. 1, pp. 1–18, 2023, doi: 10.1016/j.ejor.2022.07.011.
[15]
J. Miranda, P. A. Rey, and J. M. Robles, “udpSkeduler: A web architecture based decision support system for course and classroom scheduling,” Decision Support Systems, vol. 52, no. 2, pp. 505–513, 2012, doi: 10.1016/j.dss.2011.10.011.
[16]
M. D. Bailey and D. Michaels, “An optimization-based DSS for student-to-teacher assignment: Classroom heterogeneity and teacher performance measures,” Decision Support Systems, vol. 119, pp. 60–71, 2019, doi: 10.1016/j.dss.2019.02.006.
[17]
Ö. Aygül, T. Hellgren, S. Azizi, and A. C. Trapp, “A predict-and-prescribe framework for dynamic course scheduling toward strategic university scaling,” Omega, vol. 138, p. 103406, 2026, doi: 10.1016/j.omega.2025.103406.
[18]
A. T. Ernst, H. Jiang, M. Krishnamoorthy, and D. Sier, “Staff scheduling and rostering: A review of applications, methods and models,” European Journal of Operational Research, vol. 153, no. 1, pp. 3–27, 2004, doi: 10.1016/S0377-2217(03)00095-X.
[19]
J. Van den Bergh, J. Beliën, P. De Bruecker, E. Demeulemeester, and L. De Boeck, “Personnel scheduling: A literature review,” European Journal of Operational Research, vol. 226, no. 3, pp. 367–385, 2013, doi: 10.1016/j.ejor.2012.11.029.
[20]
O. Czibula, H. Gu, A. Russell, and Y. Zinder, “A multi-stage IP-based heuristic for class timetabling and trainer rostering,” Annals of Operations Research, vol. 252, no. 2, pp. 305–333, 2017, doi: 10.1007/s10479-015-2090-3.
[21]
B. A. Schwendimann et al., “Perceiving learning at a glance: A systematic literature review of learning dashboard research,” IEEE Transactions on Learning Technologies, vol. 10, no. 1, pp. 30–41, 2017, doi: 10.1109/TLT.2016.2599522.
[22]
R. Kaliisa, K. Misiejuk, S. López-Pernas, M. Khalil, and M. Saqr, “Have learning analytics dashboards lived up to the hype? A systematic review of impact on students’ achievement, motivation, participation and attitude,” in Proceedings of the 14th learning analytics and knowledge conference, 2024, pp. 295–304, doi: 10.1145/3636555.3636884.
[23]
G. Post, L. Di Gaspero, J. H. Kingston, B. McCollum, and A. Schaerf, “The third international timetabling competition,” Annals of Operations Research, vol. 239, no. 1, pp. 69–75, 2016, doi: 10.1007/s10479-013-1340-5.
[24]
T. Müller, H. Rudová, and Z. Müllerová, “Real-world university course timetabling at the international timetabling competition 2019,” Journal of Scheduling, vol. 28, no. 2, pp. 247–267, 2025, doi: 10.1007/s10951-023-00801-w.
[25]
A. Bonutti, F. De Cesco, L. Di Gaspero, and A. Schaerf, “Benchmarking curriculum-based course timetabling: Formulations, data formats, instances, validation, visualization, and results,” Annals of Operations Research, vol. 194, no. 1, pp. 59–70, 2012, doi: 10.1007/s10479-010-0707-0.
[26]
G. H. G. Fonseca, H. G. Santos, E. G. Carrano, and T. J. R. Stidsen, “Integer programming techniques for educational timetabling,” European Journal of Operational Research, vol. 262, no. 1, pp. 28–39, 2017, doi: 10.1016/j.ejor.2017.03.020.
[27]
D. S. Holm, R. O. Mikkelsen, M. Sørensen, and T. J. R. Stidsen, “A graph-based MIP formulation of the international timetabling competition 2019,” Journal of Scheduling, vol. 25, no. 4, pp. 405–428, 2022, doi: 10.1007/s10951-022-00724-y.
[28]
N. Pillay, “A survey of school timetabling research,” Annals of Operations Research, vol. 218, no. 1, pp. 261–293, 2014, doi: 10.1007/s10479-013-1321-8.
[29]
J. S. Tan, S. L. Goh, G. Kendall, and N. R. Sabar, “A survey of the state-of-the-art of optimisation methodologies in school timetabling problems,” Expert Systems with Applications, vol. 165, p. 113943, 2021, doi: 10.1016/j.eswa.2020.113943.
[30]
A. W. Siddiqui, S. A. Raza, and Z. M. Tariq, “A web-based group decision support system for academic term preparation,” Decision Support Systems, vol. 114, pp. 1–17, 2018, doi: 10.1016/j.dss.2018.08.005.
[31]
T. H. Hultberg and D. M. Cardoso, “The teacher assignment problem: A special case of the fixed charge transportation problem,” European Journal of Operational Research, vol. 101, no. 3, pp. 463–473, 1997, doi: 10.1016/S0377-2217(96)00082-3.
[32]
B. Domenech and A. Lusa, “A MILP model for the teacher assignment problem considering teachers’ preferences,” European Journal of Operational Research, vol. 249, no. 3, pp. 1153–1160, 2016, doi: 10.1016/j.ejor.2015.08.057.
[33]
J. Johnes, “Operational research in education,” European Journal of Operational Research, vol. 243, no. 3, pp. 683–696, 2015, doi: 10.1016/j.ejor.2014.10.043.
[34]
M. J. Davis, Y. Lu, M. Sharma, M. S. Squillante, and B. Zhang, “Stochastic optimization models for workforce planning, operations, and risk management,” Service Science, vol. 10, no. 1, pp. 40–57, 2018, doi: 10.1287/serv.2017.0199.
[35]
I. Senthooran, P. Le Bodic, and P. J. Stuckey, “Optimising training for service delivery,” Leibniz International Proceedings in Informatics, vol. 210, pp. 48:1–48:15, 2021, doi: 10.4230/LIPIcs.CP.2021.48.
[36]
A. F. Gilmour, “Teacher certification area and the academic outcomes of students with learning disabilities or emotional/behavioral disorders,” The Journal of Special Education, vol. 54, no. 1, pp. 40–50, 2020, doi: 10.1177/0022466919849905.
[37]
J. J. Kirksey and M. Lloydhauser, “Dual certification in special and elementary education and associated benefits for students with disabilities and their teachers,” AERA Open, vol. 8, 2022, doi: 10.1177/23328584211071096.
[38]
S. L. Woulfin and B. Jones, “Special development: The nature, content, and structure of special education teachers’ professional learning opportunities,” Teaching and Teacher Education, vol. 100, p. 103277, 2021, doi: 10.1016/j.tate.2021.103277.
[39]
T. L. Fisher, P. T. Sindelar, D. Kramer, and E. Bettini, “Are paraprofessionals being hired to replace special educators? A study of paraprofessional employment,” Exceptional Children, vol. 88, no. 3, pp. 302–315, 2022, doi: 10.1177/00144029211062595.
[40]
W. Matcha, N. A. Uzir, D. Gašević, and A. Pardo, “A systematic review of empirical studies on learning analytics dashboards: A self-regulated learning perspective,” IEEE Transactions on Learning Technologies, vol. 13, no. 2, pp. 226–245, 2020, doi: 10.1109/TLT.2019.2916802.
[41]
O. Viberg, M. Hatakka, O. Bälter, and A. Mavroudi, “The current landscape of learning analytics in higher education,” Computers in Human Behavior, vol. 89, pp. 98–110, 2018, doi: 10.1016/j.chb.2018.07.027.
[42]
L. Paulsen and E. Lindsay, “Learning analytics dashboards are increasingly becoming about learning and not just analytics: A systematic review,” Education and Information Technologies, 2024, doi: 10.1007/s10639-023-12401-4.
[43]
S. R. Vemula and M. Moraes, “Learning analytics dashboards for advisors – a systematic literature review.” 2024, [Online]. Available: https://arxiv.org/abs/2402.01671.
[44]
A. F. Wise and Y. Jung, “Teaching with analytics: Towards a situated model of instructional decision-making,” Journal of Learning Analytics, vol. 6, no. 2, 2019, doi: 10.18608/jla.2019.62.4.
[45]
A. Larrabee Sønderlund, E. Hughes, and J. Smith, “The efficacy of learning analytics interventions in higher education: A systematic review,” British Journal of Educational Technology, vol. 50, no. 5, pp. 2594–2618, 2019, doi: 10.1111/bjet.12720.
[46]
R. Kaliisa, K. Misiejuk, S. López-Pernas, M. Khalil, and M. Saqr, “Have learning analytics dashboards lived up to the hype? A systematic review of impact on students’ achievement, motivation, participation and attitude.” 2023, [Online]. Available: https://arxiv.org/abs/2312.15042.
[47]
R. Kaliisa, I. Jivet, and P. Prinsloo, “A checklist to guide the planning, designing, implementation, and evaluation of learning analytics dashboards,” International Journal of Educational Technology in Higher Education, vol. 20, no. 1, 2023, doi: 10.1186/s41239-023-00394-6.
[48]
OECD, “Education at a glance 2025: OECD indicators.” 2025, doi: 10.1787/1c0d9c79-en.