July 04, 2026
Bangladesh has an estimated 1.17 mental-health professionals per 100,000 population and only six child psychiatrists nationwide. No Bengali-language, culturally adapted tool exists for early screening of abuse-related psychological trauma in children. We present ShishuRaksha AI, a decision-support (not diagnostic) framework that fuses four screening modalities: validated questionnaires (SDQ, CPSS), Bengali narrative text, House-Tree-Person (HTP) drawing features, and facial affect. The fusion is training-free and clinically weighted, uses cross-modal attention, and includes a single-modality override rule. Every risk score is explained through clinically weighted, perturbation-based additive attribution and rendered as a bilingual (Bangla/English) report with referral routing to national child-protection services (OCC, DSS, NMHH) under the Children Act 2013. No clinical dataset of abused children can be collected ethically at this stage, so we introduce a noise-aware synthetic benchmark (500 cases, 116 positive [23.2%], four deliberate noise layers, literature-grounded HTP priors) and evaluate tree-ensemble surrogates of the fusion design (facial channel excluded) under 5-fold stratified cross-validation. The fused model reaches an AUC of 0.874 [0.834–0.908], against 0.756 [0.705–0.803] for an SDQ-only baseline, with ablation, operating-point, subgroup, and calibration analyses. We state all limitations openly, including synthetic-only data, no held-out set, text-feature circularity, and an urban–rural subgroup gap. This work is a feasibility study and a design contribution toward ethically deployable child-protection screening in low-resource settings.
child protection, psychological trauma screening, multimodal fusion, explainable AI, Bengali NLP, synthetic benchmark, decision support, Bangladesh
A child in a remote haor district such as Sunamganj who stops speaking after a violent episode at home will, in most cases, never see a specialist. Bangladesh has an estimated 1.17 mental-health professionals per 100,000 population and, per WHO figures, only six child psychiatrists in the entire country [1]. The epidemiology is not reassuring: community surveys using the Bangla SDQ have found psychiatric disorder in substantial fractions of 5–10-year-olds across rural, urban, and slum settings alike [2]. Abuse-related trauma sits in the worst corner of this gap. It is stigmatized and rarely disclosed, and outside a handful of urban clinics it is almost never screened. The instruments that do exist are paper questionnaires that need trained administrators, written in English or translated without any automation.
Machine learning cannot close this gap, and we do not claim it can. What it might do is triage: help a school counselor, an NGO field worker, or a shishu welfare officer decide which children need referral to the scarce professionals who exist. That framing shapes every design choice here. ShishuRaksha AI1 produces a prioritization signal with a human-readable justification, in Bangla, routed to Bangladesh’s real child-protection bodies: One-Stop Crisis Centres (OCC), the Department of Social Services (DSS), and the national mental-health helpline, under the Children Act 2013 [3]. It is decision support; it does not diagnose, and it is built so that it cannot quietly pretend to.
There is an obvious ethical problem with building such a system: training or validating on records of abused children cannot be justified at the proof-of-concept stage, and no ethics board was asked to permit it. We take a synthetic-first path instead. A noise-aware generator produces a benchmark cohort grounded in published HTP and screening literature, and the framework is evaluated on it with the full set of checks a real validation would require: ablation, operating points, subgroups, and calibration. The numbers are feasibility evidence, nothing more; Section 7 is deliberately the longest discussion in the paper.
The contributions: a Bangladesh-specific four-modality screening framework; training-free attention fusion with clinically specified weights and a single-modality override; bilingual explainable reporting wired to statutory referral routing; a noise-aware synthetic benchmark methodology for domains where real data collection is ethically infeasible; and a feasibility evaluation of the whole design.
The Strengths and Difficulties Questionnaire (SDQ) [4] is the most widely deployed brief screening instrument for child mental health, with well-established psychometric properties [5]. A Bangla adaptation was validated by Mullick and Goodman on 261 Bangladeshi children, distinguishing clinic from community samples and between diagnostic groups [6], the SDQ multi-informant algorithm has been evaluated in Dhaka clinics alongside London [7], and Bangla-SDQ epidemiology has documented substantial prevalence of child psychiatric disorder across rural, urban, and slum communities in Bangladesh [2]. For trauma specifically, the Child PTSD Symptom Scale (CPSS) [8] is a standard self-report measure. All of these instruments, however, assume access to trained administrators and, in Bangladesh, confront a system with only six child psychiatrists nationwide [1]; none provides automated triage support in Bengali.
The House-Tree-Person (HTP) test [9] and related drawing techniques [10] have long been used with children precisely because they demand little verbal ability. The evidence behind them is mixed: Allen and Tussey’s systematic review found that projective drawings could not reliably detect whether a child had experienced sexual or physical abuse [11], and Lin et al.showed that deep neural networks trained on 4,196 children’s HTP drawings failed to predict depression, casting doubt on HTP as a stand-alone diagnostic signal [12]. At the same time, computational drawing analysis continues to advance: object-detection pipelines now generate structured HTP score tables automatically [13], machine learning over digitized drawings has been used to identify psychological trauma among Syrian refugee children for early intervention [14], and recent computer-vision systems automate feature extraction from children’s drawings for screening referral [15]. We take the adverse validity evidence seriously as a design constraint rather than ignoring it: in ShishuRaksha AI the drawing modality is (i) one weak signal among four, (ii) down-weighted by clinically specified weights, and (iii) used only for screening prioritization, never diagnosis.
Machine learning has been applied across mental-health detection, diagnosis support, and public-health surveillance [16], and multimodal approaches (for example, fusing audio and text of clinical interviews for depression detection [17]) generally outperform unimodal baselines. Existing multimodal systems, however, target adults, assume high-resource languages, and rely on learned fusion whose parameters resist clinical audit. To our knowledge, no prior system combines child-appropriate modalities for a Bengali-language, Bangladesh-specific screening context.
Additive feature-attribution methods such as SHAP [18] and LIME [19] dominate post-hoc explanation, and studies of clinician needs emphasize that explanations must map onto actionable, domain-meaningful factors rather than raw features [20]; explainability is increasingly framed as a precondition for the responsible clinical use of AI [21]. Rudin [22] further argues that high-stakes decisions call for models that are interpretable by design rather than post-hoc explanations of black boxes. Our training-free, clinically weighted fusion follows that principle. Prior clinical XAI work assumes English-language reporting and established referral infrastructure; none produces bilingual (Bangla/English) explanations routed to a specific national child-protection pathway.
Gap. No existing work combines (a) child-appropriate multimodal screening, (b) training-free, clinically auditable fusion, (c) Bengali-language explainable reporting, and (d) referral routing grounded in Bangladesh’s child-protection institutions (OCC, DSS, NMHH; Children Act 2013 [3]). ShishuRaksha AI is designed to fill this gap as a decision-support framework, with the synthetic-benchmark methodology of Section IV providing a feasibility evaluation in the absence of ethically obtainable clinical data.
Four input channels feed the system. Validated questionnaires carry the largest clinical weight, \(w_q{=}0.40\): the SDQ [4] and a DSM-5-adapted CPSS [8], with subscale scores mapped into a 32-slot feature vector. They sit highest because they are the best-tested instruments of the four. A Bengali free-text narrative, encoded with BanglaBERT [23] (768-d CLS embedding), carries \(w_t{=}0.25\). The HTP drawing channel (\(w_d{=}0.20\)) concatenates 1,280 EfficientNet-B0 features [24] with 20 binary HTP markers (omitted figures, dark shading, heavy line pressure, encapsulation, aggressive imagery, and similar literature-derived indicators). Facial affect is the fourth, weakest-weighted channel (\(w_f{=}0.15\)), specified for interview settings but excluded from this paper’s evaluation. Fixing these weights costs accuracy, potentially a great deal of it. What it buys is auditability: a psychologist can read the weights, disagree with them, and change them. No learned attention matrix allows that.
Each modality vector (32-d questionnaire, 768-d text, 1,300-d drawing, 32-d facial) passes through a seeded random projection into a shared 128-d space; the projection matrices are generated from per-modality deterministic seeds, so the entire mapping is reproducible from the codebase alone. Scaled dot-product attention is then computed across the four projected representations, a residual connection and tanh nonlinearity are applied, and the four attended 128-d vectors are concatenated into the 512-d fused representation. No parameters are learned anywhere in this path; every weight is either clinically specified or seed-derived, which makes the fusion fully auditable. A learned variant (fine-tuned encoders, 8-head fusion) is specified in the codebase but neither implemented nor evaluated here.
The fused score maps to four tiers: LOW \([0,0.25)\), MODERATE \([0.25,0.50)\), HIGH \([0.50,0.75)\), and CRITICAL \([0.75,1]\). Each tier binds to concrete actions and a referral target drawn from Bangladesh’s existing infrastructure: MODERATE routes to a school counsellor, HIGH to the DSS national child helpline (1098) within 24 hours, CRITICAL to a One-Stop Crisis Centre (hotline 16767) immediately. One safeguard sits above the fusion: if any single modality’s score exceeds 0.85, the case is raised to at least HIGH regardless of the composite. A frightening drawing should not be averaged away by a calm questionnaire. Whether 0.85 is the right trigger is an open clinical question; the mechanism is the point.
Every risk score is decomposed by clinically weighted, perturbation-based additive attribution in the spirit of SHAP [18], rendered as a waterfall: which modality pushed the score up, which pulled it down, and by how much. The choice of a perturbation scheme over exact Shapley computation trades theoretical guarantees for speed on commodity hardware, a trade we consider acceptable at the screening tier; a worked example appears in Fig. 5.
The output is not a number. It is a report generated in Bangla and English side by side, stating the risk tier, the attribution waterfall in plain language, the tier-specific action list, and the exact referral contact with hotline number (1098 for DSS, 16767 for OCC), consistent with the reporting duties set by the Children Act 2013 [3]. The report ends with a fixed disclaimer, in both languages, that the tool does not diagnose.
Collecting drawings and narratives from abused children to validate an unproven prototype is not an option we considered. The benchmark is generated instead, from a single fixed seed (42). Questionnaire scores are sampled per class from SDQ and CPSS subscale distributions consistent with published cutoffs (SDQ total difficulties 0–40, CPSS DSM-5 subscales); drawing cases are 20-dimensional binary HTP marker vectors with class-conditional marker prevalences taken from the HTP literature, plus a continuous marker-burden score; Bengali narratives are template-generated in trauma and non-trauma variants.
Clean synthetic data would prove nothing, so the generator deliberately corrupts its own output through four layers. Questionnaire scores receive \(\pm\)15% measurement error, 10% reporter misclassification, and 5% missing subscales. Drawing markers suffer 20% detection errors with \(\pm\)0.1 jitter on the burden score. Narratives flip trauma\(\to\)non-trauma in 15% of cases (modeling partial disclosure) and the reverse in 10%. Labels themselves carry 8% inter-rater disagreement flips. Each layer models a failure mode we expect in the field; whether the four together approximate real-world messiness is untestable until real data exist.
Labels are drawn at a 20% target prevalence and then subjected to the 8% inter-rater noise layer, yielding a realized cohort of 500 cases with 116 positives (23.2%). No real children and no clinical data were used at any stage. One circularity must be disclosed plainly: three of the five text features used in evaluation (emotional density, trauma keywords, disclosure score) are constructed as noisy label-correlated proxies rather than extracted from the generated narratives, so the text channel’s measured contribution is optimistic by construction. Results involving that modality should be read with this in mind.
The training-free fusion engine produces a fixed score, not a fitted probability, so discrimination is evaluated through tree-ensemble surrogates over the same modality features. The fused configuration and the text-only configuration use gradient-boosted classifiers [25] (200 and default estimators, respectively); the SDQ-only and drawing-only baselines use random forests (100 trees); all use median imputation, a fixed random state of 42, and the fused configuration additionally receives the clinically-weighted composite score as a feature. Stratified 5-fold cross-validation over the 500 cases yields out-of-fold predictions for every case, and all reported metrics are computed on the pooled out-of-fold predictions. There is no held-out test set. With 500 cases, we judged that setting one aside would add more statistical noise than rigor, but readers should weigh the cross-validated numbers with that in mind. Confidence intervals are 1,000-resample bootstraps. The facial modality is excluded from evaluation entirely: no facial component was implemented, and we prefer reporting nothing over reporting a placeholder.
| Model | AUC | 95% CI |
|---|---|---|
| SDQ only (baseline) | 0.756 | [0.705, 0.803] |
| Text only | 0.768 | [0.716, 0.815] |
| Drawing only | 0.774 | [0.718, 0.827] |
| ShishuRaksha AI (fusion) | 0.874 | [0.834, 0.908] |
Fig. 2 and Table 1 tell a consistent story. The three unimodal configurations land close together, between 0.756 and 0.774, and none clearly beats the SDQ-only baseline. Fusion changes that: 0.874, with a confidence interval whose lower bound sits above every unimodal point estimate. On synthetic data, this pattern hints that the modalities carry partly independent signal, which is what the clinical-weighting design assumes. It cannot show that real children’s data would behave the same way. The fused curve stays above the others across most of the false-positive range, with the largest margin below FPR \(\approx\) 0.3, exactly the region a screening deployment would occupy.
Ablation (Fig. 3a) removes one modality at a time from the fused configuration. Dropping the text channel costs the most (\(-\)0.099, to 0.775). Removing the questionnaire costs \(-\)0.052, and removing the drawing channel only \(-\)0.030. We are careful not to celebrate the text result. It is exactly the modality carrying the label-proxy circularity disclosed in Section 4, so its large ablation drop is partly an artifact of how its features were built. The honest reading is that the questionnaire and drawing channels each add modest, plausibly real signal, while the text channel’s true value will only be known once features are extracted from real narratives.
A weight-sensitivity check adds robustness and a caveat at once. Re-running the fused evaluation with equal weights, swapped questionnaire and text weights, a drawing-heavy setting, and with the composite feature removed entirely yields AUCs of 0.865–0.875. The headline number does not depend on the chosen clinical weights; it also means the weights add little measured discrimination, and their value lies in the auditability of the deployed engine, not here.
At the MODERATE tier boundary (0.25), the fused model reaches sensitivity 0.707 at specificity 0.878 (PPV 0.636, NPV 0.908). This is the setting a broad screening pass would use: it accepts more false alarms in order to miss fewer children. The HIGH boundary (0.50) trades down to sensitivity 0.603 at specificity 0.930 (PPV 0.722), and the CRITICAL boundary (0.75) reaches specificity 0.956 at sensitivity 0.457. With only 116 positive cases, per-cell counts behind these estimates are small and threshold-level numbers should be treated as unstable; the same caveat applies to the four-tier confusion matrix, which is additionally degenerate because the synthetic ground truth is binary (LOW/HIGH only) while predictions span all four tiers.
Fig. 3b shows fused-model AUC across age, gender, and location subgroups. Age and gender bands stay within roughly 0.05 of the overall estimate with overlapping intervals. Location does not: urban cases (Dhaka/Chittagong divisions, \(n{=}119\)) score 0.780 [0.648–0.885] against 0.902 [0.863–0.938] for the rest of the country. The intervals overlap, so with this sample size the gap is suggestive rather than established. Still, a 0.122 point difference in the direction of worse urban performance is exactly the kind of gap that would mean unequal screening quality in deployment, and we flag it as a first-order item for any real-data validation.
Discrimination is only half the requirement; a score that overstates risk weakens screener trust. Fig. 4 shows acceptable global calibration (ECE \(=0.074\)) with one consistent failure: in the highest bins the model predicts risk near 0.85–0.97 while observed prevalence sits near 0.65–0.83. Raw scores must therefore not be read as probabilities; thresholds should be set on ranked scores, or the output recalibrated before any probability is shown to a human. Mid-range bins hold only 9–26 cases each, so we avoid reading meaning into them.
Fig. 5 shows the attribution output for a synthetic test case (age 12, risk score 0.618, tier HIGH). The CPSS PTSD total dominates, followed by the composite risk score and the questionnaire modality as a whole; drawing and facial channels contribute least. This ordering is what the clinical weights would predict: reassuring for auditability, and a reminder that the explanation reflects the design’s assumptions as much as the case’s evidence. The report renders this waterfall with Bangla-first labels and the DSS 1098 referral line.
Everything in this paper rests on synthetic data. No result says anything direct about real Bangladeshi children; a system at AUC 0.874 on a generated cohort could sit near chance on a clinical one. The evaluation also measures tree-ensemble surrogates rather than the training-free fusion engine itself; the surrogates consume the same features, but the two are not the same computation, and the gap between them is unquantified. There is no held-out test set. The text modality inherits a circularity from the generation process, as disclosed in Section IV, and its contribution should be discounted accordingly. The facial modality exists only on paper. Calibration degrades exactly where it matters most, at high predicted risk. And subgroup performance is uneven, with urban cases scoring 0.122 AUC below rural ones (Fig. 3b), a gap that would mean unequal screening quality across communities if deployed as-is.
Some of these problems have a clear path. Prospective validation under proper ethical review would replace the synthetic evidence or refute it; the likely route is a partnership with a child-protection NGO and clinical co-investigators, with assent and guardian-consent protocols designed before any case is collected. Earlier still, Bangladeshi child psychologists scoring a sample of synthetic cases would test whether the benchmark’s notion of risk matches clinical intuition. Recalibration, an evaluated facial channel, and the specified deep extension all wait on data that does not yet ethically exist. We consider that ordering correct; collecting data first and asking ethics questions later is how this domain gets systems that should never have been built.
ShishuRaksha AI is a design and a feasibility argument, not a validated screening tool. On a noise-aware synthetic benchmark, fusing four child-appropriate modalities under clinically specified weights outperformed every unimodal alternative by a margin large enough to justify the next, harder step. The framework’s value may lie less in its numbers than in its scaffolding: auditable fusion, bilingual explanation, and statutory referral routing. That scaffolding survives even if the specific weights do not. Whether any of this holds on real children is precisely the question a synthetic study cannot answer, and the one we now intend to ask properly.
This study uses exclusively synthetic data; no real children, clinical records, or personally identifiable information were involved. ShishuRaksha AI is designed strictly as a decision-support tool for trained professionals and must not be used for diagnosis. The full codebase, the synthetic benchmark, and the evaluation pipeline that produced every number in this paper are available at https://github.com/shkoli/shishuraksha-ai.
Code and the full synthetic benchmark: https://github.com/shkoli/shishuraksha-ai↩︎