Variational Mixture of Graph Neural Experts for Alzheimer’s Disease Recognition across Frequency Bands in EEG Brain Networks

Jun-En Ding1, Anna Zilverstand2, Shihao Yang1, Albert Chih-Chieh Yang3, Feng Liu1,1
1Department of Industrial and Systems Engineering, Rutgers University, Piscataway, NJ, USA
2Department of Psychiatry and Behavioral Sciences, University of Minnesota, Minneapolis, MN, USA
3Institute of Brain Science, College of Medicine, National Yang-Ming Chiao Tung University, Taipei City, Taiwan


Abstract

Dementia disorders such as Alzheimer’s disease (AD) and frontotemporal dementia (FTD) exhibit overlapping electrophysiological signatures in EEG that challenge accurate diagnosis. Existing EEG-based methods are limited by full-band frequency analysis, which hinders precise differentiation of dementia subtypes and severity stages. To address this limitation, we propose a Variational Mixture of Graph Neural Experts (VMoGE) framework that integrates multi-band EEG analysis with variational graph neural networks and a mixture-of-experts architecture. Each expert specializes in a specific EEG frequency band and models brain connectivity using a Gaussian Markov Random Field prior, while a variational gating mechanism adaptively integrates expert outputs. This design enables the model to learn frequency-specific brain network representations while modeling latent uncertainty through variational inference. Experimental results on two EEG dementia datasets show that VMoGE achieves strong performance, with an AUC of 0.89 for HC vs. AD classification in the main comparison and competitive results across dementia subtyping and CDR staging tasks. Clinically, VMoGE offers three key translational values: the expert gating weights correlate with MMSE scores and CDR severity, slow-wave δ/θ-band contributions are associated with AD-related EEG slowing and disease progression, and spatially localized activation maps reveal posterior \(\theta\)/\(\alpha\)-band alterations and region-specific \(\beta\)-band changes, providing neurophysiologically interpretable markers aligned with known AD neuropathology.

Note to PractitionersThe clinical diagnosis of dementia often relies on physicians’ experience-based interpretation of electroencephalography (EEG) waveforms and clinical indicators. Many existing studies lack a comprehensive evaluation that simultaneously considers model performance on both EEG biomarkers and clinical indicators for Alzheimer’s disease assessment. This study proposes the variational mixture of graph neural experts (VMoGE), a mixture of graph experts integrating multi-band feature analysis with variational inference. VMoGE automatically learns brain network activity patterns across different frequency bands and visualizes disease-associated brain regions and spectral features. Additionally, VMoGE provides automated dementia biomarker identification, with each expert module corresponding to specific frequency bands. Weight variations correlate with clinical scales (e.g., MMSE), age, and disease progression, thereby enhancing the efficiency and objectivity of detecting early-stage AD biomarker changes and identifying high-risk cases.

Alzheimer’s disease, EEG, Mixture of experts, Graph neural networks, Variational inference

1 Introduction↩︎

Alzheimer’s disease (AD) represents the most prevalent form of dementia worldwide, affecting an estimated 6.9 million individuals aged 65 years and above in the United States as of 2024 [1], compared to around 0.35 million individuals in the same age group in Taiwan [2]. Specifically, frontotemporal dementia (FTD) is a major cause of young-onset dementia, typically presenting between ages 35 and 75. It is marked by early behavioral and language changes, including disinhibition, emotional blunting, stereotyped behaviors, and dietary alterations, reflecting selective degeneration of the frontal and temporal lobes [3], [4].

Traditional diagnostic approaches rely heavily on neuropsychological assessments (e.g., MMSE, MoCA), structural neuroimaging (MRI/CT), and invasive biomarker analysis (CSF Aβ42/tau ratios, amyloid PET) [5], [6]. While these methods provide valuable diagnostic information, they face critical limitations that restrict widespread implementation. Neuropsychological tests such as MMSE exhibit ceiling effects in early-stage dementia and are susceptible to practice effects [7][9]. Neuroimaging biomarkers require expensive equipment and specialized expertise that remain largely inaccessible in resource-constrained settings [10][12]. Moreover, the invasive nature of procedures such as lumbar puncture for CSF collection reduces patient compliance and precludes their routine use in longitudinal monitoring, particularly among elderly populations who may face increased procedural risks [13], [14].

The Electroencephalography (EEG) serves as the most mainstream and cost-effective measurement method, effectively reflecting cortical activities through brain signals. Several studies have demonstrated that AD is associated with disrupted brain connectivity, reduced neural flexibility, and abnormal EEG state transitions, reflecting impaired large-scale network coordination and functional degeneration [15][17]. With the increasing availability of physiological signals, recent advances in time series embedding methods further highlight that the choice of representation strategy, ranging from Fourier transforms to Transformer-based architectures, significantly impacts downstream classification performance [18]. Moreover, multimodal biosignal monitoring systems such as wearable belts combining ECG, sEMG, and EDA for assessing visceral pain in irritable bowel syndrome [19] underscore the clinical value of integrating heterogeneous physiological signals. Clinical trials that leverage objective physiological measures alongside validated scales, as demonstrated in acupuncture efficacy studies for dialysis patients [20], [21], further reinforce the importance of combining multiple assessment modalities [19].

Recent research findings indicate that analyzing time-frequency differences across bands can effectively serve as biomarkers for analysis. However, many studies show controversial results in FTD band analysis. For example, both AD and FTD typically exhibit spectral slowing, characterized by increased delta/theta power and reduced \(\alpha\)/\(\beta\) power, which correlates with cognitive decline [22], [23]. Early-onset AD often demonstrates widespread delta enhancement, whereas FTD shows more region-specific abnormalities and distinct microstate alterations [22]. Power ratios such as theta/alpha and theta/beta, along with gamma band changes, have been proposed as sensitive markers for prodromal and clinical AD [24].

The success of deep learning (DL) has greatly enhanced feature extraction from multi-band EEG signals, such as convolutional neural networks (CNNs), long short-term memory networks (LSTMs), and hybrid architectures, which can automatically learn spatiotemporal features across frequency bands [25], [26]. Specifically, Transformer architectures have demonstrated significant potential for EEG-based AD classification [27][29]. The field has progressed to convolutional-Transformer hybrid architectures such as CEEDNet and DICE-Net [30] to large-scale foundation models with cross-dataset pretraining, with models like LEAD achieving classification accuracies exceeding 93% and establishing Transformers as a pivotal direction in EEG-based dementia diagnosis [31].

Building on these advancements, and unlike traditional EEG feature extraction approaches, more recent research has explored graph neural networks (GNNs), which have emerged as a powerful architecture for modeling EEG brain connectivity and capturing complex spatial, spectral, and temporal dependencies. Furthermore, multi-path, multi-frequency, and attention-based GNN architectures demonstrate superior performance in distinguishing dementia-related disorders, offering promising tools for accurate and interpretable diagnosis [32][36].

Research on GNNs integrated with variational inference has demonstrated strong potential in handling graph data with inherent uncertainty and noise. To further enrich representation learning, the MoVGAE integrates first- and high-order neighborhood information, achieving superior results on classification, clustering, and link prediction tasks [37]. Extending to spatiotemporal settings, the DVGNN leverages diffusion processes and variational inference to uncover dynamic causal relationships with enhanced accuracy and interpretability [38]. Although these approaches often consider standard Gaussian priors as prior distributions, variational approximations remain insufficient to fully capture complex distributions, particularly brain signal networks with graph structures.

Recently, Mixture of Experts (MoE) has gained significant attention due to its key advantage of dynamic task allocation, which leverages expert specialization to significantly reduce computational costs while maintaining high performance, especially in Large Language Models (LLMs) [39] or dynamics of multimodal fusion [40], [41]. Furthermore, GNN integrated with MoE has significantly improved graph and node classification accuracy [42][44]. For example, GraphDIVE employs mixture-of-experts with semantic partitioning to address imbalanced graph classification by optimizing the evidence lower bound, effectively reducing bias toward majority classes [45]. In both molecular property prediction and multivariate time series anomaly detection, GNNs provide powerful structural representations, while MoE introduces specialized expert modules dynamically routed by gating mechanisms [46], [47].

Nevertheless, although many approaches have been developed to identify potential EEG-based biomarkers for AD, they remain limited by differences in signal preprocessing procedures [48], variations in experimental design protocols across subjects [31], [49], or focus solely on deep learning models for EEG signal feature extraction, which significantly reduces the interpretability of effective biomarkers for dementia diagnosis [50][53]. Furthermore, several existing studies employ full-band frequency analysis for AD assessment [54], wang2024ADFormer?, which may result in cross-frequency interference and hinder the precise differentiation of FTD variations among AD patients. Moreover, apart from differentiating AD from FTD, various characteristics of AD necessitate evaluation through daily behavioral patterns, including the assessment of cognitive function and neuropsychiatric symptoms using clinical scales. These limitations have constrained the development of effective, straightforward, and objective biomarkers in EEG-based diagnostic applications [55].

To bridge this gap, we first develop a multi-granularity EEG feature extractor designed to learn node-level representations that capture both different-scale temporal patterns. Subsequently, we introduce a variational mixture of graph neural experts (VMoGE) where each expert is modeled under a Gaussian Markov Random Field (GMRF) prior corresponding to a distinct EEG frequency band. By embedding structured variational inference into the mixture-of-experts paradigm, our VMoGE adaptively assigns responsibilities among experts while maintaining the intrinsic graph topology and effectively representing epistemic uncertainty, which contributes to the classification and staging of dementia. The main contributions of this paper are as follows:

  • This study presents a multi-granularity transformer that captures multi-scale temporal patterns across four EEG frequency bands by integrating 1-D convolutional neural network (CNN) modules at multiple temporal granularities with FFT-based spectral features, yielding rich frequency-specific node-level representations for graph-based dementia classification.

  • We propose a variational mixture of graph neural experts where each expert is governed by a band-specific GMRF prior. Through a closed-form KL divergence formulation, VMoGE enforces neurophysiologically meaningful spatial smoothness constraints, improving generalization under limited-sample clinical conditions.

  • Through adaptive gating and evidence lower bound optimization, VMoGE dynamically assigns responsibilities to frequency-specific experts, providing interpretable insights into EEG biomarkers, including expert weights that correlate with clinical indicators (e.g., MMSE scores, age, and disease progression) and spatial patterns aligned with neuropathological signatures of AD and FTD.

The structure of this study is as follows. Section 2 reviews related work. Section 3 presents the problem formulation, MGT-NFE module, GMRF prior modeling, and variational graph convolutional encoder. Section 4 details the VMoGE framework, including the gating mechanism and ELBO optimization. Section 5 describes the experimental setup, dataset descriptions, baseline comparisons, and ablation studies. Section 6 evaluates model robustness under varying noise conditions. Section 7 provides explainable diagnosis analysis via gating weights, clinical indicators, and spatial brain activation patterns. Section 8 discusses the frequency-specific biomarkers, spatial patterns, and neurophysiological mechanisms revealed by VMoGE, as well as its potential for multimodal extension.

2 Related Work↩︎

2.1 Variational Inference Framework in EEG↩︎

Existing variational inference methods have established a robust framework for EEG analysis, providing principled solutions for representation learning, uncertainty quantification, and modeling individual differences in brain activity patterns. For example, EEG2Vec uses a conditional variational autoencoder that learns both generative and discriminative representations from affective EEG data, achieving 68.49% accuracy in emotion classification while generating synthetic EEG sequences that preserve low-frequency spectral characteristics [56]. Building on graph-based approaches, [57] proposed variational pathway reasoning (VPR), which uses random walks within brain regions to model emotion-related pathways and Bayesian processes to learn pathway importance, achieving 94.3% accuracy on the SEED dataset with interpretable visualizations. To address individual variability, the Variational Instance-Adaptive Graph (V-IAG) combines deterministic instance-adaptive graphs with probabilistic variational graphs to capture both subject-specific dependencies and underlying uncertainties [58]. Complementing these approaches, [59] explored variational predictive coding through hierarchical recurrent state-space models that balance active learning and active inference, enabling multi-level EEG prediction while maintaining biological plausibility. Moreover, the EEG graph combines a variational spatial encoder with a Gaussian temporal encoder to capture both structural brain priors and delayed cross-regional temporal dependencies [60].

2.2 Mixture-of-Experts for Brain Network Analysis↩︎

Generally, MoE architectures enable efficient neural network scaling through conditional computation, where specialized expert subnetworks process different input types via gating mechanisms. The core formulation follows \(P(y|x) = \sum_{i} g_i(x)P(y|x,\theta_i)\), where \(g_i(x)\) represents gating weights and \(P(y|x,\theta_i)\) denotes expert outputs [61]. Recent developments in MoE have established them as a robust solution for modeling the complex heterogeneity observed in brain signals from EEG [62][65] and fMRI [66], [67]. By incorporating specialized expert modules and adaptive gating strategies, these frameworks successfully accommodate the inherent subject variability, diverse task requirements, and multi-scale dynamics present in neuroimaging applications [68]. Specifically, MoE models have shown notable promise in neurophysiological signal processing. In psychiatric and neurological disorder research, cognition-aware MoE frameworks utilize functional brain atlases to capture the dynamics of cognitive regions [69], while disease-specific routing strategies incorporate specialized sub-networks to construct comprehensive whole-brain embeddings [70]. Similarly, Seizure-MoE and Mix-MoE for EEG-based seizure subtype classification effectively mitigate class imbalance and integrate manual features, while Neuro-MoBRE further demonstrates robust multi-task decoding across subjects through brain-regional expert specialization, including applications in epileptic seizure diagnosis [64], [67].

Despite the promising ability of MoE models to capture features across distinct dynamic regions of EEG brain networks, current approaches seldom integrate more fine-grained frequency-based graph structures as prior distributions. This limitation is particularly relevant in FTD research, where electrophysiological signatures differ from those observed in AD. For instance, AD is often characterized by rostral dominance in fractal complexity, whereas FTD exhibits caudal dominance, alongside distinct alterations in slow-frequency activity [71]. In this work, we demonstrate that incorporating structure-specific graph priors that reflect such frequency-band asymmetries can substantially improve the interpretability of EEG biomarkers, enabling MoE frameworks to better distinguish between FTD and AD cohorts.

Figure 1: Diagram of MGT-NFE for node feature extraction, incorporating multi-granularity hierarchical feature extraction and spatial positional encoding at different granularities.

3 Methods↩︎

Figure 2: Overview of the VMoGE framework for AD biomarker identification and prediction. VMoGE framework first extracts node features from multi-channel EEG signals (channels C1–C19) by integrating spatial and frequency band features through a 1D-CNN and FFT-based MGT-NFE module. The prior graph structure is constructed using a GMRF, and a variational router models the latent distribution to capture structural correlations across multiple frequency bands. Finally, the framework performs a weighted summation of four expert networks based on k-th gating probabilities to obtain the final classification result.

3.1 Problem Formulation↩︎

In this work, we implement the common feature-extraction pipeline in two datasets. The feature extraction process was implemented through a systematic epoching and spectral analysis pipeline. We define each instance of raw EEG timeseries as \(\boldsymbol{X}_{\text{raw}} \in \mathbb{R}^{ C \times T}\) with \(C\) channels and time step of \(T\). The EEG data were first segmented into non-overlapping epochs by calculating the epoch length \(T^{\prime}\) in samples and determining the total number of epochs that could be extracted. For each epoch, the algorithm extracted temporal segments across all channels and performed channel-wise spectral analysis. We compute the power spectral density (PSD) with Welch’s method and define the total power over 0.5–45 Hz. We then compute relative band power (RBP) in four non-overlapping bands \(\delta\) (0.5–4 Hz), \(\theta\) (4–8 Hz), \(\alpha\) (8–13 Hz), and \(\beta\) (13–45 Hz), using left-closed, right-open intervals so that bands partition the total power.

Given the extracted feature set \(\{(\mathbf{X}_i, y_i)\}_{i=1}^{N}\), where \(\mathbf{X}_i \in \mathbb{R}^{K \times C \times F}\) denotes the RBP feature tensor for the \(i\)-th sample, composed of \(K = 4\) frequency-band feature matrices \(\mathbf{X}_i^{(k)} \in \mathbb{R}^{C \times F}\) across \(C = 19\) channels, with \(F\) being the length of the RBP-derived feature sequence, and \(y_i \in \{0, 1\}\) is the corresponding binary class label.

3.2 Multi-Granularity Transformer for Node Feature Extraction (MGT-NFE)↩︎

Aiming to capture multi-scale temporal patterns in RBP-derived EEG features while extracting meaningful node representations for each channel, the MGFormer  [72] architecture was designed to more effectively capture multi-level features encompassing both local and long-range temporal signals. As shown in Fig. 1 of the MGT-NFE components, these include self-attention to capture long-range temporal dependencies, a one-dimensional convolutional neural network (1-D CNN) to extract local temporal details, and Fast Fourier Transform (FFT) features.

3.2.1 Multi-Granularity Token Encoder↩︎

For each channel–band pair \((c, k)\) with \(k \in \{\delta, \theta, \alpha, \beta\}\), the corresponding RBP feature sequence \(\mathbf{x}_{c,k} \in \mathbb{R}^{F}\) is defined as the \(c\)-th row vector of \(\mathbf{X}_i^{(k)}\). Multi-scale temporal patterns are then extracted from this RBP feature sequence using convolutional layers with kernel sizes tailored to different temporal granularities. Let \(\mathcal{G}\) denote the set of temporal granularity indices, and each granularity \(g \in \mathcal{G}\) is associated with a one-dimensional convolutional kernel of length \(\Phi(g)\), a stride \(s(g)\), and a pooling operator \(P_g(\cdot)\). Given the RBP feature sequence \(\mathbf{x}_{c,k} \in \mathbb{R}^{F}\), the temporal token at granularity \(g\) is defined as \[\mathbf{T}^{(c,k)}_{g} = P_g\!\Big( \sigma\!\big( \mathrm{Conv1D} ( \mathbf{x}_{c,k};\,\Phi(g),\,s(g) ) \big) \Big) \in \mathbb{R}^{L_g\times D_T},\] where \(\sigma(\cdot)\) is a nonlinear activation function, such as LeakyReLU, \(D_T\) is the number of convolutional filters, and \(L_g\) is the resulting token length after convolution, stride, and pooling. By concatenating the tokens across all granularities, we obtain \[\mathbf{T}^{(c,k)} = \mathrm{Concat}_{g\in\mathcal{G}} \left[ \mathbf{T}^{(c,k)}_{g} \right] \in \mathbb{R}^{L\times D_T}, \qquad L=\sum_{g\in\mathcal{G}} L_g .\]

3.2.2 Transformer Encoding↩︎

To preserve temporal ordering, we add a positional encoding \(\mathbf{PE}\in\mathbb{R}^{L\times D_T}\) to the concatenated tokens before feeding them into the Transformer layers: \[\begin{align} \mathbf{E}^{(0)} &= \mathbf{T}^{(c,k)} + \mathbf{PE},\\ \mathbf{E}^{(l+1)} &= \mathrm{TransformerBlock}\left(\mathbf{E}^{(l)}\right), \quad l=0,1,\ldots,L_{\text{tr}}-1, \end{align}\] where \(L_{\text{tr}}\) denotes the total number of Transformer layers. Each \(\mathrm{TransformerBlock}(\cdot)\) consists of multi-head self-attention and position-wise feed-forward networks, producing outputs of the same shape \(\mathbb{R}^{L\times D_T}\).

After \(L_{\mathrm{tr}}\) transformer layers, we obtain the final encoded features \(\mathbf{E}^{(L_{\mathrm{tr}},c,k)}\) for each channel–band pair. We then extract a representative per-channel, per-band node feature from the final transformer output: \[\mathbf{h}_{c}^{(k)} = \operatorname{Aggregate}\!\left(\mathbf{E}^{(L_{\mathrm{tr}},c,k)}\right) \in \mathbb{R}^{d_h},\] where \(\operatorname{Aggregate}(\cdot)\) can be mean pooling, max pooling, or learned attention pooling, and \(d_h\) denotes the per-band node feature dimension.

Applying the MGT-NFE independently to each frequency band and stacking the resulting per-channel features yields the band-specific node feature matrix: \[\label{eq:per95band95H} \mathbf{H}^{(k)} = [\mathbf{h}_{1}^{(k)}; \mathbf{h}_{2}^{(k)}; \ldots; \mathbf{h}_{C}^{(k)}] = \operatorname{MGT\text{-}NFE}(\mathbf{X}^{(k)}_{i}),\tag{1}\] where \(\mathbf{H}^{(k)} \in \mathbb{R}^{C \times d_{h}}\) represents the node feature matrix for frequency band \(k\). To further obtain a unified cross-band descriptor, we concatenate the band-wise features of each channel: \[\label{eq:concat95h} \mathbf{h}_{c} = \operatorname{Concat}[\mathbf{h}_{c}^{(\delta)}, \mathbf{h}_{c}^{(\theta)}, \mathbf{h}_{c}^{(\alpha)}, \mathbf{h}_{c}^{(\beta)}],\tag{2}\] where \(\mathbf{h}_c \in \mathbb{R}^{D^{\prime}}\) with \(D^{\prime}=4d_h\) being the total node feature dimension. The stacking over channels gives the unified node feature matrix \(\mathbf{H} = [\mathbf{h}_{1}; \mathbf{h}_{2}; \ldots; \mathbf{h}_{C}] \in \mathbb{R}^{C \times D'}\).

3.3 GMRF Prior Modeling↩︎

In this section, we formulate the extracted node features \(\mathbf{H}^{(k)}\) for each band \(k\) as a corresponding band graph \(G^{(k)}=\{(\mathcal{E}^{(k)},\mathcal{V}^{(k)}), \mathbf{A}^{(k)}\}\) with vertex \(\mathcal{V}^{(k)} = \{v_{i}\}_{i=1}^{C}\) and adjacency matrix \(\mathbf{A}^{(k)} \in \mathbb{R}^{C \times C}\). As shown in Fig. 2, our goal is to establish a gating network from a variational perspective as the most suitable MoE weighting mechanism. Existing EEG research has demonstrated that the characteristic differences across different frequency bands in early AD can serve as clinical biomarkers for diagnostic interpretation [23], [73]. To characterize and encode different band graph features, we introduce latent variables \(\mathbf{Z}^{(k)} \in \mathbb{R}^{C \times d_{z}}\) for each band \(k\), where \(d_z\) is the latent dimension. We denote by \(\mathbf{z}^{(k)}_{d} \in \mathbb{R}^{C}\) the \(d\)-th column of \(\mathbf{Z}^{(k)}\), which collects the values of all \(C\) nodes along the \(d\)-th latent dimension (\(d = 1, \dots, d_z\)). We model each latent band graph as having a prior distribution that follows a Gaussian Markov Random Field (GMRF). The precision matrix \(\mathbf{Q}^{(k)} \in \mathbb{R}^{C \times C}\) encodes the graph structure and satisfies the Markov property, has edges \((i, j)^{(k)} \in \mathcal{E}^{(k)} \Longleftrightarrow Q_{ij}^{(k)} \neq 0\) for all \(i \neq j\). In practice, the precision matrix \(\mathbf{Q}^{(k)}\) is derived from the symmetric normalized graph Laplacian \(\mathbf{L}\) with an added diagonal loading term:

\[\label{eq:precision95matrix} \mathbf{Q}^{(k)} = \mathbf{L}^{(k)} + \lambda_{Q}\mathbf{I}, \qquad \mathbf{L}^{(k)} = \mathbf{I} - \mathbf{D}^{(k)^{-1/2}}\mathbf{A}^{(k)}\mathbf{D}^{(k)^{-1/2}},\tag{3}\] where \(\mathbf{I} \in \mathbb{R}^{C \times C}\) is the identity matrix and positive semi-definite, \(\mathbf{D}^{(k)} \in \mathbb{R}^{C \times C}\) is the degree matrix of \(\mathbf{A}^{(k)}\), and \(\lambda_{Q} > 0\) is a small diagonal loading constant. The term \(\lambda_{Q}\mathbf{I}\) ensures that \(\mathbf{Q}^{(k)}\) is positive definite, so that its inverse \((\mathbf{Q}^{(k)})^{-1}\) exists and \(\log|\mathbf{Q}^{(k)}|\) is well defined. The structured GMRF prior is then applied independently to each latent dimension:

\[\label{eq:gmrf95prior} p(\mathbf{Z}^{(k)} \mid \mathbf{A}^{(k)}) = \prod_{d=1}^{d_z}\mathcal{N}(\mathbf{z}^{(k)}_{d}; \mathbf{0}, (\mathbf{Q}^{(k)})^{-1}),\tag{4}\] in which the prior distribution encodes the smoothness assumption of the graph structure, where two nodes strongly connected in \(\mathbf{A}^{(k)}\) are encouraged to have similar latent representations in \(\mathbf{Z}^{(k)}\), enabling the model to leverage EEG brain connectivity patterns specific to each frequency band.

3.4 Variational Graph Convolutional Encoder↩︎

In this work, we design a simple two-layer graph convolutional network (GCN) to encode the node features \(\mathbf{H}^{(k)}\) from the MGT-NFE of the \(k\)-th frequency band into latent representations. Specifically, given the adjacency matrix \(\mathbf{A}^{(k)}\) and node feature matrix \(\mathbf{H}^{(k)}\), the output of the GCN encoder is computed as:

\[\tilde{\mathbf{H}}^{(k)}= \text{GCN}^{(k)}(\mathbf{H}^{(k)}, \mathbf{\hat{A}}^{(k)})\] where \(\hat{\mathbf{A}}^{(k)} = \mathbf{D}^{(k)^{-1/2}} \mathbf{A}^{(k)} \mathbf{D}^{(k)^{-1/2}}\) is the symmetrically normalized adjacency matrix. These GCN embeddings \(\tilde{\mathbf{H}}^{(k)}=[\tilde{\mathbf{h}}^{(k)}_{1},\tilde{\mathbf{h}}^{(k)}_{2},...,\tilde{\mathbf{h}}^{(k)}_{C}]\) serve as inputs to a node-level variational encoder, which approximates the posterior distribution over latent variables \(\mathbf{Z}^{(k)} = [\mathbf{z}_1^{(k)}, \dots, \mathbf{z}_C^{(k)}]^{\top}\), where \(\mathbf{z}_i^{(k)} \in \mathbb{R}^{d_z}\) denotes the latent vector of node \(i\). The graph encoder produces a mean-field Gaussian variational posterior that is factorized over nodes:

\[\label{eq:posterior95node} q_{\phi}(\mathbf{Z}^{(k)} \mid \mathbf{H}^{(k)}, \mathbf{A}^{(k)}) = \prod_{i=1}^{C} \mathcal{N} \left( \mathbf{z}_i^{(k)} \mid \boldsymbol{\mu}_i^{(k)}, \operatorname{diag} \left( (\boldsymbol{\sigma}_i^{(k)})^2 \right) \right),\tag{5}\] where the mean \(\boldsymbol{\mu}_i^{(k)} \in \mathbb{R}^{d_z}\) and log standard deviation \(\log \boldsymbol{\sigma}_i^{(k)} \in \mathbb{R}^{d_z}\) are computed from the node embedding \(\tilde{\boldsymbol{h}}^{(k)}_{i}\) of the \(k\)-th expert via simple MLP mappings:

\[\boldsymbol{\mu}_i^{(k)} = \mathbf{W}_\mu \, \tilde{\mathbf{h}}_i^{(k)}, \quad \log \boldsymbol{\sigma}_i^{(k)} = \mathbf{W}_{\log\sigma} \, \tilde{\mathbf{h}}_i^{(k)},\] where \(\mathbf{W}_{\mu}, \mathbf{W}_{\log\sigma} \in \mathbb{R}^{d_z \times H_{\text{gcn}}}\) project the GCN output \(\tilde{\mathbf{h}}_{i}^{(k)} \in \mathbb{R}^{H_{\text{gcn}}}\) to the latent dimension \(d_z\).

To enable gradient-based optimization through stochastic sampling, we apply the reparameterization trick [74] to obtain:

\[\mathbf{z}_i^{(k)} = \boldsymbol{\mu}_i^{(k)} + \boldsymbol{\sigma}_i^{(k)} \odot \boldsymbol{\epsilon}_i, \quad \boldsymbol{\epsilon}_i \sim \mathcal{N}(\mathbf{0}, \mathbf{I}),\] where \(\odot\) denotes element-wise multiplication. This guarantees the sampling process to be differentiable and supports efficient backpropagation during training. The resulting latent representations \(\mathbf{z}_i^{(k)}\) are then used for downstream tasks, such as graph-level classification through mean pooling.

4 Variational Mixture of Graph-Structured Experts↩︎

4.0.1 Graph-Level Representation and Expert Decoder↩︎

For classification, we aggregate the node-level latent representations using mean pooling: \[\bar{\mathbf{z}}^{(k)} = \frac{1}{C} \sum_{i=1}^{C} \mathbf{z}_i^{(k)} \in \mathbb{R}^{d_z}\] where \(\bar{\mathbf{z}}^{(k)}\) is then passed through an expert-specific decoder network to produce the \(k\)-th expert’s logit output: \[\hat{y}^{(k)} = \text{MLP}(\bar{\mathbf{z}}^{(k)}; \boldsymbol{\theta}_k)\] where \(\text{MLP}\) is a multi-layer perceptron and \(\boldsymbol{\theta}_k\) represents the learnable parameters of the \(k\)-th expert decoder.

4.0.2 Mixture of Experts with Gating Network↩︎

In contrast to conventional fixed mixture models [75], [76], we propose an adaptive gating mechanism grounded in graph structural priors within the MoE framework, which enables each input sample to dynamically modulate the relative contribution of individual frequency-band experts based on the underlying graph-prior information. Formally, the gating coefficients for the \(k\)-th expert are obtained through a softmax operation applied over all experts, and can be formulated as:

\[\pi^{(k)}(\mathbf{H}')= \frac{ \exp\left( \mathbf{w}_k^{\top} f_{\boldsymbol{\psi}}(\mathbf{H}') \right) }{ \sum_{j=1}^{K} \exp\left( \mathbf{w}_j^{\top} f_{\boldsymbol{\psi}}(\mathbf{H}') \right) }.\] where \(f_{\boldsymbol{\psi}}(\mathbf{H}')\) denotes a gating feature mapping applied to the concatenated input \(\mathbf{H}^{\prime} = \text{Concat}([\mathbf{H}^{(1)}, \mathbf{H}^{(2)}, \ldots, \mathbf{H}^{(K)}])\), and \(\mathbf{w}_k\) is the parameter vector associated with the \(k\)-th expert. Then, the final prediction is computed as a weighted combination of all expert outputs:

\[\label{eq:mixture95output} \hat{y} = \sum_{k=1}^{K} \pi^{(k)}(\mathbf{H}^{\prime}) \cdot \hat{y}^{(k)}\tag{6}\] where the softmax ensures \(\sum_{k=1}^{K} \pi^{(k)}(\mathbf{H}^{\prime})= 1\) and \(\pi^{(k)}(\mathbf{H}^{\prime}) \in [0, 1]\).

4.1 Predictive Distribution for VMoGE↩︎

Our predictive distribution over the output label \(y\) integrates over the latent variables from all frequency band experts, weighted by the input-dependent gating mechanism. Formally, the predictive distribution can be expressed as:

\[\begin{align} \label{eq:mixture} p_{\theta}(y \mid \{ \mathbf{H}^{(k)},\mathbf{A}^{(k)}\}_{k=1}^{K}; \boldsymbol{\Theta}) &= \sum_{k=1}^{K} \pi^{(k)}(\mathbf{H}^{\prime}) \notag \\ &\quad\times \mathbb{E}_{q_{\boldsymbol{\phi}}(\mathbf{Z}^{(k)} \mid \mathbf{H}^{(k)}, \mathbf{A}^{(k)})} \left[ p(y \mid \mathbf{Z}^{(k)}; \boldsymbol{\theta}_k) \right], \end{align}\tag{7}\] where \(\boldsymbol{\Theta} = \{\{\boldsymbol{\theta}_k\}_{k=1}^{K}, \boldsymbol{\phi}\}\) denotes the complete set of model parameters, including the expert-specific parameters \(\boldsymbol{\theta}_k\), and encoder parameters \(\boldsymbol{\phi}\). Here, gating weights \(\pi^{(k)}(\mathbf{H}^{\prime})\) are computed dynamically for each input. Here, the likelihood term \(p(y \mid \mathbf{Z}^{(k)}; \boldsymbol{\theta}_k)\) represents the conditional probability of the class label given the latent representation for the \(k\)-th expert.

4.2 Optimizing Variational Structured Inference Lower Bounds↩︎

Unlike traditional optimization methods that use lower bounds, we incorporate the structured properties of each band’s relational graph into the variational lower bound objective. We aim to effectively approximate the latent posterior distributions and preserve the structural dependencies embedded in the graph data. Specifically, we formulate our training objective based on the evidence lower bound (ELBO), which balances two critical components: predictive performance and latent regularization. Our goal is to maximize the ELBO defined as:

\[\begin{align} \label{eq:ELBO} \max_{\boldsymbol{\Theta}} \;\mathcal{L}_{\mathrm{ELBO}} =&\; \mathbb{E}_{q_{\boldsymbol{\phi}}}\left[\log p_{\boldsymbol{\theta}} \left( y \mid \{\mathbf{Z}^{(k)}\}_{k=1}^{K}, \{\mathbf{H}^{(k)}\}_{k=1}^{K}; \boldsymbol{\Theta} \right)\right] \notag \\ &- \lambda \cdot \sum_{k=1}^{K} D_{\mathrm{KL}}^{(k)}\left(q_{\phi}(\mathbf{Z}^{(k)} \mid \mathbf{H}^{(k)},\mathbf{A}^{(k)}) \;\| \;p(\mathbf{Z}^{(k)} \mid \mathbf{A}^{(k)})\right ) \notag \\ \end{align}\tag{8}\] where \(\lambda\) is a regularization parameter, and the first term maximizes the expected log-likelihood of the target label under the joint variational posterior distribution over the band-specific latent representations. We assume that this posterior factorizes across the \(K\) frequency bands as \(q_{\boldsymbol{\phi}} \left( \{\mathbf{Z}^{(k)}\}_{k=1}^{K} \mid \{\mathbf{H}^{(k)},\mathbf{A}^{(k)}\}_{k=1}^{K} \right) = \prod_{k=1}^{K} q_{\boldsymbol{\phi}}^{(k)} \left( \mathbf{Z}^{(k)} \mid \mathbf{H}^{(k)},\mathbf{A}^{(k)} \right)\). The second term imposes a Kullback–Leibler (KL) divergence that regularizes the learned latent representations toward the conditional mixture of GMRF priors \(p(\mathbf{Z}^{(k)} \mid \mathbf{A}^{(k)})\) defined in Eq. 4 , ensuring that the latent space respects the structural dependencies encoded in each band-specific graph. The coefficient \(\lambda\) controls the strength of this structural regularization.

4.2.0.1 GMRF Expert KL Divergence

To model structured prior information from band-specific relational graphs, we construct a structured variational posterior distribution \(q_{\phi}(\mathbf{Z}^{(k)} \mid \mathbf{H}^{(k)}, \mathbf{A}^{(k)})\) and a prior distribution \(p(\mathbf{Z}^{(k)} \mid \mathbf{A}^{(k)})\) to formulate the KL divergence \(D_{KL}(q(\cdot) \mid\mid p(\cdot))\). As established in Eq. 5 and Eq. 4 , the posterior factorizes over nodes while the GMRF prior factorizes over the \(d_z\) latent dimensions. Both formulations correspond to a matrix-variate Gaussian over \(\mathbf{Z}^{(k)}\) with independent columns. We therefore evaluate the KL divergence per latent dimension and sum over the \(d_z\) dimensions. Specifically, for the \(d\)-th latent dimension, the variational posterior over the \(C\)-node column \(\mathbf{z}_d^{(k)} \in \mathbb{R}^{C}\) admits the marginal variational posterior \[\label{eq:kl95posterior95dim} q_{\phi}(\mathbf{z}_{d}^{(k)}) = \mathcal{N}\!\left(\mathbf{z}_{d}^{(k)}; \boldsymbol{\mu}_{d}^{(k)}, \boldsymbol{\Sigma}_{d}^{(k)}\right),\tag{9}\] where covariance \(\boldsymbol{\Sigma}_{d}^{(k)} = \operatorname{diag}\!\left([(\sigma_{1,d}^{(k)})^2, \ldots, (\sigma_{C,d}^{(k)})^2]\right) \in \mathbb{R}^{C \times C}\) is diagonal by the mean-field assumption, with its \(i\)-th entry \((\sigma_{i,d}^{(k)})^2\) denoting the marginal variance of node \(i\) along the \(d\)-th latent dimension. The expert-specific GMRF prior incorporates the structural information from the adjacency matrix \(\mathbf{A}^{(k)}\) through the precision matrix \(\mathbf{Q}^{(k)}\): \[\label{eq:kl95prior95dim} p(\mathbf{z}_{d}^{(k)}) = \mathcal{N}(\mathbf{z}_{d}^{(k)}; \mathbf{0}, (\mathbf{Q}^{(k)})^{-1}),\tag{10}\] where \(\mathbf{Q}^{(k)} \in \mathbb{R}^{C \times C}\) is the positive-definite precision matrix from Eq. 3 . The KL divergence between \(q_{\phi}(\mathbf{z}_d^{(k)})\) and \(p(\mathbf{z}_d^{(k)})\) for a single dimension is given by

\[\begin{align} D_{KL}^{(d)} &= \mathbb{E}_{q_{\phi}(\mathbf{z}_d^{(k)})} \left[ \log \frac{q_{\phi}(\mathbf{z}_d^{(k)})}{p(\mathbf{z}_d^{(k)})} \right] \nonumber \\ &= \mathbb{E}_{q_{\phi}(\mathbf{z}_d^{(k)})} \left[ -\frac{1}{2}\log|\boldsymbol{\Sigma}_d^{(k)}| \right. \nonumber \\ &\quad \left. - \frac{1}{2}(\mathbf{z}_d^{(k)} - \boldsymbol{\mu}_d^{(k)})^{\top} (\boldsymbol{\Sigma}_d^{(k)})^{-1} (\mathbf{z}_d^{(k)} - \boldsymbol{\mu}_d^{(k)}) \right. \nonumber \\ &\quad \left. + \frac{1}{2}\log|(\mathbf{Q}^{(k)})^{-1}| + \frac{1}{2}(\mathbf{z}_d^{(k)})^{\top} \mathbf{Q}^{(k)} \mathbf{z}_d^{(k)} \right], \end{align}\] More specifically, we compute the expectations of the quadratic forms. For the centered quadratic form under \(q_{\phi}(\mathbf{z}_d^{(k)})\), since \(\mathbf{z}_d^{(k)} \sim \mathcal{N}(\boldsymbol{\mu}_d^{(k)}, \boldsymbol{\Sigma}_d^{(k)})\):

\[\begin{align} \label{eq:first95quadratic95term} &\mathbb{E}_{q_{\phi}(\mathbf{z}_d^{(k)})} \left[(\mathbf{z}_d^{(k)} - \boldsymbol{\mu}_d^{(k)})^{\top} (\boldsymbol{\Sigma}_d^{(k)})^{-1} (\mathbf{z}_d^{(k)} - \boldsymbol{\mu}_d^{(k)})\right] \nonumber \\ &\qquad = \mathbb{E}_{q_{\phi}(\mathbf{z}_d^{(k)})}\!\left[\text{tr}\!\left((\boldsymbol{\Sigma}_d^{(k)})^{-1} (\mathbf{z}_d^{(k)} - \boldsymbol{\mu}_d^{(k)})(\mathbf{z}_d^{(k)} - \boldsymbol{\mu}_d^{(k)})^{\top}\right)\right] \nonumber \\ &\qquad = \text{tr}(\mathbf{I}_{C}) = C. \end{align}\tag{11}\]

For the non-centered quadratic form involving the prior precision, we obtain \[\begin{align} \label{eq:quad95form952} \mathbb{E}_{q_{\phi}(\mathbf{z}_d^{(k)})}[(\mathbf{z}_d^{(k)})^{\top} \mathbf{Q}^{(k)} \mathbf{z}_d^{(k)}] &= \mathbb{E}_{q_{\phi}(\mathbf{z}_d^{(k)})}[\text{tr}(\mathbf{Q}^{(k)} \mathbf{z}_d^{(k)}(\mathbf{z}_d^{(k)})^{\top})] \nonumber \\ &= \text{tr}(\mathbf{Q}^{(k)} \boldsymbol{\Sigma}_d^{(k)}) + (\boldsymbol{\mu}_d^{(k)})^{\top} \mathbf{Q}^{(k)} \boldsymbol{\mu}_d^{(k)}. \end{align}\tag{12}\] Combining Eq. 11 and Eq. 12 , and summing over all \(d_z\) latent dimensions, we obtain the closed-form expression for the \(k\)th expert-specific KL divergence:

\[\begin{align} D^{(k)}_{\mathrm{KL}} &= \sum_{d=1}^{d_z} D_{KL}^{(k,d)} \nonumber \\ &= \frac{1}{2} \sum_{d=1}^{d_z} \bigg[ \mathrm{tr} \!\left( \mathbf{Q}^{(k)} \boldsymbol{\Sigma}_d^{(k)} \right) + (\boldsymbol{\mu}_d^{(k)})^{\top} \mathbf{Q}^{(k)} \boldsymbol{\mu}_d^{(k)} - C \nonumber \\ &\quad - \log \left| \boldsymbol{\Sigma}_d^{(k)}\mathbf{Q}^{(k)} \right| \bigg], \label{eq:final95kl} \end{align}\tag{13}\] where the trace term \(\mathrm{tr}(\mathbf{Q}^{(k)} \boldsymbol{\Sigma}_d^{(k)})\) penalizes posterior uncertainty according to the graph-structured prior precision, the quadratic term \((\boldsymbol{\mu}_d^{(k)})^{\top} \mathbf{Q}^{(k)} \boldsymbol{\mu}_d^{(k)}\) captures the deviation of the posterior mean from the prior mean under the graph-structured precision, and the log-determinant of the diagonal posterior covariance is \(\log|\boldsymbol{\Sigma}_d^{(k)}| = \sum_{i=1}^{C} \log (\sigma_{i,d}^{(k)})^2\). Because \(\mathbf{Q}^{(k)}\) is shared across the \(d_z\) dimensions, the term \(\log|\mathbf{Q}^{(k)}|\) is computed once and reused; in practice we evaluate it via a Cholesky factorization \(\mathbf{Q}^{(k)} = \mathbf{L}\mathbf{L}^{\top}\) as \(\log|\mathbf{Q}^{(k)}| = 2\sum_{i} \log L_{ii}\) for numerical stability.

4.2.0.2 Overall Optimization

In practice, we implement the negative ELBO as our overall loss function for the training stage as follows: \[\label{eq:total95loss} \mathcal{L}_{\text{total}} = \mathcal{L}_{\text{cls}} + \lambda \cdot \mathcal{L}_{\text{KL}}.\tag{14}\] where \(\mathcal{L}_{\text{cls}}\) denotes the binary classification loss computed from the weighted mixture of expert outputs in Eq. 6 , and \(\mathcal{L}_{\text{KL}} = \sum_{k=1}^{K} D_{\text{KL}}^{(k)}\) aggregates the expert-specific KL divergences in Eq. 13 over all \(K\) experts. The hyperparameter \(\lambda\) controls the strength of the structural regularization.

5 Experiments↩︎

To evaluate our proposed VMoGE on two EEG datasets, we designed our experiments to address our research questions as follows:

  • RQ1: Does the proposed VMoGE framework improve AUC and accuracy for EEG-based AD and FTD diagnosis compared to state-of-the-art feature extraction and graph-based models?

  • RQ2: How does integrating multiple frequency-band experts within the MoE framework influence the discriminative power and robustness of EEG-based dementia classification?

  • RQ3: Can the incorporation of GMRF priors enhance model stability and generalization across heterogeneous datasets, especially under limited-sample conditions?

  • RQ4: Do the expert gating weights learned by VMoGE reflect clinically meaningful associations with cognitive indicators (e.g., MMSE scores) and age-related EEG spectral changes, thus providing interpretable biomarkers for dementia progression?

5.1 Data Acquisition↩︎

  • Open AD dataset: [77] The EEG dataset contained resting-state recordings from 88 elderly subjects: 36 AD patients, 23 FTD patients, and 29 healthy controls (HC). EEG signals were acquired using a 19-channel clinical system (10-20 electrode placement) at 500 Hz during eyes-closed conditions, with recording durations averaging 12-14 minutes per subject. Participants underwent Mini-Mental State Examination (MMSE) with mean scores of 17.75 (AD), 22.17 (FTD), and 30 (HC). The dataset includes both raw and preprocessed data with bandpass filtering (0.5-45 Hz), artifact subspace reconstruction, and independent component analysis for denoising.

  • Session-based AD: To validate the proposed method, this study enrolled 123 participants. We excluded cases with other secondary causes of dementia or significant psychiatric history. Patients were categorized into three CDR types: CDR=0 (\(n=15\)), CDR=1 (\(n=84\)), and CDR\(\geq\)​2 (\(n=24\)). The AD patients were diagnosed with probable Alzheimer’s disease according to the NINCDS-ADRDA diagnostic criteria [78], with a mean age of 78.0 years (SD=8.6). EEG data were recorded using the Nicolet EEG system. Participants underwent a five-minute adaptation period before the test, followed by three segments of alternating 10- to 20-second closed-eye and open-eye periods, along with a light stimulation test. Electrodes were placed according to the international 10–20 system, totaling 19 leads. The sampling frequency was 256 Hz, with filtering set to high-pass 0.05 Hz and low-pass 70 Hz, supplemented by a 60 Hz notch filter. Electrode impedance was maintained below 3 k\(\Omega\). All signals used the earlobe as the reference point. Technicians monitored subjects’ alertness throughout testing, prompting them to stay awake when drowsiness signs appeared. After manual review, three 10-second segments free of significant noise were extracted from the closed-eye phase of the raw EEG data for subsequent data collection.

5.2 EEG Data Preprocessing↩︎

The raw EEG signals underwent a systematic multi-stage preprocessing pipeline to ensure signal quality and remove artifacts. Initially, band-pass filtering between 0.05 and 70 Hz was applied using a finite impulse response (FIR) filter with Hamming window design, followed by a 60 Hz notch filter to eliminate power line interference. The filtered signals were re-referenced to averaged bilateral mastoid electrodes (A1/A2), with average referencing employed when mastoid electrodes were unavailable. Artifact removal was conducted in two stages: first, Artifact Subspace Reconstruction (ASR) with a 0.5-second sliding window and 17 standard deviation cutoff threshold was applied to correct high-amplitude transient artifacts; second, Independent Component Analysis (ICA) using the Infomax algorithm was performed, with ICLabel-classified ocular and myogenic components automatically excluded from signal reconstruction. Following artifact removal, the cleaned continuous EEG data were segmented into consecutive, non-overlapping 6-second epochs, with incomplete terminal segments discarded. Signals were decomposed into frequency bands under two configurations: (1) four canonical bands: \(\delta\) (0.5–4 Hz), \(\theta\) (4–8 Hz), \(\alpha\) (8–13 Hz), and \(\beta\) (13–45 Hz); and (2) an extended five-band configuration in which the \(\beta\) band is split into \(\beta_1\) (13–25 Hz) and \(\gamma\) (25–45 Hz), as used in the experiments reported in Table. [tbl:table:overall].

The PSD was estimated for each epoch and channel using Welch’s method with 4-second FFT windows spanning 0.5 to 45 Hz. RBP features were computed by normalizing band-specific power by total power within 0.5 to 45 Hz. Extracted features were organized into subject-level tensors with zero-padding applied to equalize epoch counts across participants, and age and MMSE scores were aligned with EEG features for multimodal analyses.

Table 1: Hyper-parameter settings of VMoGE.
Parameter Values
EEG channels \(C = 19\)
Epoch length \(T = 6\,\mathrm{s}\)
Kernel ratios \(\mathcal{G}\) =\(\{0.04, 0.06, 0.08\}\times SR^{\dagger}\)
CNN kernel size \(k_{\mathrm{cnn}} = 11\)
Temporal filters \(D_T = 32\)
Pooling stride \(s_p = 4\)
Embedding dim. \(d = 64\)
Transformer depth \(L_{\mathrm{tr}} = 2\)
Attention heads 8
GCN hidden dim. \(H_{\mathrm{gcn}} = 64\)
Node feature dim. \(D' = 768\) per channel
Training stage Adam \((\eta = 10^{-3})\), \(E = 200\) epochs

3pt

\(^{\dagger}\) denotes the sampling rate (SR) of the EEG signal.

max width=

Five EEG frequency bands: \(\delta\) (0.5–4 Hz), \(\theta\) (4–8 Hz), \(\alpha\) (8–13 Hz), \(\beta_{1}\) (13–25 Hz), and \(\gamma\) (25–45 Hz).

Four EEG frequency bands: \(\delta\) (0.5–4 Hz), \(\theta\) (4–8 Hz), \(\alpha\) (8–13 Hz), and \(\beta\) (13–45 Hz).

Table 2: Comparison of EEG-based dementia classification methods on the Open AD dataset.
Study Model type Graph-based method Task(s) Interpretability Datasets Performance\(^{\dagger}\)
Miltiadous et al. (2021) [79] ML methods No AD vs FTD Decision rules Open AD AD: 78.5%, FTD: 86.3% Acc
Mootoo et al. (2023) [80] GDFT + SVM Yes AD vs HC Graph-frequency domain interpretation Open AD AD: 85% Acc
Parihar et al. (2023) [81] Wavelet + Ensemble ML No AD vs HC Physiological interpretation of EEG frequency bands Open AD AD: 86.9% Acc
Chen et al. (2023) [82] CNN + ViT fusion No AD vs FTD vs NC Attention map Open AD 87.33% Acc, 88.19% AUC
Lal et al. (2024) [83] ML methods No AD vs HC Feature importance AHEPA dataset AD: 91%; FTD: 93%; AD vs FTD: 91% Acc
Tawhid et al. (2025) [84] CNN-based No AD vs HC Layer-wise feature visualization Open AD AD: 97.08%; FTD: 98.14% Acc
Siuly et al. (2025) [85] STFT + CNN No AD vs HC Grad-CAM Open AD AD: 95.59%; FTD: 95.75% Acc
VMoGE Graph neural network and Mixture-of-Experts Yes HC vs AD Frequency band × Brain network × Clinical indicators Open AD and Session-based AD: \(83\%\) Acc, \(89\%\) AUC

2pt


\(^{\dagger}\) Best performance from recent studies on the same classification task.

5.3 Baseline Methods↩︎

In this work, we follow the experimental settings in Table. [tbl:tab:parameters95setting] and evaluate our proposed VMoGE on two diverse datasets using five-fold cross-validation, comparing it against a wide range of baseline approaches, including traditional EEG feature extraction models, transformer-based deep learning methods, and GNN models combined with MoE mechanisms. For EEG signal classification, traditional methods include EEGNet and EEGViT, where the former employs a lightweight CNN design to extract spatiotemporal and spectral features, while the latter introduces the vision transformer architecture to capture long-range dependencies in EEG sequences. In addition, methods such as DTCA-Net, Deformer, and ADFormer are representative transformer-based frameworks that focus on channel selection and multi-scale feature modeling, and have been widely applied to AD-related EEG analysis. Meanwhile, MGformer leverages a multi-granularity encoder design to extract temporal patterns across frequency bands, thereby enhancing the discriminative power of node representations within brain networks.

For graph-based comparative methods, we incorporated several Graph-MoE models, including GraphMoRE, MoGE, Mowst, and GraphDIVE. Additionally, we compare against recent graph neural network models, combined graph neural networks with expert routing strategies to effectively model the heterogeneity of brain networks, and demonstrate robustness in cross-subject classification tasks.

Table 3: Temporal granularity configurations.
Configuration Granularities
Fine [0.02, 0.03, 0.04]
Medium [0.04, 0.06, 0.08]
Coarse [0.08, 0.10, 0.12]
Mixed-1 [0.02, 0.06, 0.10]
Mixed-2 [0.03, 0.07, 0.11]
Single-Fine [0.02]
Single-Medium [0.06]
Single-Coarse [0.10]
Figure 3: Ablation study of MGT-NFE module using different granularity components for EEG feature extraction.
Table 4: The \(\lambda\) value ablation study results showing AUC scores across different tasks and datasets. The red values indicate the highest AUC, and the second best values are shown in blue for each prior type within each task.
Open AD
HC vs FTD HC vs AD FTD vs AD
\(0.1\) \(0.2\) \(0.6\) \(0.8\) \(1.0\) \(0.1\) \(0.2\) \(0.6\) \(0.8\) \(1.0\) \(0.1\) \(0.2\) \(0.6\) \(0.8\) \(1.0\)
No Prior 0.76 0.68 0.67 0.60 0.62 0.83 0.78 0.75 0.70 0.77 0.68 0.66 0.68 0.61 0.79
GMRF (\(\textbf{L} + \lambda_{Q} \textbf{I}\)) 0.73 0.72 0.68 0.66 0.70 0.77 0.78 0.81 0.63 0.74 0.69 0.75 0.76 0.71 0.60
GMRF (\(\textbf{L}{\textit{norm}} + \lambda_{Q} \textbf{I}\)) 0.74 0.71 0.69 0.67 0.47 0.80 0.78 0.83 0.72 0.74 0.74 0.70 0.67 0.68 0.72
GMRF 0.77 0.58 0.69 0.66 0.65 0.78 0.77 0.78 0.77 0.78 0.70 0.68 0.67 0.84 0.74
Session-based AD
CDR=0 vs CDR=1 CDR=0 vs CDR=2 CDR=1 vs CDR=2
\(0.1\) \(0.2\) \(0.6\) \(0.8\) \(1.0\) \(0.1\) \(0.2\) \(0.6\) \(0.8\) \(1.0\) \(0.1\) \(0.2\) \(0.6\) \(0.8\) \(1.0\)
No Prior 0.66 0.65 0.66 0.67 0.61 0.66 0.73 0.71 0.77 0.68 0.59 0.60 0.47 0.58 0.56
GMRF (\(\textbf{L} + \lambda_{Q} \textbf{I}\)) 0.55 0.61 0.63 0.73 0.66 0.70 0.72 0.71 0.67 0.64 0.59 0.59 0.59 0.61 0.62
GMRF (\(\textbf{L}{\textit{norm}} + \lambda_{Q} \textbf{I}\)) 0.69 0.65 0.64 0.60 0.70 0.73 0.75 0.70 0.70 0.75 0.56 0.57 0.59 0.63 0.56
GMRF 0.64 0.66 0.64 0.61 0.59 0.72 0.76 0.81 0.67 0.79 0.56 0.47 0.56 0.48 0.52

5.3.1 Comparison With EEG Feature Extraction Models↩︎

Table. [tbl:table:overall] presents the VMoGE performance, including Area Under the Curve (AUC) and Accuracy (ACC), which respectively reflect discriminative capability and overall performance for EEG classification. Conventional CNN or Transformer-based EEG feature extraction models (e.g., EEGNet, EEGViT, DTCA-Net, Deformer, ADFormer, MGFormer) achieved limited performance in most tasks. For instance, in the HC vs AD setting with the Open AD dataset, Deformer and ADFormer achieved AUCs of 0.84 and 0.85, respectively, whereas our VMoGE achieved 0.89, corresponding to improvements of 0.05 and 0.04. In the HC vs FTD task, the best baseline (Deformer, AUC = 0.75) was still lower than our VMoGE (AUC = 0.78), with an improvement of 0.03. Moreover, in the Session-based CDR=1 vs CDR=2 classification, conventional methods achieved only limited performance, with EEGNet attaining the best AUC of 0.59, whereas our VMoGE reached an AUC of 0.65, yielding an improvement of 0.06.

5.3.2 Comparison With Graph-MoE Models↩︎

Among GNN architectures incorporating MoE mechanisms (e.g., GraphMoRE, MoGE, Mowst, GraphDIVE), VMoGE achieved superior performance across all tasks on the Open AD dataset. In HC vs FTD, VMoGE attained an AUC of 0.78, surpassing the best competitor Mowst (AUC = 0.64) by 0.14. For HC vs AD, VMoGE reached an AUC of 0.89, outperforming GraphMoRE and MoGE (both AUC = 0.85) by 0.04. In the challenging FTD vs AD classification task, VMoGE led with an AUC of 0.79, exceeding Mowst (AUC = 0.72) by 0.07. Such results show that conventional Graph-MoE models lack frequency-specific priors to distinguish dementia subtypes, while VMoGE’s GMRF-guided expert specialization consistently improves discrimination.

On the Session-based dataset, VMoGE maintained its advantage across staging tasks. In CDR=0 vs CDR=1, VMoGE achieved the highest AUC of 0.67, outperforming GraphDIVE (AUC = 0.53) by 0.14. For CDR=0 vs CDR=2, VMoGE obtained an AUC of 0.73, remaining competitive with GraphDIVE (AUC = 0.68) by 0.05. In CDR=1 vs CDR=2, VMoGE reached an AUC of 0.65, surpassing GraphDIVE (AUC = 0.58) by 0.07 and substantially exceeding Mowst (AUC = 0.50). This advantage confirms that VMoGE’s structured variational inference and graph-prior regularization enable robust staging even with scarce, noisy EEG signal.

5.3.3 Comparison With Recent EEG-based AD Classification Studies↩︎

To provide a comparative perspective on VMoGE, Table. [tbl:tab:compare95methods] reports representative recent EEG-based dementia classification studies evaluated using the Open AD dataset. Although several prior works report higher single-task accuracy than VMoGE, their experimental settings differ considerably and are often limited to a single dataset, precluding a direct comparison. First, the studies listed in Table. [tbl:tab:compare95methods] differ substantially in preprocessing pipelines, epoch length, validation protocols, and task formulation, all of which strongly affect reported performance [48], [79]. In particular, the epoch-level cross-validation adopted in [84] and [85] can yield accuracies above 95%, as is common with epoch-level splitting, epochs from the same subject are not separated at the subject level, potentially resulting in overly optimistic estimates of generalization performance. In contrast, the subject-independent five-fold cross-validation used in our work provides a more clinically realistic and inherently more conservative evaluation. Second, most existing methods focus exclusively on a single binary task such as HC vs. AD, whereas VMoGE simultaneously addresses AD classification and three CDR-based dementia staging tasks, demonstrating broader clinical applicability. Third, our framework prioritizes interpretability and biomarker identification by linking expert gating weights to specific frequency bands, brain regions, and clinical indicators (e.g., MMSE, age, CDR severity), whereas most CNN or wavelet based approaches in Table. [tbl:tab:compare95methods] operate as black box predictors with limited mechanistic insight into AD or FTD neuropathology. Finally, VMoGE is validated on two independent datasets, including a clinically collected Session-based dataset with CDR staging annotations, which is rarely reported in prior studies.

5.4 Computational Complexity for Clinical Implementation↩︎

Following the previous experimental setting, our VMoGE demonstrates the complete advantages of MoE in terms of both training cost and inference efficiency [86], [87]. As shown in Table. [tbl:tab:computational95efficiency], to ensure a fair comparison with classic transformer-based models (e.g., ADFormer, EEGFormer) that have higher computational complexity, we computed the latency, training time, and inference time as evaluation metrics. By exploiting expert-wise conditional computation, VMoGE avoids the heavy quadratic complexity typical of monolithic Transformer architectures and instead exhibits near-linear scaling with respect to the number of frequency-band experts. We reported that VMoGE achieved training times of 10.43–11.29s per epoch on the Open AD dataset, reducing the per-epoch training time by 2.92–3.63s compared with ADFormer (14.06-14.21s) and by 9.26–11.10 s compared with EEGDeformer (19.69-22.39s). More critically for clinical applications, VMoGE maintains remarkably low inference latency of 0.0017-0.0030s per sample, which is 2.8-6.4× faster than baseline methods.

@llcccc@ Task Group & Model & Dataset & Latency (ms) & Train (s) & Infer (s)
& & Open AD & \(0.97 \pm 0.02\) &\(10.43 \pm 0.03\) & \(0.0017 \pm 0.0000\)
& & Session & \(0.93 \pm 0.01\) &\(17.66 \pm 0.12\) & \(0.0029 \pm 0.0000\)
& & Open AD & \(0.47 \pm 0.15\) & \(14.06 \pm 0.08\) & \(0.0048 \pm 0.0014\)
& & Session & \(0.48 \pm 0.06\) & \(26.41 \pm 1.00\) & \(0.0113 \pm 0.0012\)
& & Open AD & \(0.41 \pm 0.07\) & \(19.69 \pm 3.21\) & \(0.0046 \pm 0.0005\)
& & Session & \(0.62 \pm 0.25\) & \(43.73 \pm 4.80\) & \(0.0146 \pm 0.0049\)
& & Open AD & \(1.00 \pm 0.01\) & \(11.28 \pm 0.26\) & \(0.0018 \pm 0.0001\)
& & Session & \(0.97 \pm 0.01\) & \(18.70 \pm 0.13\) & \(0.0030 \pm 0.0000\)
& & Open AD & \(0.57 \pm 0.14\) & \(14.27 \pm 0.18\) & \(0.0060 \pm 0.0016\)
& & Session & \(0.49 \pm 0.01\) & \(28.25 \pm 0.28\) & \(0.0117 \pm 0.0004\)
& & Open AD & \(0.43 \pm 0.02\) & \(21.89 \pm 0.39\) & \(0.0048 \pm 0.0003\)
& & Session & \(0.85 \pm 0.23\) & \(46.90 \pm 0.95\) & \(0.0191 \pm 0.0044\)
& & Open AD & \(1.01 \pm 0.00\) & \(11.29 \pm 0.13\) & \(0.0017 \pm 0.0001\)
& & Session & \(0.99 \pm 0.01\) & \(18.90 \pm 0.33\) & \(0.0030 \pm 0.0000\)
& & Open AD & \(0.51 \pm 0.19\) & \(14.21 \pm 0.08\) & \(0.0053 \pm 0.0019\)
& & Session & \(0.46 \pm 0.07\) & \(27.77 \pm 0.39\) & \(0.0110 \pm 0.0014\)
& & Open AD & \(0.56 \pm 0.24\) & \(22.39 \pm 0.31\) & \(0.0062 \pm 0.0029\)
& & Session & \(0.73 \pm 0.23\) & \(48.46 \pm 1.59\) & \(0.0164 \pm 0.0045\)

5.5 Ablation study for MGT-NFE using different granularity components for EEG feature extraction.↩︎

5.5.1 Granularities Module Ablation↩︎

We assessed the contribution of the MGT-NFE layer to EEG feature extraction under various temporal granularity configurations, as detailed in Table. 3, and evaluated classification performance across two datasets as shown in Fig. 3. Each configuration corresponds to a distinct temporal resolution. The Fine setting, defined by temporal intervals of 0.02, 0.03, and 0.04 seconds, is designed to capture short-term neural events. The Medium configuration, consisting of time windows of 0.04, 0.06, and 0.08 seconds, provides a balance between temporal stability and discriminability. In contrast, the Coarse configuration, with intervals of 0.08, 0.10, and 0.12 seconds, emphasizes long-term rhythmic dependencies in EEG dynamics. Furthermore, the Mixed-1 and Mixed-2 configurations integrate multiple temporal scales to model cross-frequency interactions. Single-scale variants, including Single-Fine, Single-Medium, and Single-Coarse, are employed as baselines to isolate the effects of individual temporal resolutions.

As shown in Fig. 3 (a), on the Open AD dataset, the Medium granularity achieved the highest F1-score of 0.860 in the HC vs AD classification task, with an overall mean of 0.783. This indicates that the medium-scale temporal range of 0.04 to 0.08 seconds effectively captures stable and representative EEG dynamics associated with cognitive decline. For the HC vs  FTD and FTD vs  AD tasks, the Mixed-2 configuration, defined by time windows of 0.03, 0.07, and 0.11 seconds, and the Single-Fine configuration, using a 0.02-second window, yielded the best performance (F1 \(\approx\) 0.589–0.783). These results suggest that incorporating both fine- and coarse-scale temporal cues enhances cross-scale discriminability among dementia subtypes. In the Session-based dataset in Fig. 3 (b), the Coarse granularity, which applies time windows between 0.08 and 0.12 seconds, achieved the highest F1-score of 0.712 in the CDR=0 vs CDR=2 task. This finding indicates that longer temporal windows are more effective for capturing extended cognitive progression patterns, particularly under small-sample and high-variance conditions.

5.5.2 Sensitivity Analysis for GMRF Priors Influence↩︎

To investigate the impact of the regularization parameter \(\lambda_{Q}\) on the prior of GMRF. We compare two GMRF prior specifications, the no-prior baseline and the structured prior defined in Eq. 3 , with analysis focusing on precision matrices of the form \(\boldsymbol{L} + \lambda_{Q} \boldsymbol{I}\) and \(\boldsymbol{L}_{\mathrm{norm}} + \lambda_{Q} \boldsymbol{I}\). Table. 4 presents the AUC scores for \(\lambda_{Q}\) values ranging from 0.1 to 1.0 across three types of binary classification tasks on both Open and Session-based AD datasets, with red highlighting indicating the optimal \(\lambda_{Q}\) value for each prior type within each task. We report that no prior GMRF achieves optimal performance at small \(\lambda_{Q}\) values on the Open AD (AUC = 0.83 at \(\lambda_{Q}\) = 0.1 for HC vs AD), indicating that weaker regularization benefits data-driven learning when spatial structure is not enforced. Conversely, GMRF methods exhibit distinct optimal ranges such as standard GMRF (\(\boldsymbol{L}+ \lambda_{Q} \boldsymbol{I}\)) performs best at moderate values (\(\lambda_{Q}\) = 0.2 to 0.6), normalized GMRF ( \(\boldsymbol{L}_{norm} + \lambda_{Q} \boldsymbol{I}\)) achieves the highest overall score (AUC = 0.83 at \(\lambda_{Q}\) = 0.6 for HC vs AD), and pure GMRF excels at both extremes, reaching the highest score with 0.84 at \(\lambda_{Q}\) = 0.8 for FTD vs AD across all configurations.

On the Session-based dataset, GMRF methods significantly improved classification performance for CDR=0 versus CDR=2 (pure GMRF: 0.81 vs no prior: 0.77) at \(\lambda_{Q} = 0.6\) and \(\lambda_{Q} = 0.8\). Performance progressively increased as \(\lambda_{Q}\) increased from 0.1 to 0.6, suggesting that spatial regularization provides greater benefits for datasets with limited sample sizes. More broadly, the optimal \(\lambda_{Q}\) values reflect the distinct characteristics of different datasets: small values (0.1-0.2) are suitable for FTD and AD with larger samples, moderate values (0.6) for balanced scenarios with clear spatial structure, and large values (0.8-1.0) for small-sample settings requiring stronger graph-prior regularization. These results indicate that spatial priors provide substantial gains when EEG brain networks exhibit inherent spatial structure.

Table 5: Ablation study results comparison: AUC and ACC across different tasks.
Open AD
Model HC vs FTD HC vs AD FTD vs AD
AUC ACC AUC ACC AUC ACC
S-E (\(\delta\)) 0.78±0.16 0.63±0.17 0.85±0.14 0.65±0.17 0.78±0.18 0.54±0.12
S-E (\(\theta\)) 0.69±0.17 0.57±0.09 0.89±0.10 0.73±0.21 0.85±0.16 0.63±0.10
S-E (\(\alpha\)) 0.74±0.09 0.54±0.14 0.80±0.18 0.57±0.13 0.79±0.20 0.74±0.15
S-E (\(\beta\)) 0.67±0.20 0.54±0.09 0.76±0.16 0.60±0.16 0.74±0.13 0.52±0.13
VMoGE 0.77±0.08 0.65±0.15 0.92±0.10 0.78±0.12 0.87±0.12 0.70±0.08
Session-based
Model CDR=0 vs CDR=1 CDR=0 vs CDR=2 CDR=1 vs CDR=2
AUC ACC AUC ACC AUC ACC
S-E (\(\delta\)) 0.66±0.11 0.58±0.07 0.78±0.11 0.57±0.09 0.68±0.10 0.58±0.06
S-E (\(\theta\)) 0.59±0.08 0.47±0.06 0.78±0.06 0.64±0.08 0.52±0.08 0.54±0.08
S-E (\(\alpha\)) 0.66±0.06 0.56±0.05 0.82±0.06 0.68±0.11 0.63±0.09 0.50±0.00
S-E (\(\beta\)) 0.64±0.10 0.51±0.07 0.74±0.16 0.53±0.10 0.55±0.12 0.54±0.12
VMoGE 0.69±0.11 0.60±0.16 0.72±0.09 0.62±0.12 0.68±0.09 0.60±0.07

3pt

S-E: Single-Expert using individual frequency band networks.

Figure 4: Box plots showing the distribution of mixture weights across different bands under three subtyping comparison conditions. The weights represent the contribution of different components in distinguishing between HC, FTD, and AD in the Open AD dataset in Fig. (a), and Session-based AD staging based on CDR=0, CDR=1, and CDR=2 in Fig.(b).
Figure 5: The expert gating weight analysis for different dementia indicator analysis. The top row shows age-based comparisons, and the bottom row shows MMSE-based comparisons. Each column represents different pairwise comparisons between HC and FTD, HC and AD, and FTD and AD.

5.5.3 Single-Expert and Full-Expert Ablation↩︎

In this section, we observe that flexible prior gating facilitates optimal allocation of frequency band expert outputs. In the experiments presented in Table. [tbl:table:MoE95comparision], we compare the Single-Expert (S-E) model, which uses only a single frequency band expert, with the Full-Expert (VMoGE) model, which integrates all frequency bands. In the Open AD, the performance of S-E models varied substantially across frequency bands. For example, \(\delta\) (AUC = 0.85) and \(\theta\) (AUC = 0.89) showed advantages in the HC vs AD task, while \(\alpha\) achieved the highest accuracy in the FTD vs AD task (ACC = 0.74). In contrast, \(\beta\) generally underperformed, suggesting that high-frequency activity alone is insufficient for robust classification. In the Session-based dataset, the \(\alpha\) band performed best in CDR=0 vs CDR=2 (AUC = 0.82), but overall accuracies were low and task-dependent fluctuations indicated the limitations of single-band experts in heterogeneous clinical data.

By contrast, VMoGE demonstrated overall stable and competitive performance across both datasets. In the Open AD dataset, it achieved the best results in HC vs AD (AUC = 0.92, ACC = 0.78) and maintained strong performance in FTD vs AD (AUC = 0.87). On the Session-based dataset, VMoGE achieved balanced results across the three CDR staging tasks, with AUCs of 0.69 (CDR=0 vs 1), and 0.68 (CDR=1 vs 2), demonstrating robustness against heterogeneous and limited-sample EEG settings.

Figure 6: Robustness evaluation of VMoGE under additive white Gaussian noise across SNR levels of 10, 20, and 30 dB. Panels A–C and D–F report AUC and accuracy, respectively, for three classification tasks on the Open AD and Session-based datasets.
Figure 7: Diagram of electrode spatial positions and corresponding brain regions in the EEG 10–20 system [88], where different colors represent distinct cortical areas.

a

b

Figure 8: Combined visualization of heatmaps (upper) and topographical maps (lower), showing spatial distributions and activation patterns across different conditions..

6 Noise in EEG Signal Analysis↩︎

Since EEG signals are inherently susceptible to noise interference, we introduced a framework with graph priors and evaluated its robustness under realistic clinical conditions by incorporating additive white Gaussian noise at signal Noise Ratio (SNR) levels of 10, 20 and 30 dB. As shown in Fig. 6, the Open AD dataset (solid lines) consistently outperformed the Session-based dataset (dashed lines) across all SNR levels, reflecting its higher signal quality and larger sample size. In the HC vs. AD task (panels B and E), both datasets maintained relatively stable performance across the tested SNR range, with Open AD sustaining AUC values between 0.83 and 0.89, indicating that the discriminative contrast between healthy controls and AD patients is sufficiently robust against moderate noise levels. For the HC vs. FTD task (panels A and D), Open AD showed moderate sensitivity to noise with AUC ranging from 0.63 to 0.75, while the Session-based AD exhibited lower and more variable performance, suggesting that α-band features distinguishing FTD from healthy controls are more susceptible to noise contamination. Notably, VMoGE maintained stable performance as SNR increased on the Open AD dataset without consistent monotonic degradation, suggesting that the GMRF-based structural priors contribute to noise resilience by enforcing graph-smoothness constraints that suppress spurious perturbations. These results demonstrate that VMoGE is capable of sustaining clinically acceptable performance under complex and realistic clinical environments.

7 Explainable Diagnosis↩︎

7.1 Gating Weights for Biomarker Analysis↩︎

In Fig. 4, we further examine the dynamic allocation of expert weights across frequency bands in the VMoGE model. For cross-group discrimination in Open AD tasks, the model assigned markedly higher weights to the \(\delta\) and \(\theta\)-bands in the HC vs AD group, while reducing the contribution of the \(\alpha\)- and \(\beta\)-bands. This aligns with the well-established pathological slowing in AD, characterized by enhanced slow-wave (\(\delta\slash\theta\)) activity and attenuated \(\alpha\slash\beta\) rhythms. In contrast, for HC vs FTD group, the \(\alpha\)-band dominated the gating weights that highlight the diagnostic relevance of \(\alpha\)-rhythm alterations in distinguishing frontotemporal dementia from normal aging. When comparing FTD vs AD, the gating weights were more balanced, with \(\delta\), \(\beta\), and \(\alpha\) receiving relatively higher contributions, indicating that a combination of slow-wave and high-frequency abnormalities best captures the differences between these two dementia subtypes.

For disease progression analysis in Session-based AD tasks, the weight distribution demonstrated a different pattern. As the disease progressed to CDR=2, the \(\theta\) and \(\alpha\)-bands exhibited further elevated weights, reflecting the progressive prominence of slow-wave activity, while these bands displayed greater variability, indicating that in some individuals, high-frequency abnormalities remained pronounced. Overall, across the progression from CDR=0 to CDR=2, the model gradually shifted from balanced reliance on multiple bands toward emphasizing slow-wave (\(\delta/\theta\)) signals, consistent with the clinical trajectory from mild to severe dementia.

7.2 Effects of Cognitive Function and Age on Expert Weight Allocation↩︎

In this section, we further investigate the interpretability of the VMoGE in AD diagnosis by analyzing the associations between the gating weights of frequency band-specific experts and clinical variables (e.g., age and MMSE scores) as presented in Fig. 5. In analyzing the Open AD dataset, the influence of age on expert weight allocation, we found that no frequency bands demonstrated significant correlations in the HC vs FTD group, indicating that the EEG features learned by the model are relatively stable and unaffected by age factors. In HC vs AD comparison, the \(\alpha\)-band showed higher diagnostic value in elderly subjects, as evidenced by increased discriminative weights assigned by the model (\(r = 0.300\), \(p = 0.0431\)). This likely reflects the diagnostic contrast between preserved \(\alpha\) characteristics in normal aging and AD-related pathological changes (e.g., \(\alpha\) slowing and desynchronization). In FTD vs AD, the \(\beta\)-band exhibited a significant negative correlation (\(r = -0.313\), \(p = 0.0341\)), indicating that it possesses greater discriminative capacity in younger patients.

To better measure cognitive function using MMSE scores, we observe that in HC vs AD task, the \(\delta\)-band exhibited a significant negative correlation (\(r = -0.336\), \(p = 0.0226\)), indicating that patients with poorer cognitive function showed higher discriminative value in this frequency band. This finding is consistent with the pathological increase in \(\delta\) waves observed in AD patients, suggesting that \(\delta\) waves become a more prominent neurophysiological marker in the late stages of the disease. In contrast, the \(\beta\)-band demonstrated a potential value in early diagnosis (\(r = 0.303\), \(p = 0.0404\)) suggesting that patients with better cognitive function retained more \(\beta\) wave characteristics.

7.3 Gating Weights for Spatial Brain Attribute Activation Analysis↩︎

Following prior research [88], [89], we aligned the channel clustering of electrodes’ spatial regions across major cortical regions, including Frontal Lobe (FL), Central Region (C), Occipital Lobe (OL), Parietal Lobe (PL), and Temporal Lobe (TL), as shown in Fig. 7. This alignment enables us to analyze and compare the spatial gating weights of different bands according to their spatial heatmaps, as illustrated in Fig. 8 (A) and topographical maps in Fig. 8 (B). In the Open AD dataset, the HC vs.AD comparison showed elevated \(\theta\)- and \(\alpha\)-band spatial weights over posterior regions (OL: \(\theta = 0.377\), \(\alpha = 0.375\); PL: \(\theta = 0.353\), \(\alpha = 0.335\)), with the highest regional weight observed for the \(\theta\) band in the temporal lobe (\(0.415\)), consistent with parieto-occipital network involvement in AD. For FTD vs AD, the \(\beta\)-band dominated the weight allocation and was concentrated in central and occipital regions (C3, Cz, C4, O1, O2), revealing the diagnostic value of high-frequency oscillatory abnormalities.

In the Session-based AD, disease progression from CDR=0 to CDR=2 was accompanied by the spatial expansion of brain region activation. For CDR=0 vs CDR=1, the \(\beta\) band received the highest weights across most regions, particularly in the temporal lobe (TL: \(0.432\)) and central region (C: \(0.406\)). For CDR=0 vs CDR=2, the weighting shifted toward slow-wave activity, with the \(\delta\) band dominating the central region (\(0.374\)) and the \(\theta\) band dominating posterior regions (OL: \(0.374\); PL: \(0.363\)). For CDR=1 vs CDR=2, the \(\beta\) band again dominated, with the largest weights observed in the temporal and central regions (TL: \(0.526\); C: \(0.522\)), while the \(\theta\) band remained moderately weighted over the posterior region (OL: \(0.381\)). These findings suggest that slow-wave features are most informative for distinguishing normal cognition from severe dementia, whereas central–temporal \(\beta\)-band activity contributes more strongly to discriminating adjacent severity stages.

8 Discussion↩︎

In this section, we further discuss VMoGE across three dimensions: (1) frequency-specific biomarkers in EEG that reveal distinct pathophysiological signatures, (2) spatial brain network alterations, and (3) structured graph modeling. The gating weight analysis in Fig. 4 shows that VMoGE primarily emphasized slow-wave activity for HC vs. AD classification, especially the \(\delta\)- and \(\theta\)-band contributions. This finding is consistent with AD-related EEG slowing [73], [90], [91], in which slow-wave activity becomes more prominent as cognitive impairment progresses. As shown in the age and cognitive function analysis in Fig. 5, the significant negative association between \(\delta\)-band weights and MMSE scores (\(r = -0.336\), \(p = 0.0226\)) further supports the role of \(\delta\) activity as a marker of more severe cognitive decline: as the condition worsens and MMSE scores decline [92]. In contrast, \(\beta\) power is positively related to cognitive performance (\(r = 0.303\), \(p = 0.0404\)), emphasizing its importance in cognitive processing [93], [94] and its contribution to working memory, executive control, and motor regulation [95][97]. Notably, the \(\beta\)-band demonstrated stronger discriminative capacity in younger patients (\(r = -0.313\), \(p = 0.0341\)), pointing to its potential as a biomarker for early-onset dementia in FTD vs AD.

At the spatial level, Fig. 8 shows that the \(\theta\)- and \(\alpha\)-band weights were elevated over posterior regions for HC vs.AD classification, consistent with disrupted posterior rhythms and impaired parietal–occipital cognitive networks [98], [99]. In the Session-based AD dataset, spatial weight allocation varied across disease stages rather than following a monotonic pattern [100]. Slow-wave contributions were most prominent when distinguishing normal cognition from severe dementia (CDR0= vs CDR=2), with the \(\delta\) band dominating the central region and the \(\theta\) band dominating posterior regions [92]. In contrast, adjacent-stage discrimination in CDR=1 vs CDR=2 relied primarily on \(\beta\)-band weights over the central (C3, Cz, and C4) and temporal (T7 and T8) regions. These findings indicate that \(\beta\)-band alterations are region-specific rather than uniformly attenuated and provide staging-relevant information that complements slow-wave markers [101], [102].

By integrating frequency-specific GMRF priors into the variational inference framework, this study successfully captured the structured dependencies of EEG brain networks. Ablation experiments shown in Table. 4 demonstrated that GMRF priors yielded significant performance improvements on the sample-limited session-based dataset, with AUC increasing from \(0.77\) (CDR=0 vs CDR=2, \(\lambda_{Q}=0.8\)) to \(0.81\) (CDR=0 vs CDR=2, \(\lambda_{Q}=0.6\)). This confirmed the regularization benefits of graph-structured priors in small-sample scenarios. The task-dependent optimal \(\lambda_{Q}\) values also reflected the heterogeneity of spatial smoothness in brain networks across different pathological states, with FTD/AD tasks favored \(\lambda_{Q}\) ranging from \(0.2\) to \(0.8\), while CDR=1/CDR=2 staging tasks favored \(\lambda_{Q}\) ranging from \(0.6\) to \(0.8\).

While VMoGE achieves strong subtype classification and severity staging from EEG alone, the multi-level neurodegenerative processes of AD are difficult to capture through a single modality. Recent work shows that fusing EEG with structural MRI via attention improves AD classification [103][105], that combining neuroimaging with genomic SNP data captures complementary disease signatures [106], that gene–brain interaction graph priors enhance both accuracy and interpretability [107], and that supervised contrastive learning across SNPs, proteomics, and MRI mitigates modality heterogeneity [108]. VMoGE is naturally suited to such extensions: each expert could represent a distinct modality (e.g., fMRI functional connectivity or genotype-informed networks), with the variational gating mechanism dynamically weighting modality-specific contributions, and the GMRF prior generalized to encode cross-modal dependencies such as structure–function coupling from simultaneous EEG-fMRI or gene–brain networks [109]. This multimodal extensibility positions VMoGE to enable more precise biomarker identification, more sensitive early diagnosis, and personalized disease risk stratification.

9 Conclusions↩︎

In this paper, we propose the VMoGE framework, a variational mixture of graph neural experts framework that integrates an MGT-NFE extractor capable of capturing multi-scale EEG features with GMRF-based structural graph priors. Through structured variational inference, the model dynamically allocates expert weights, enabling adaptive integration of multi-frequency band information to achieve precise dementia diagnosis based on frequency-specific EEG biomarkers. Furthermore, VMoGE reveals critical neurophysiological insights across two AD diagnostic datasets. These frequency-specific and spatially-localized biomarkers not only align closely with established neuropathological mechanisms but also demonstrate the effectiveness of the structured variational inference framework in capturing brain network heterogeneity. This work provides a deep learning solution that combines both accuracy and clinical interpretability for early dementia screening and disease subtype differentiation.

References↩︎

[1]
“2025 alzheimer’s disease facts and figures,” Alzheimer’s & Dementia, vol. 21, no. 4, p. e70235, Apr. 2025, doi: 10.1002/alz.70235.
[2]
Ministry of Health and Welfare, Taiwan, Accessed: 2025-08-30“MOHW announces latest epidemiological survey results on dementia in taiwan communities.” https://www.mohw.gov.tw/cp-16-78102-1.html, 2024.
[3]
N. D. Weder, R. Aziz, K. Wilkins, and R. R. Tampi, “Frontotemporal dementias: A review,” Annals of general psychiatry, vol. 6, no. 1, p. 15, 2007.
[4]
D. Bathgate, J. Snowden, A. Varma, A. Blackshaw, and D. Neary, “Behaviour in frontotemporal dementia, alzheimer’s disease and vascular dementia,” Acta neurológica scandinavica, vol. 103, no. 6, pp. 367–378, 2001.
[5]
O. Hansson, S. Lehmann, M. Otto, H. Zetterberg, and P. Lewczuk, “Advantages and disadvantages of the use of the CSF amyloid \(\beta\) (a\(\beta\)) 42/40 ratio in the diagnosis of alzheimer’s disease,” Alzheimer’s research & therapy, vol. 11, no. 1, p. 34, 2019.
[6]
A. Leuzy et al., “Considerations in the clinical use of amyloid PET and CSF biomarkers for alzheimer’s disease,” Alzheimer’s & Dementia, vol. 21, no. 3, p. e14528, 2025.
[7]
S. Gluhm, J. Goldstein, K. Loc, A. Colt, C. Van Liew, and J. Corey-Bloom, “Cognitive performance on the mini-mental state examination and the montreal cognitive assessment across the healthy adult lifespan,” Cognitive and Behavioral Neurology, vol. 26, no. 1, pp. 1–5, 2013.
[8]
R. O’Caoimh, S. Timmons, and D. W. Molloy, “Screening for mild cognitive impairment: Comparison of ‘MCI specific’ screening instruments,” Journal of Alzheimer’s disease, vol. 51, no. 2, pp. 619–629, 2016.
[9]
G. D’Ignazio et al., “Is the MMSE enough for MCI? A narrative review of the usefulness of the MMSE,” Frontiers in Psychology, vol. 16, p. 1727738, 2025.
[10]
A. G. Vrahatis, K. Skolariki, M. G. Krokidis, K. Lazaros, T. P. Exarchos, and P. Vlamos, “Revolutionizing the early detection of alzheimer’s disease through non-invasive biomarkers: The role of artificial intelligence and deep learning,” Sensors, vol. 23, no. 9, p. 4184, 2023.
[11]
P. M. McMahon, S. S. Araki, E. A. Sandberg, P. J. Neumann, and G. S. Gazelle, “Cost-effectiveness of PET in the diagnosis of alzheimer disease,” Radiology, vol. 228, no. 2, pp. 515–522, 2003.
[12]
N. Mattsson-Carlgren et al., “Plasma biomarker strategy for selecting patients with alzheimer disease for antiamyloid immunotherapies,” JAMA neurology, vol. 81, no. 1, pp. 69–78, 2024.
[13]
D. Baldaranov et al., “Safety and tolerability of lumbar puncture for the evaluation of alzheimer’s disease,” Alzheimer’s & Dementia: Diagnosis, Assessment & Disease Monitoring, vol. 15, no. 2, p. e12431, 2023.
[14]
L. VandeVrede and S. E. Schindler, “Clinical use of biomarkers in the era of alzheimer’s disease treatments,” Alzheimer’s & Dementia, vol. 21, no. 1, p. e14201, 2025.
[15]
S. M. Razavi, T. M. Tran, S. E. Wilson, and S. Khanmohammadi, “Brain state transition disruptions in alzheimer’s disease: Insights from EEG state dynamics,” in 2025 IEEE 22nd international symposium on biomedical imaging (ISBI), 2025, pp. 1–5.
[16]
J. Cruzat et al., “Temporal irreversibility of large-scale brain dynamics in alzheimer’s disease,” Journal of Neuroscience, vol. 43, no. 9, pp. 1643–1656, 2023.
[17]
M. Sandonı́s-Fernández et al., “Disrupted temporal structure of the m/EEG meta-states sequencing in alzheimer’s disease,” NeuroImage, p. 121555, 2025.
[18]
H. Irani, Y. Ghahremani, A. Kermani, and V. Metsis, “Time series embedding methods for classification tasks: A review,” Expert Systems, vol. 42, no. 11, p. e70148, 2025.
[19]
A. M. Karimi Forood, S. M. Pandian, R. Q. McNaboe, T. De Carvalho Lachos, D. O. Lantigua, and H. F. Posada-Quintero, “Development and validation of a multimodal wearable belt for abdominal biosignal monitoring with application to irritable bowel syndrome,” Micromachines, vol. 16, no. 11, p. 1255, 2025.
[20]
L. S. M. Jahromi et al., “Efficacy of acupuncture on pain severity and frequency of calf cramps in dialysis patients: A randomized clinical trial,” Innovations in Acupuncture and Medicine, vol. 17, no. 2, pp. 47–54, 2024.
[21]
M. Correia de Carvalho, J. Nunes de Azevedo, P. Azevedo, C. Pires, M. Laranjeira, and J. P. Machado, “Effect of acupuncture on functional capacity in patients undergoing hemodialysis: A patient-assessor blinded randomized controlled trial,” in Healthcare, 2022, vol. 10, p. 1947.
[22]
N. Lin, J. Gao, C. Mao, H. Sun, Q. Lu, and L. Cui, “Differences in multimodal electroencephalogram and clinical correlations between early-onset alzheimer’s disease and frontotemporal dementia,” Frontiers in Neuroscience, vol. 15, p. 687053, 2021.
[23]
B. Jiao et al., “Neural biomarker diagnosis and prediction to mild cognitive impairment and alzheimer’s disease using EEG technology,” Alzheimer’s research & therapy, vol. 15, no. 1, p. 32, 2023.
[24]
C. A. Chetty et al., “EEG biomarkers in alzheimer’s and prodromal alzheimer’s: A comprehensive analysis of spectral and connectivity features,” Alzheimer’s Research & Therapy, vol. 16, no. 1, p. 236, 2024.
[25]
W. Khan et al., “An explainable and efficient deep learning framework for EEG-based diagnosis of alzheimer’s and frontotemporal dementia,” Frontiers in Medicine, vol. 12, p. 1590201, 2025.
[26]
K. Rezaee and M. Zhu, “Diagnose alzheimer’s disease and mild cognitive impairment using deep CascadeNet and handcrafted features from EEG signals,” Biomedical Signal Processing and Control, vol. 99, p. 106895, 2025.
[27]
B. Cansiz, H. ILHAN, N. Aydin, and G. Serbes, “Olfactory EEG based alzheimer disease classification through transformer based feature fusion with tunable q-factor wavelet coefficient mapping,” Frontiers in Neuroscience, vol. 19, p. 1638922, 2025.
[28]
S. Latif, N. U. Islam, Z. Uddin, K. M. Cheema, S. S. Ahmed, and M. F. Khan, “Deep ensemble learning with transformer models for enhanced alzheimer’s disease detection,” Scientific Reports, vol. 15, no. 1, p. 24720, 2025.
[29]
M. Şeker and M. S. Özerdem, “Investigating convolutional and transformer-based models for classifying mild cognitive impairment using 2D spectral images of resting-state EEG,” Biomedical Signal Processing and Control, vol. 105, p. 107667, 2025.
[30]
A. Miltiadous, E. Gionanidis, K. D. Tzimourta, N. Giannakeas, and A. T. Tzallas, “DICE-net: A novel convolution-transformer architecture for alzheimer detection in EEG signals,” IEEe Access, vol. 11, pp. 71840–71858, 2023.
[31]
Y. Wang, N. Huang, N. Mammone, M. Cecchi, and X. Zhang, “LEAD: Large foundation model for EEG-based alzheimer’s disease detection,” arXiv preprint arXiv:2502.01678, 2025.
[32]
D. Zhang and C. Zhu, “A dual path graph neural network framework for dementia diagnosis,” Scientific Reports, vol. 15, no. 1, p. 23319, 2025.
[33]
Q. Xu, L. An, H. Yang, and K.-S. Hong, “A multi-graph convolutional network method for alzheimer’s disease diagnosis based on multi-frequency EEG data with dual-mode connectivity,” Frontiers in Neuroscience, vol. 19, p. 1555657, 2025.
[34]
Z. Zhou et al., “A novel graph neural network method for alzheimer’s disease classification,” Computers in Biology and Medicine, vol. 180, p. 108869, 2024.
[35]
A. T. Adebisi, H.-W. Lee, and K. C. Veluvolu, “EEG-based brain functional network analysis for differential identification of dementia-related disorders and their onset,” IEEE Transactions on Neural Systems and Rehabilitation Engineering, vol. 32, pp. 1198–1209, 2024.
[36]
D. Klepl, F. He, M. Wu, D. J. Blackburn, and P. Sarrigiannis, “Adaptive gated graph convolutional network for explainable diagnosis of alzheimer’s disease using EEG data,” IEEE Transactions on Neural Systems and Rehabilitation Engineering, vol. 31, pp. 3978–3987, 2023.
[37]
S. J. Ahn and M. Kim, “Variational graph normalized autoencoders,” in Proceedings of the 30th ACM international conference on information & knowledge management, 2021, pp. 2827–2831.
[38]
G. Liang, P. Tiwari, S. Nowaczyk, S. Byttner, and F. Alonso-Fernandez, “Dynamic causal explanation based diffusion-variational graph neural network for spatiotemporal forecasting,” IEEE Transactions on Neural Networks and Learning Systems, 2024.
[39]
W. Cai, J. Jiang, F. Wang, J. Tang, S. Kim, and J. Huang, “A survey on mixture of experts in large language models,” IEEE Transactions on Knowledge and Data Engineering, 2025.
[40]
C. Wu, Z. Shuai, Z. Tang, L. Wang, and L. Shen, “Dynamic modeling of patients, modalities and tasks via multi-modal multi-task mixture of experts,” in The thirteenth international conference on learning representations, 2025.
[41]
B. Wang et al., “Dynamical multimodal fusion with mixture-of-experts for localizations,” arXiv preprint arXiv:2507.01337, 2025.
[42]
Y. Xu, J. Wang, M. Guang, and C. Jiang, “Graph multi-convolution and attention pooling for graph classification,” IEEE Transactions on Pattern Analysis and Machine Intelligence, 2024.
[43]
Y. Wang, Z. Yang, and X. Che, “A hierarchical mixture-of-experts framework for few labeled node classification,” Neural Networks, vol. 188, p. 107285, 2025.
[44]
Y. Shi, Y. Wang, W. Liang, J. Zhang, P. Dong, and A. Li, “Mixture of experts for node classification,” in Proceedings of the 2025 international conference on multimedia retrieval, 2025, pp. 1154–1162.
[45]
F. Hu, L. Wang, S. Wu, L. Wang, and T. Tan, “Graph classification by mixture of diverse experts,” arXiv preprint arXiv:2103.15622, 2021.
[46]
V. Y. Shirasuna et al., “A multi-view mixture-of-experts based on language and graphs for molecular properties prediction,” in ICML 2024 AI for science workshop, 2024.
[47]
X. Huang, W. Chen, B. Hu, and Z. Mao, “Graph mixture of experts and memory-augmented routers for multivariate time series anomaly detection,” in Proceedings of the AAAI conference on artificial intelligence, 2025, vol. 39, pp. 17476–17484.
[48]
F. Del Pup, A. Zanola, L. F. Tshimanga, A. Bertoldo, and M. Atzori, “The more, the better? Evaluating the role of EEG preprocessing for deep learning applications.” IEEE Transactions on Neural Systems and Rehabilitation Engineering, 2025.
[49]
S. Rezaei et al., “Future of alzheimer’s detection: Advancing diagnostic accuracy through the integration of qEEG and artificial intelligence,” NeuroImage, p. 121373, 2025.
[50]
M. Lopes, R. Cassani, and T. H. Falk, “Using CNN saliency maps and EEG modulation spectra for improved and more interpretable machine learning-based alzheimer’s disease diagnosis,” Computational Intelligence and Neuroscience, vol. 2023, no. 1, p. 3198066, 2023.
[51]
W. Xia, R. Zhang, X. Zhang, and M. Usman, “A novel method for diagnosing alzheimer’s disease using deep pyramid CNN based on EEG signals,” Heliyon, vol. 9, no. 4, 2023.
[52]
A. Kowshiga, T. Pavithra, V. Priyanka, et al., “Deep learning into the future: Hybrid CNN-RNN for early detection of alzheimer’s disease,” in 2024 5th international conference on electronics and sustainable communication systems (ICESC), 2024, pp. 940–946.
[53]
J. Sreedhar et al., “Fuzzy RNN model-based classification of alzheimer’s disease and dementia using brain EEG signals,” IEEE Transactions on Consumer Electronics, 2025.
[54]
Y. Liu, L. An, H. Yang, and S. S. Ge, “Multi-frequency EEG and multi-functional connectivity graph convolutional network based detection method of patients with alzheimer’s disease,” Complex & Intelligent Systems, vol. 11, no. 8, pp. 1–17, 2025.
[55]
A. C. Yang et al., “Cognitive and neuropsychiatric correlates of EEG dynamic complexity in patients with alzheimer’s disease,” Progress in Neuro-Psychopharmacology and Biological Psychiatry, vol. 47, pp. 52–61, 2013.
[56]
D. Bethge et al., “EEG2Vec: Learning affective EEG representations via variational autoencoders,” in 2022 IEEE international conference on systems, man, and cybernetics (SMC), 2022, pp. 3150–3157.
[57]
T. Zhang, Z. Cui, C. Xu, W. Zheng, and J. Yang, “Variational pathway reasoning for EEG emotion recognition,” in Proceedings of the AAAI conference on artificial intelligence, 2020, vol. 34, pp. 2709–2716.
[58]
T. Song et al., “Variational instance-adaptive graph for EEG emotion recognition,” IEEE Transactions on Affective Computing, vol. 14, no. 1, pp. 343–356, 2021.
[59]
A. Ofner and S. Stober, “Balancing active inference and active learning with deep variational predictive coding for EEG,” in 2020 IEEE international conference on systems, man, and cybernetics (SMC), 2020, pp. 3839–3844.
[60]
C. Liu et al., “VSGT: Variational spatial and gaussian temporal graph models for EEG-based emotion recognition,” in Proceedings of the thirty-third international joint conference on artificial intelligence, 2024, pp. 3078–3086.
[61]
L. Xu, M. Jordan, and G. E. Hinton, “An alternative model for mixtures of experts,” Advances in neural information processing systems, vol. 7, 1994.
[62]
M. Liu et al., “EMPT: A sparsity transformer for EEG-based motor imagery recognition,” Frontiers in Neuroscience, vol. 18, p. 1366294, 2024.
[63]
Y. Gui, M. Chen, Y. Su, G. Luo, and Y. Yang, “EEGMamba: Bidirectional state space model with mixture of experts for EEG multi-task classification,” arXiv preprint arXiv:2407.20254, 2024.
[64]
Z. Du, R. Peng, W. Liu, W. Li, and D. Wu, “Mixture of experts for EEG-based seizure subtype classification,” IEEE Transactions on Neural Systems and Rehabilitation Engineering, vol. 31, pp. 4781–4789, 2023.
[65]
E. D. Übeyli, “Wavelet/mixture of experts network structure for EEG signals classification,” Expert systems with applications, vol. 34, no. 3, pp. 1954–1962, 2008.
[66]
X. Zhu, Y. Chai, R. Li, M. Lan, and L. Gao, “Decoding the moving mind: Multi-subject fMRI-to-video retrieval with MLLM semantic grounding,” bioRxiv, pp. 2025–04, 2025.
[67]
D. Wu, Y. Jia, S. Li, S. Zhao, J. Yang, and M. Sawan, “Neuro-MoBRE: Exploring multi-subject multi-task intracranial decoding via explicit heterogeneity resolving,” arXiv preprint arXiv:2508.04128, 2025.
[68]
L. Yang et al., “EEG emotion recognition via identity based multi-gate mixture-of-experts network,” in 2022 IEEE international conference on bioinformatics and biomedicine (BIBM), 2022, pp. 2498–2505.
[69]
X. Zhu, Y. Wang, Y. Ma, Z. Zhang, and Y. Gao, “CognitMoE: A cognition-aware collaborative multi-expert network for bipolar disorder diagnosis,” Neural Networks, p. 107854, 2025.
[70]
J. Zhang et al., “BrainNet-MoE: Brain-inspired mixture-of-experts learning for neurological disease identification,” arXiv preprint arXiv:2503.07640, 2025.
[71]
K. Ghassemkhani, K. S. Saroka, and B. T. Dotta, “Evaluating EEG complexity and spectral signatures in alzheimer’s disease and frontotemporal dementia: Evidence for rostrocaudal asymmetry,” npj Aging, vol. 11, no. 1, p. 50, 2025.
[72]
R. Rahman, A. K. al Azad, and S. Momen, “MGFormer: A lightweight multi-granular transformer for subject-independent alzheimer’s classification,” Biomedical Signal Processing and Control, vol. 110, p. 108230, 2025.
[73]
A. H. Meghdadi et al., “Resting state EEG biomarkers of cognitive decline associated with alzheimer’s disease and mild cognitive impairment,” PloS one, vol. 16, no. 2, p. e0244180, 2021.
[74]
D. P. Kingma and M. Welling, arXiv:1312.6114“Auto-encoding variational bayes,” in Proceedings of the 2nd international conference on learning representations (ICLR), 2014.
[75]
R. A. Jacobs, M. I. Jordan, S. J. Nowlan, and G. E. Hinton, “Adaptive mixtures of local experts,” Neural computation, vol. 3, no. 1, pp. 79–87, 1991.
[76]
F. Hu, L. Wang, S. Wu, L. Wang, and T. Tan, “Graph classification by mixture of diverse experts,” arXiv preprint arXiv:2103.15622, 2021.
[77]
A. Miltiadous et al., “A dataset of scalp EEG recordings of alzheimer’s disease, frontotemporal dementia and healthy subjects from routine EEG,” Data, vol. 8, no. 6, p. 95, 2023.
[78]
B. Dubois et al., “Research criteria for the diagnosis of alzheimer’s disease: Revising the NINCDS–ADRDA criteria,” The Lancet Neurology, vol. 6, no. 8, pp. 734–746, 2007.
[79]
A. Miltiadous et al., “Alzheimer’s disease and frontotemporal dementia: A robust classification method of EEG signals and a comparison of validation methods,” Diagnostics, vol. 11, no. 8, p. 1437, 2021.
[80]
X. S. Mootoo, A. Fours, C. Dinesh, M. Ashkani, A. Kiss, and M. Faltyn, “Detecting alzheimer disease in EEG data with machine learning and the graph discrete fourier transform,” medRxiv, pp. 2023–11, 2023.
[81]
A. Parihar and P. D. Swami, “EEG classification of alzheimer’s disease, frontotemporal dementia and control normal subjects using supervised machine learning algorithms on various EEG frequency bands,” Int J Sci Technol Manag, vol. 12, no. 6, pp. 82–89, 2023.
[82]
Y. Chen, H. Wang, D. Zhang, L. Zhang, and L. Tao, “Multi-feature fusion learning for alzheimer’s disease prediction using EEG signals in resting state,” Frontiers in neuroscience, vol. 17, p. 1272834, 2023.
[83]
U. Lal, A. V. Chikkankod, and L. Longo, “A comparative study on feature extraction techniques for the discrimination of frontotemporal dementia and alzheimer’s disease with electroencephalography in resting-state adults,” Brain sciences, vol. 14, no. 4, p. 335, 2024.
[84]
M. N. A. Tawhid, S. Siuly, E. Kabir, and Y. Li, “Advancing alzheimer’s disease detection: A novel convolutional neural network based framework leveraging EEG data and segment length analysis: MN ahad tawhid et al.” Brain Informatics, vol. 12, no. 1, p. 13, 2025.
[85]
S. Siuly, M. N. A. Tawhid, Y. Li, R. Acharya, M. T. Sadiq, and H. Wang, “Investigating brain lobe biomarkers to enhance dementia detection using EEG data,” Cognitive Computation, vol. 17, no. 2, p. 90, 2025.
[86]
N. Shazeer et al., “Outrageously large neural networks: The sparsely-gated mixture-of-experts layer,” arXiv preprint arXiv:1701.06538, 2017.
[87]
W. Fedus, B. Zoph, and N. Shazeer, “Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity,” Journal of Machine Learning Research, vol. 23, no. 120, pp. 1–39, 2022.
[88]
J. Sanchis, S. Garcı́a-Ponsoda, M. A. Teruel, J. Trujillo, and I.-Y. Song, “A novel approach to identify the brain regions that best classify ADHD by means of EEG and deep learning,” Heliyon, vol. 10, no. 4, 2024.
[89]
J. N. Acharya, A. J. Hani, J. Cheek, P. Thirumala, and T. N. Tsuchida, “American clinical neurophysiology society guideline 2: Guidelines for standard electrode position nomenclature,” The Neurodiagnostic Journal, vol. 56, no. 4, pp. 245–252, 2016.
[90]
M. Kopčanová et al., “Resting-state EEG signatures of alzheimer’s disease are driven by periodic but not aperiodic changes,” Neurobiology of Disease, vol. 190, p. 106380, 2024.
[91]
N. Benz et al., “Slowing of EEG background activity in parkinson’s and alzheimer’s disease with early cognitive dysfunction,” Frontiers in aging neuroscience, vol. 6, p. 314, 2014.
[92]
Y. T. Kwak, “Quantitative EEG findings in different stages of alzheimer’s disease,” Journal of clinical neurophysiology, vol. 23, no. 5, pp. 457–462, 2006.
[93]
H. Azami et al., “Beta to theta power ratio in EEG periodic components as a potential biomarker in mild cognitive impairment and alzheimer’s dementia,” Alzheimer’s research & therapy, vol. 15, no. 1, p. 133, 2023.
[94]
M. Torabinikjeh, V. Asayesh, M. Dehghani, A. Kouchakzadeh, H. Marhamati, and S. Gharibzadeh, “Correlations of frontal resting-state EEG markers with MMSE scores in patients with alzheimer’s disease,” The Egyptian Journal of Neurology, Psychiatry and Neurosurgery, vol. 58, no. 1, p. 31, 2022.
[95]
B. Güntekin, D. D. Emek-Savaş, P. Kurt, G. G. Yener, and E. Başar, “Beta oscillatory responses in healthy subjects and subjects with mild cognitive impairment,” NeuroImage: Clinical, vol. 3, pp. 39–46, 2013.
[96]
M. A. Klados, C. Styliadis, C. A. Frantzidis, E. Paraskevopoulos, and P. D. Bamidis, “Beta-band functional connectivity is reorganized in mild cognitive impairment after combined computerized physical and cognitive training,” Frontiers in neuroscience, vol. 10, p. 55, 2016.
[97]
R. Schmidt, M. H. Ruiz, B. E. Kilavik, M. Lundqvist, P. A. Starr, and A. R. Aron, “Beta oscillations in working memory, executive control of movement and thought, and sensorimotor function,” Journal of Neuroscience, vol. 39, no. 42, pp. 8231–8238, 2019.
[98]
H. J. Choong, E. Kim, and F. He, “Evaluating brain electroencephalogram signal dynamics across cognitive disorders using information geometry,” PLOS Complex Systems, vol. 2, no. 7, p. e0000059, 2025.
[99]
K. Baik et al., “Implication of EEG theta/alpha and theta/beta ratio in alzheimer’s and lewy body disease,” Scientific Reports, vol. 12, no. 1, p. 18706, 2022.
[100]
K. Kudo et al., “Neurophysiological trajectories in alzheimer’s disease progression,” Elife, vol. 12, p. RP91044, 2024.
[101]
U. Smailovic and V. Jelic, “Neurophysiological markers of alzheimer’s disease: Quantitative EEG approach,” Neurology and therapy, vol. 8, no. Suppl 2, pp. 37–55, 2019.
[102]
S. Iqbal, H. Nisar, and K. H. Yeap, “Quantitative EEG signatures of power and functional connectivity alterations in alzheimer’s disease and frontotemporal dementia,” Scientific Reports, vol. 16, no. 1, p. 12158, 2026.
[103]
J. Liu et al., “Multimodal diagnosis of alzheimer’s disease based on resting-state electroencephalography and structural magnetic resonance imaging,” Frontiers in Physiology, vol. 16, p. 1515881, 2025.
[104]
A. K. Paduvilan et al., “Attention-driven hybrid deep learning and SVM model for early alzheimer’s diagnosis using neuroimaging fusion,” BMC Medical Informatics and Decision Making, vol. 25, no. 1, p. 219, 2025.
[105]
Y. Chen et al., “Toward robust early detection of alzheimer’s disease via an integrated multimodal learning approach,” in ICASSP 2025-2025 IEEE international conference on acoustics, speech and signal processing (ICASSP), 2025, pp. 1–5.
[106]
J. Venugopalan, L. Tong, H. R. Hassanzadeh, and M. D. Wang, “Multimodal deep learning models for early detection of alzheimer’s disease stage,” Scientific reports, vol. 11, no. 1, p. 3254, 2021.
[107]
Z. Zhang et al., “DEG-BRIN-GCN: Interpretable graph convolutional framework with differentially expressed genes brain region interaction network prior for AD diagnosis,” Frontiers in Neuroscience, vol. 19, p. 1697528, 2025.
[108]
X. Xie et al., “Multi-modal fusion with supervised contrastive learning model for early alzheimer’s disease diagnosis and multi-modal biomarker identification,” Interdisciplinary Sciences: Computational Life Sciences, pp. 1–16, 2026.
[109]
Y. Sun et al., “Structure–function coupling reveals the brain hierarchical structure dysfunction in alzheimer’s disease: A multicenter study,” Alzheimer’s & Dementia, vol. 20, no. 9, pp. 6305–6315, 2024.

  1. Corresponding author. Email: feng.liu@rutgers.edu↩︎