June 29, 2026
Spiking neural networks (SNNs) have emerged as a promising paradigm for energy-efficient brain-inspired computing. However, existing online unsupervised SNN learning suffers from low training accuracy and poor scalability. Although current online supervised learning algorithms perform well on large-scale datasets and networks, the non-hardware-friendly operations hinder efficient edge deployment. In this work, we propose SpikON, the first algorithm-hardware co-design framework for efficient and scalable end-to-end online supervised SNN learning. We first propose the learnable threshold through time and scaled weight centralization through time techniques to address the inefficiency of traditional algorithms. Moreover, to reduce latency and energy consumption, we introduce the novel training dataflow and cascade computation reuse scheme for SNNs that allows concurrent forward-backward computation and temporal reuse across timesteps. We further design the dedicated SNN accelerator with a dual-parallel engine and customized SIMD-based SNN core for efficient end-to-end online learning. Experiments show that the SpikON algorithm achieves 32.2% and 35.0% reductions in training latency and energy consumption over the baseline, without sacrificing accuracy. Moreover, the SpikON co-design achieves 7.2x (11.5x) and 26.8x (15.8x) training throughput (energy efficiency) compared with the edge Apple M4 GPU and TPU-like accelerator, respectively. The code is available at GitHub.
<ccs2012> <concept> <concept_id>10010583.10010633.10010640</concept_id> <concept_desc>Hardware Application-specific VLSI designs</concept_desc> <concept_significance>500</concept_significance> </concept> </ccs2012>
Spiking neural networks (SNNs) emerge as a bio-inspired computing paradigm, bridging the gap between artificial neural networks (ANNs) and human brains [1]. SNNs mimic the dynamics of biological neurons and utilize binary spikes to communicate between layers [2], [3]. Due to the inherent sparsity, SNNs are regarded as more energy-efficient than ANNs and are widely used in low-power edge computing [4]–[6]. Meanwhile, backpropagation through time (BPTT) with surrogate gradient enables the training of deep SNNs that achieve accuracy comparable to ANN counterparts [7], [8].
However, BPTT suffers from significant memory cost during training, since it requires storing intermediate data at all timesteps [2], [9], [10]. BPTT is also inconsistent with online learning, where gradients are computed immediately at each timestep [2], [7], [9], [11]. Previous work modifies BPTT to enable forward-in-time SNN learning by tracking presynaptic activities [2], [7] or ignoring unimportant temporal gradients [9]. Despite their success on large-scale datasets, we identify three main challenges in deploying scalable online supervised SNN learning algorithms on edge devices.
(1) Non-hardware-friendly normalization. In online SNN learning scenarios, training samples are provided to the network sequentially (one single sample at a time) [7], [11], [12]. However, batch normalization (BN) [13] becomes ineffective with batch size \({=}\) 1, as it relies on large batches to estimate the dataset statistics [14]. To solve this issue, prior work proposes the scaled weight standardization (sWS) to replace BN [2], [7], [9], [14]. Although online SNN training with sWS can achieve comparable accuracy with BPTT, it introduces \(2.0\times\) and \(2.4\times\) higher training time and energy on average (Fig. 1 (a) and (b)) compared to the vanilla network (without BN and sWS). (2) Inherent multi-timestep computation. Compared to ANNs that process inputs in a single shot, SNNs require computation over multiple timesteps during both forward and backward passes [15]. This inherent temporal property of SNNs increases both training latency and energy consumption due to the repetitive computations across timesteps. Moreover, the overall training overhead of SNNs scales proportionally with the number of timesteps [16]. (3) Lack of efficient hardware support. Most online learning systems are designed for the unsupervised spike timing-dependent plasticity (STDP) algorithm. However, STDP suffers from low training accuracy and poor scalability for large networks and datasets [11], [17]–[19]. Although one recent work [20] designs hardware for its proposed online supervised SNN learning algorithm, it remains constrained to small-scale datasets such as MNIST [21].
To address the aforementioned challenges, we develop an algorithm-hardware co-design framework named SpikON, featuring high-throughput and energy-efficient training. To the best of our knowledge, SpikON is the first system that enables efficient and scalable online supervised SNN learning. Specifically, the main contributions are summarized as follows:
We design the dedicated accelerator to support the end-to-end online SNN training, with a dual-parallel learning engine and customized SIMD-based SNN core.
We propose the hardware-friendly learnable threshold through time (LTTT) and scaled weight centralization through time (sWCTT) approaches to overcome the inefficiency of sWS, without sacrificing training accuracy.
We propose the bi-directional temporal parallel (BTP) training dataflow that enables the concurrent forward and backward computation across all timesteps, effectively reducing SNN training latency.
We introduce the cascade temporal computation reuse (CTCR) scheme that utilizes the spike activation similarity between adjacent timesteps to reuse partial results, thus reducing redundant operations and training energy.
BPTT and online SNN learning follow the same forward propagation flow (Fig. 2 (a)), which is implemented based on leaky integrate-and-fire (LIF) neurons [22] modeled with three stages: charging (\(u^{l+1}_t\)), firing (\(s^{l+1}_t\)), and resetting (\(v^{l+1}_t\)). However, BPTT backpropagates gradients through the temporal and spatial directions, as shown in Fig. 2 (b). This process can be formulated as follows: \[\frac{\partial L}{\partial w^l} = \sum_{t=0}^{T-1} \sum_{j \leq t} \frac{\partial L}{\partial w^l_{j}},with\text{ }w^{l}_{j}=w^{l}\;\forall j\text{},\] where \(T\) is the timestep. When \(j\) is equal to \(t{-}1\), we obtain \[{!}{ \frac{\partial L}{\partial w^l_{t{-}1}}{=}\smash\{\frac{\partial L}{\partial s^{l+1}_{t-1}}\textcolor{red}{\frac{\partial s^{l+1}_{t-1}}{\partial u^{l+1}_{t-1}}}{+}\frac{\partial L}{\partial u^{l+1}_t} \textcolor{lightgreen}{\frac{\partial u^{l+1}_t}{\partial v^{l+1}_{t-1}}}( \frac{\partial v^{l+1}_{t-1}}{\partial u^{l+1}_{t-1}}{+}\textcolor{lightblue}{\frac{\partial v^{l+1}_{t-1}}{\partial s^{l+1}_{t-1}}}\textcolor{red}{\frac{\partial s^{l+1}_{t-1}}{\partial u^{l+1}_{t-1}}})\smash\}\frac{\partial u^{l+1}_{t-1}}{\partial w^{l}_{t-1}},}\] where \(\textcolor{lightgreen}{\frac{\partial u^{l+1}_t}{\partial v^{l+1}_{t-1}}}\in(0,1)\) is the leaky constant \(\beta\). Note that a surrogate function [1], [23] is used to approximate the non-differentiable \(\textcolor{red}{\frac{\partial s^{l+1}_{t-1}}{\partial u^{l+1}_{t-1}}}\). We can observe that BPTT cannot compute the gradient instantly at each timestep. To solve this issue, prior work [9] shows that temporal gradients are unimportant during backpropagation and proposes the spatial learning through time (SLTT) method. Therefore, weight gradients can be calculated by \(\frac{\partial L}{\partial w^l}{=}\sum_{t=0}^{T-1} \frac{\partial L}{\partial s^{l+1}_{t-1}}\\\textcolor{red}{\frac{\partial s^{l+1}_{t-1}}{\partial u^{l+1}_{t-1}}}\frac{\partial u^{l+1}_{t-1}}{\partial w^{l}}\) (Fig. 2 (c)). In this way, gradients at each timestep are computed independently without relying on \(\textcolor{lightgreen}{\frac{\partial u^{l+1}_t}{\partial v^{l+1}_{t-1}}}\), which aligns with the property of online SNN learning.
BN is introduced to handle the internal covariate shift issue in deep network training [13]. However, BN fails to work well when applied to online SNN learning (Fig. 1 (a)). To overcome this limitation, sWS [9], [14] is proposed to replace BN by normalizing weights rather than activations. The sWS is defined as: \[\hat{w}_{i,j} = \gamma \cdot \frac{w_{i,j} - \mu_{w_{i, .}}}{\sigma_{w_{i, .}} \sqrt{N}}, \label{equation2}\tag{1}\] where \(w_{i,j}\), \(\hat{w}_{i,j}\), \(N\), and \(\gamma\) denote the original, standardized weights, fan-in size, and fixed scale value, respectively. The mean (\(\mu_{w_{i, .}}\)) and variance (\(\sigma_{w_{i, .}}\)) are computed along the input channel dimension. Although previous work (e.g., SLTT) demonstrates that sWS can improve training accuracy under batch size \({=}1\), it introduces substantial latency and energy overheads (Fig. 1). This motivates us to design a more efficient algorithm tailored for online SNN learning.
Prior online SNN training algorithms (e.g., SLTT) directly adopt sWS, which was originally designed for deep ANNs, to replace BN. However, sWS overlooks the unique LIF neuronal dynamics and temporal dependencies of SNNs. This motivates us to rethink normalization for SNNs: can we leverage the intrinsic characteristics of SNNs to achieve normalization-free and efficient online learning? To this end, we propose two complementary approaches on SNN activations and weights, namely LTTT and sWCTT.
Unlike previous SNN works [4], [24]–[26] that use a fixed or dynamic threshold parameter (\(\theta\)) shared across all timesteps, LTTT assigns each timestep in the LIF neuron an independent learnable \(\theta\) (Fig. 3). This design approximates activation normalization (e.g., BN) by controlling spike firing and balancing membrane potential through time.
As illustrated in Fig. 3, a set of \(T\) learnable thresholds (\(\theta\) for \(t{=}0,\dots,T{-}1\)) is introduced in the LIF neuron. The forward computation can be formulated as follows: \[\begin{align} \textit{firing: }s_{t}^{l+1} =H(u_t^{l+1}{-}\theta_t^{l+1})&= \begin{cases} 1,&\text{if } u_t^{l+1} \geq \theta_t^{l+1} \\ 0,&\text{otherwise}, \end{cases} \tag{2} \\ \textit{charging: }u_t^{l+1} &= \beta v_{t-1}^{l+1} + w^{l} s_t^l, \tag{3} \\ \textit{resetting: }v_t^{l+1} &= u_t^{l+1} - s_{t}^{l+1}\theta_t^{l+1} \tag{4} , \end{align}\] where \(\theta_t^{l+1}\) is the learnable threshold of layer \(l{+}1\) at timestep \(t\), and \(H\) denotes the Heaviside step function. Each timestep in LTTT utilizes its own \(\theta_t^{l+1}\), enabling temporal adaptation during training. The threshold updates can be calculated as: \[\begin{align} \nabla\theta_{t}^{l+1} &{=} \frac{\partial L}{\partial s_t^{l+1}}\frac{\partial s_t^{l+1}}{\partial \theta_{t}^{l+1}}{=}{-}\frac{\partial L}{\partial s_t^{l+1}}H^{'}(u_t^{l+1}{-}\theta_t^{l+1}), \label{equation6}\\ \theta_t^{l+1} &= \theta_t^{l+1} {-} \eta\nabla\theta_{t}^{l+1}, \end{align}\tag{5}\] where \(\eta\) represents the learning rate. Due to the non-differentiability of \(H\), we adopt triangle surrogate function to approximate \(H^{'}(u_t^{l+1}{-}\theta_t^{l+1})\). Moreover, equation (5 ) focuses on the backward pass 1⃝ in Fig. 3 because the temporal gradient is ignored in online learning. Note that the overhead introduced by LTTT is minimal, as it only adds \(T\) trainable parameters per LIF neuron, and the required \(\frac{\partial L}{\partial s_t^{l+1}}\) can be directly reused from the weight update process. Therefore, both the parameter and computation costs are negligible.
Network parameters \(w^l\), \(\theta^l_t\), \(\alpha^l_t\); Total layer number \(N\); SNN timestep \(T\); Learning rate \(\eta\); Other required training data. Initialize: \(\nabla w^l\)=0. Forward pass: Calculate \(u^l_t\), \(v^l_t\), and \(s^l_t\) using equation (3 ), (4 ), (2 ) (weights are rescaled by equation (6 )); Backward pass: Calculate the instantaneous loss \(L\); Calculate instantaneous gradients \(\nabla w^l_t\), \(\nabla \theta^l_t\), and \(\nabla \alpha^l_t\) by equation (7 ), (5 ), and (8 ); Weight gradient accumulation: \(\nabla w^l\)=\(\nabla w^l\) + \(\nabla w^l_t\); Parameter update: \(w^l\) = \(w^l\) \(-\) \(\eta\) \(\nabla w^l\), \(\theta^l_t\) = \(\theta^l_t\) \(-\) \(\eta\) \(\nabla \theta^l_t\), \(\alpha^l_t\) = \(\alpha^l_t\) \(-\) \(\eta\) \(\nabla \alpha^l_t\); Trained SNN parameters \(w^l\), \(\theta^l_t\), \(\alpha^l_t\).
Similar to how LTTT approximates activation normalization, sWCTT aims to achieve the effect of weight normalization. Specifically, the weights are first centralized by subtracting their mean value to make the distribution zero-centered. Moreover, sWCTT introduces a learnable scaling factor \(\alpha_t^l\) for each timestep to rescale the centralized weights \(\overline{w}^l\) (Fig. 3). This mechanism regulates the weight magnitude over time and stabilizes the weight distribution during training. The sWCTT can be formulated as follows: \[\begin{align} \textit{forward: }\overline{w}^l_{t,i,j}&{=}w^l_{t,i,j}{-}\mu_{w^l_{t,i,.}},\hat{w}^l_t{=}\alpha^l_t\overline{w}^l_{t},w^{l}_{t}{=}w^{l}\;\forall t \tag{6}\\ \textit{backward: }\frac{\partial L}{\partial w^l_{t,i,j}}&{=}\alpha^l_t(g^l_{t,i,j}{-}\mu_{g^l_{t,i,.}}), \tag{7} \\ \frac{\partial L}{\partial \alpha^l_{t}}&{=}\sum_{i,j}g^l_{t,i,j}\overline{w}^l_{t,i,j},\textit{ with } g^l_{t}{=}\frac{\partial L}{\partial \hat{w}^l_t}, \tag{8} \end{align}\] where \(w^l_t\), \(\overline{w}^l_{t}\), and \(\hat{w}^l_t\) denote the original, centralized, and scaled weights, respectively, and \(g^l_t\) is the scaled weight gradient. The \(i\) and \(j\) represent the output and input neuron indices, respectively. Different from sWS that applies a fixed pre-computed \(\gamma\) (equation (1 )), sWCTT assigns an independent learnable \(\alpha^l_t\) to each timestep, allowing more flexible weight adjustment over time. Moreover, sWCTT eliminates the need to calculate variance and simplifies the weight gradient computation, making it more hardware-friendly for efficient online SNN learning.
Algorithm [algorithm1] outlines the overall training process of SpikON, which combines the proposed LTTT and sWCTT for efficient online SNN learning. Specifically, during the forward pass, the timestep-dependent \(\theta^l_t\) controls the update of potentials and spike states (Line 4). Once the forward propagation at a given timestep is completed, the backward pass computes the instantaneous loss (\(L\)) and corresponding gradients \(\nabla w^l_t\), \(\nabla \theta^l_t\), and \(\nabla \alpha^l_t\) (Line 7, 8). The gradients obtained from all timesteps are subsequently accumulated and applied to update the network parameters (Line 9, 12).
The proposed SpikON accelerator is designed to support efficient and scalable online SNN learning. As shown in Fig. 4, SpikON consists of five components: bi-temporal parallel engine (BTPE), customized SIMD-based SNN core, global SRAM, memory controller, and top controller. The BTPE performs the vector-matrix multiplications (VMM) in both forward and backward passes, including \(\nabla w\) and error signal computations. It supports bi-directional temporal parallel (BTP) dataflow and cascade temporal computation reuse (CTCR) scheme to reduce training latency and energy. Moreover, the SIMD-based SNN core executes element-wise and vector-wise operations, such as average pooling, softmax, and LIF neuron update. The memory controller arbitrates concurrent accesses from the SNN core and BTPE to the global SRAM, while the top controller manages the overall operation by controlling start signals for their coordinated execution.
BTP training dataflow. We design the BTP dataflow to improve online SNN training throughput while effectively handling temporal dependencies. Unlike traditional training flow that processes each timestep’s forward and backward pass in order, the BTP dataflow enables concurrent forward-backward computation across all timesteps (Fig. 5 (a)). Specifically, in the forward pass, computations of adjacent timesteps are interleaved to exploit temporal parallelism. For example, the \(Conv.0\) layer of T0 starts first, and the \(Conv.0\) layer of T1 begins when T0 proceeds to \(LIF0_{t=0}\) layer. Once the potential \(v^0_0\) in \(LIF0_{t=0}\) is updated, it can be utilized by the next timestep, allowing T0’s \(Conv.1\), \(LIF0_{t=1}\), and T2’s \(Conv.0\) to execute concurrently. This staggered execution continues across timesteps, achieving high forward parallelism. Different from the forward process, the backward pass in the SpikON algorithm does not involve temporal error propagation across timesteps (Section 3). Thus, all timesteps can be executed independently, enabling fully parallel backward computation.
Dual-parallel learning engine. To support the proposed BTP training dataflow, we design the specialized BTPE module, as shown in the right part of Fig. 4. The BTPE consists of 24 BTP-dataflow lanes, buffer, scheduler, data rearranger, and 23 local SRAMs. Each lane in BTPE is assigned to process one timestep of SNN, supporting up to 24 timesteps in parallel. When \(T\) is smaller than 24, multiple lanes can be grouped to handle the same timestep. For example, for SNN with \(T=6\), lanes 0, 6, 12, and 18 are used to process \(T0\). The local scheduler uses the finite state machine to control data loading, distribution, result write-back, etc. Moreover, local SRAMs are utilized to enable the CTCR method by storing lane-level intermediate results for reuse (Section 4.3).
BTP-dataflow lane. To better leverage the inherent sparsity of SNNs, the BTP-dataflow lane (Fig. 5 (b)) in the BTPE is implemented to support both sparse and dense computation modes during training. Specifically, each lane consists of input buffers A and B, data pre-processor, four processing unit (PU) arrays, output aggregator, and local controller. The PU array contains \(4\times 4\) PUs that adopt an output-stationary dataflow, where each PU computes one element of the output feature map, allowing for 64-element parallel processing per lane. The output aggregator (OA) accumulates the partial results of current PU arrays with outputs from the previous timestep. Moreover, the local controller coordinates data access to local SRAMs within BTPE and generates the \(done\) signal when PU arrays and OA finish computation.
Processing unit. As shown in Fig. 5 (b), PU consists of FP32 multiplier and adder, partial sum (PS) and counting (CT) registers, control logic, etc. It operates in two computation modes: sparse and dense. During the sparse mode, the \(mode\) signal disables the dense path while activating the sparse path. The multiplexer selects among \(0\), \(-b\), or \(+b\) for computing potential \(u\) and \(\nabla w\). Note that \(-b\) indicates the negative number utilized in the CTCR method. In contrast, during the dense mode, the \(mode\) signal enables accumulation using results from the dense path for error signal computation. Once the processed element pairs reach the preset \(data\) \(volume\), the PU sets \(done=1\) to indicate computation completion.
| Dataset | Method | Network | Params | Timesteps | Accuracy | Training Time/Epoch (s) | Training Energy/Epoch (J) |
|---|---|---|---|---|---|---|---|
| CIFAR-10 | OTTT | VGG11 (sWS) | 9.23M | 6 | 92.47% | 2,829.10 | 525,395.45 |
| SLTT | VGG11 (BN) | 9.23M | 6 | 74.30% | 1,219.80 | 189,402.19 | |
| SLTT | VGG11 (sWS) | 9.23M | 6 | 92.47% | 2,048.03 | 397,490.18 | |
| SpikON (ours) | VGG11 (LTTT) | 9.23M | 6 | 90.88% | 1,028.20 | 163,299.09 | |
| SpikON (ours) | VGG11 (LTTT+sWCTT) | 9.23M | 6 | 92.24% | 1,395.98 | 245,739.04 | |
| OTTT | VGG11 (sWS) | 9.27M | 6 | 70.16% | 2,886.71 | 524,567.27 | |
| SLTT | VGG11 (BN) | 9.27M | 6 | 29.67% | 1,223.17 | 185,793.66 | |
| SLTT | VGG11 (sWS) | 9.27M | 6 | 70.29% | 2,095.40 | 401,829.07 | |
| SpikON (ours) | VGG11 (LTTT) | 9.27M | 6 | 56.68% | 1,052.81 | 164,892.59 | |
| SpikON (ours) | VGG11 (LTTT+sWCTT) | 9.27M | 6 | 69.84% | 1,387.75 | 239,152.99 | |
| DVS-CIFAR10 | OTTT | VGG11 (sWS) | 9.23M | 10 | 76.80% | 962.6 | 180,017.09 |
| SLTT | VGG11 (BN) | 9.23M | 10 | 42.80% | 429.95 | 71,574.49 | |
| SLTT | VGG11 (sWS) | 9.23M | 10 | 78.70% | 707.94 | 128,350.70 | |
| SpikON (ours) | VGG11 (LTTT) | 9.23M | 10 | 76.10% | 387 | 61,699.13 | |
| SpikON (ours) | VGG11 (LTTT+sWCTT) | 9.23M | 10 | 79.00% | 496.1 | 86,394.05 | |
| OTTT | VGG11 (sWS) | 9.23M | 20 | 97.22% | 348.15 | 72,046.33 | |
| SLTT | VGG11 (BN) | 9.23M | 20 | 33.33% | 135.57 | 31,873.54 | |
| SLTT | VGG11 (sWS) | 9.23M | 20 | 97.92% | 208.54 | 47,133.58 | |
| SLTT | VGG11 (sWS) | 9.23M | 10 | 96.88% | 120.74 | 25,468.13 | |
| SpikON (ours) | VGG11 (LTTT) | 9.23M | 20 | 97.57% | 120.63 | 28,646.23 | |
| SpikON (ours) | VGG11 (LTTT+sWCTT) | 9.23M | 20 | 97.22% | 156.02 | 36,603.03 | |
| SpikON (ours) | VGG11 (LTTT+sWCTT) | 9.23M | 10 | 97.57% | 80.84 | 18,203.50 |
Spike similarity. We compare the spike differences between adjacent timesteps in the forward pass by calculating the cosine similarity (Fig. 6). For CIFAR-100, all layers exhibit consistently high cosine similarity (\({>}80\%\)), as the dataset is static and each timestep receives identical input. In contrast, the inputs vary dynamically across timesteps for DVS-CIFAR10, leading to lower similarity in certain layers. However, multiple layers (e.g., layers 2, 4, 6, 7, and 8) still preserve more than \(55\%\) similarity across timesteps.
CTCR strategy. Owing to the high temporal similarity of spike activations, we propose the CTCR method to reuse intermediate results across adjacent timesteps (green circles and arrows in Fig. 5 (a)). Specifically, the convolution result of \(s^l_t\) with \(w^l\) (denoted as \(s^l_t{\ast} w^l\)) can be reused for computing \(s^l_{t+1}{\ast} w^l\), where only the difference term \((s^l_{t+1}{-}s^l_t){\ast}w^l\) needs to be processed. Since \((s^l_{t+1}{-}s^l_t)\) introduces higher sparsity than \(s^l_{t+1}\), the proposed PU (Fig. 5 (b)) can utilize this property to reduce computation energy. Note that the same reuse mechanism can also be applied to FC layers. The two theorems that support our CTCR strategy are presented below: \[\begin{align} \textit{FC layer: } y^l_t&{=}w^ls^{l-1}_t{=}w^ls^{l-1}_{t-1}{+}w^l(s^{l-1}_t{-}s^{l-1}_{t-1}),\\ \textit{Conv. layer: } y^l_t&{=}w^l{\ast}s^{l-1}_t{=}w^l{\ast}s^{l-1}_{t-1}{+}w^l{\ast}(s^{l-1}_t{-}s^{l-1}_{t-1}), \end{align}\] where \(y^l_t\) represents the intermediate result. To verify the sparsity improvement introduced by the CTCR strategy, we conduct the following theoretical analysis. \[\begin{align} &\textit{Assume the shape of s_t, s_{t+1} is (1, D), and the non-zero ratio of s_t} \\ &\textit{and s_{t+1} is m and n, respectively. Define \Delta{=}s_{t+1}{-}s_t}.\\ &\textit{Similarity: \cos(s_t, s_{t+1}){=}\frac{\|s_t\|^2{+}\|s_{t+1}\|^2{-}\|s_{t+1}{-}s_t\|^2}{2 \|s_t\| \|s_{t+1}\|}{=}c.} \\ &\textit{Given \|s_t\| {=} \sqrt{mD} and \|s_{t+1}\| {=} \sqrt{nD}, we will have}\\ &\textit{c{=}\frac{mD {+} nD {-} \|\Delta\|^2}{2D\sqrt{mn}}, \|\Delta\|^2 {=} (m {+} n)D {-} 2cD\sqrt{mn}.} \end{align}\] Therefore, the non-zero ratio of \(\Delta\) is \((m {+} n) {-} 2c\sqrt{mn}\). To guarantee the sparsity improvement of CTCR, it needs to satisfy \(\sqrt{m} {\leq} 2c\sqrt{n}\).
To efficiently support neuron operations and ensure the generality of the SpikON accelerator, we design a SIMD-based SNN core with a five-stage pipeline and a lightweight custom ISA. As shown in Fig. 5, SNN core includes the
\(1024{\times}32bit\) instruction SRAM, \(32{\times}512bit\) register file, executor with 64 FP32 units for parallel computation, controller, etc. Each FP32 unit can execute general
arithmetic (e.g., add, sub, mult, div, sqrt, max, min, etc.) and SNN-specific operations. In particular, lif and reset instructions model LIF neuron dynamics in the forward pass, while sg handles surrogate gradient
computation in the backward pass. Moreover, SIMD controller manages pipeline stalls and resolves instruction dependencies to ensure correct and efficient execution. Note that although SNN core executes instructions sequentially without parallel timestep
processing as in BTPE, its latency can be effectively hidden since BTPE dominates the overall computation.
Algorithm setup. We evaluate the proposed SpikON algorithm on both static and neuromorphic datasets, such as CIFAR-10 [27], CIFAR-100 [27], DVS-CIFAR10 [28], and DVS128-Gesture [29]. We use the VGG11 SNN architecture with a lightweight classifier (64C3-128C3-AP2-256C3-256C3-AP2-512C3-512C3-AP2-512C3-512C3-GAP-FC) for all experiments [30]. The training epoch and batch size of SNNs are set to 100 and 1, respectively. The leaky constant \(\beta\) is configured as 0.09 for our SNN models. The baselines are the prior online SNN learning algorithms, OTTT [7] and SLTT [9].
Hardware setup and baseline. We design the SpikON accelerator at RTL using Verilog. To estimate the area and power of core circuits, we synthesize SpikON using Synopsys Design Compiler [31] at 1.5 GHz under the TSMC 28nm standard cell library. The on-chip SRAMs are generated by the TSMC memory compiler. The energy consumption of off-chip memory HBM2 (5.7pJ/b) is adopted from [32]. In addition, we implement a simulator to evaluate the performance of the SpikON accelerator. SpikON is benchmarked against the edge Apple M4 GPU (10 cores) [33] and server Nvidia A40 GPU [34]. We also build a TPU-like training accelerator with 1521 floating-point PEs (systolic array replacing BTPE module in Fig. 5) as a baseline and evaluate it under our simulator framework.
Comparison with OTTT and SLTT. (1) Accuracy. As shown in Table 1, SpikON (LTTT-only) exhibits an average accuracy drop of 4.54% compared to the SLTT (sWS) baseline across all datasets. However, by further incorporating the sWCTT strategy, SpikON (LTTT+ sWCTT) even surpasses the SLTT (sWS), achieving an average accuracy improvement of 0.08%. (2) Latency. We can observe that SLTT (sWS) achieves a \(30.4\%\) reduction on average in training latency over the OTTT (sWS). Moreover, our SpikON (LTTT+sWCTT) further reduces the training latency by 32.2% compared to SLTT (sWS), without sacrificing accuracy. (3) Energy. In terms of training energy, SLTT (sWS) consumes \(27.8\%\) less energy on average than OTTT (sWS). Benefiting from the normalization-free online learning, SpikON (LTTT+sWCTT) achieves a further \(35.0\%\) reduction in energy consumption relative to SLTT (sWS).
Sparsity improvement achieved by CTCR. To verify the effectiveness of CTCR, we evaluate the improved sparsity between adjacent timesteps across all layers on CIFAR-100 and DVS-CIFAR10 during training (Fig. 8). We can observe that CTCR consistently enhances the sparsity of all layers for the static CIFAR-100. In particular, the improved sparsity of \(layer0\) reaches 100% since it receives the same input at each timestep. Moreover, most layers (e.g., 0, 2, 3, 6, 7, 8) exhibit increased sparsity under CTCR on the DVS-CIFAR10, with only minor reductions observed in a few layers. Note that the computation overhead of CTCR is negligible, as it only involves subtracting activations between adjacent timesteps.
Area and power breakdown. Fig. 9 presents the area and power breakdown of the SpikON accelerator. The total area and power consumption are \(35.096\) \(mm^2\) and \(11.861\) \(W\), respectively. (1) Area. The SNN core and BTP engine occupy the majority of the chip area (32.4% and 62.1%), while the controller contributes only 0.1% of the total area. Within the BTP engine, most of the space is allocated to the 24 BTP-dataflow lanes and 23 local SRAMs, taking up 22.7% and 38.3% respectively. (2) Power. The BTP engine dominates the overall power, accounting for 83.6% of the total power. The SNN core and global SRAM consume only 3.1% and 7.4%, respectively.
Comparison with edge/server GPUs and TPU-like accelerator. We use Apple’s \(powermetrics\) tool [35] and \(nvidia\)-\(smi\) [36] to measure the power of the M4 and A40 GPUs at runtime, respectively. (1) SpikON algorithm+M4 GPU. With the proposed SpikON algorithm, the online SNN training throughput and energy efficiency of the edge M4 GPU are improved by 1.7x and 1.9x, respectively (Fig. 7). (2) SpikON co-design v.s. GPUs. With both the SpikON algorithm and accelerator, SpikON achieves 7.2x throughput and 11.5x energy efficiency on average over the M4 GPU (Fig. 7). Even compared with the server A40 GPU, SpikON still delivers 1.4x and 42.7x improvement in throughput and energy efficiency, respectively. Note that on the DVS128-Gesture dataset, the throughput of SpikON co-design (32.8x) is lower than A40 GPU (94.6x), mainly due to the large 128\(\times\)128 input feature map and the limited parallelism of each lane. However, SpikON still demonstrates a 14.4x advantage in energy efficiency over the A40 GPU. (3) SpikON co-design v.s. TPU-like accelerator. Due to the lack of an existing SOTA ASIC baseline for online supervised SNN learning, we implement a TPU-like training accelerator and keep all hardware configurations with the same number of PEs for fair comparison. As shown in Fig. 7, we observe that SpikON co-design achieves 26.8x higher throughput and 15.8x higher energy efficiency on average.
Ablation study of SpikON accelerator. We conduct an ablation study to demonstrate the advantages of the proposed methods (Fig. 10). (1) BTP training dataflow. As shown in Fig. 10 (a), with the BTP dataflow, we can achieve an average 3.8x speedup in training throughput across all datasets. (2) Sparsity-aware PU+CTCR. Fig. 10 (b) presents the energy reduction achieved in the forward pass using the proposed techniques. The sparsity-aware PU utilizes the inherent sparsity of SNNs to reduce energy by 56% on average. With the additional CTCR strategy, the energy consumption is further reduced to 71%. Notably, the CTCR achieves greater energy reduction on static datasets than on event-stream datasets, which is aligned with the improved sparsity results in Fig. 8.
We propose SpikON, the first algorithm-hardware co-design for efficient and scalable end-to-end online supervised SNN learning. Specifically, SpikON algorithm achieves 32.2% and 35.0% reductions in training latency and energy over baseline, without sacrificing accuracy. Moreover, SpikON co-design achieves 7.2x (11.5x) and 26.8x (15.8x) training throughput (energy efficiency) compared with the edge Apple M4 GPU and TPU-like accelerator, respectively.