July 05, 2024
Pedestrian detection in conventional frame-based imaging often suffers from limited temporal responsiveness and substantial data redundancy. Inspired by the biological retina, event-based vision sensing (EVS) offers ultra-low latency, high temporal resolution, wide dynamic range, and low power consumption, making it highly attractive for pedestrian perception in complex environments. This paper provides a comprehensive review of EVS and its application to pedestrian detection in intelligent transportation and surveillance scenarios. We first summarize the sensing principles, historical development, and key advantages of event-based vision in comparison with conventional frame-based imaging. We then review the major methodological components of event-based pedestrian detection, including sensing inputs, event representations, preprocessing strategies, feature extraction, detection models, datasets, and evaluation metrics. In addition, representative methods are comparatively analyzed in terms of temporal fidelity, detection accuracy, computational efficiency, and deployment complexity. Finally, we discuss the major open challenges in current EB-PD research, including benchmark standardization, event-native model design, multimodal fusion, and real-world deployment, and outline several promising directions for future development. This review aims to provide a structured and up-to-date reference for researchers working on event-based pedestrian perception and related intelligent vision systems.
A comprehensive review is presented on event-based vision sensing and its application to pedestrian detection in intelligent transportation and surveillance.
Existing event-based pedestrian detection studies are structured around sensing inputs, preprocessing strategies, feature representations, detection models, datasets, and evaluation metrics.
Representative methods are critically compared from the perspective of temporal fidelity, detection accuracy, computational efficiency, and deployment complexity.
The review highlights key trade-offs in event-based pedestrian detection, including direct event streams versus event-to-frame conversion, SNN-based versus CNN-based processing, and short versus long temporal accumulation windows.
Open challenges are discussed with emphasis on benchmark standardization, event-native perception models, multimodal learning, and deployment-oriented research.
event-based vision ,neuromorphic sensor ,dynamic vision sensing ,pedestrian detection ,spatio-temporal fusion ,event stream processing
Pedestrian detection is a fundamental task in computer vision. It aims to accurately identify [1]–[4] and localize [5]–[7] pedestrians in images or image sequences. As a core component of Intelligent Transportation Systems (ITS) [8]–[10], pedestrian detection plays an important role in active safety and autonomous driving [11], [12]. It is also closely related to pedestrian tracking [13]–[18], person re-identification [19]–[22], intelligent surveillance [23]–[25], human behavior analysis [26]–[29], assisted driving [30]–[32], public safety [33]–[35], and broader AI-driven perception systems [36], [37]. Existing pedestrian-detection methods include global feature-based approaches [38]–[42], body part-based methods [43]–[45], and stereo-vision-based techniques [46]–[50]. Despite substantial progress, pedestrian detection remains challenging under difficult illumination, high-speed motion, occlusion, and resource-constrained deployment scenarios [44], [51].
These challenges become even more pronounced in real-world traffic scenes. Pedestrian behavior is inherently difficult to predict because it depends on complex and dynamic factors, including motion state, intended destination, demographic attributes, and interactions with nearby vehicles and other pedestrians [52]–[60]. From an edge-computing perspective, pedestrian perception systems must also balance model complexity against computational efficiency and latency constraints [51]. These considerations motivate the search for sensing and perception paradigms that can respond more rapidly and robustly to dynamic environments.
Event cameras are asynchronous visual sensors that represent a major shift in image acquisition [61]. Their pixel design is inspired by the mammalian retina [62]. Event-based vision sensing (EVS) [63]–[66] and the Dynamic Active Pixel Vision Sensor (DAVIS) [67], [68] are among the most representative technologies in this family [69], [70]. The development of EVS can be traced back to the “Silicon Retina” introduced by Mahowald and colleagues in 1992 [71]. In 2006, Delbruck’s group introduced the Dynamic Vision Sensor (DVS) [64], which marked a major milestone in the maturation of event-camera technology. This was followed by the introduction of DAVIS in 2013 [72] and the DAVIS346 color variant in 2017 [67]. In 2021, Prophesee and Sony released the EVK4 platform equipped with the IMX636 event sensor. In 2025, DVSense launched the DVSync series, representing a recent attempt to integrate high-resolution RGB and EVS sensing within a unified hardware system.
The biological inspiration of EVS lies in several functional mechanisms of the retina, as illustrated in Fig. 1. In the biological retina, ON and OFF bipolar cells respond to increasing and decreasing light intensity, respectively [61], [62]. Event cameras mimic this principle by generating events when local log-intensity changes exceed a threshold, thereby emphasizing contrast changes rather than recording redundant static information [73]–[75]. In addition, retinal mechanisms related to contrast enhancement and information transmission have inspired the design of event-based sensing circuits, which encode scene dynamics with precise timestamps and polarity information [71]–[73], [75]–[77].
Each event records the pixel location, timestamp, and polarity of a brightness change [63], [72]. Because events are generated only when changes occur, event cameras offer several distinctive advantages over conventional frame-based imaging. First, they capture luminance variations with microsecond-level temporal resolution, making them highly suitable for fast motion without conventional motion blur [61], [78]–[80]. Second, their sparse output greatly reduces redundant data transmission and lowers power consumption [61]. Third, their high dynamic range (HDR) enables robust sensing under challenging illumination conditions [81]. For pedestrian detection, these properties are particularly attractive because they support faster response to environmental changes [11], [82], more reliable operation under low light [83], and improved perception of rapidly moving pedestrians [84].
Against this background, this paper reviews the field of event-based pedestrian detection (EB-PD), with emphasis on applications in autonomous driving, intelligent transportation, and human-centered visual perception. We examine the available datasets, sensing settings, event representations, preprocessing pipelines, feature-extraction strategies, and detection models used in this domain. Existing methods are organized according to their sensing and representation paradigms, including direct event-stream processing, event-to-frame conversion, and joint event-frame fusion, and their strengths and limitations are analyzed accordingly.
The primary contributions of this review are threefold. First, we provide a structured analysis of the EB-PD literature and organize existing methods into three representative processing paradigms: direct event-stream processing, event-to-frame conversion, and joint event-frame fusion, while discussing their respective advantages and limitations. Second, we summarize representative public datasets and commonly used evaluation metrics, and we compare major EB-PD-related tasks and methods across widely used benchmarks. Third, we identify current research gaps, major practical challenges, and several promising directions for future development in event-based pedestrian detection.
Given the expansive definition [61] of event-based sensors and the wide variety of related devices [85], a rigorous and unified system remains undeveloped [86]. Before presenting the search process, we therefore clarify the scope of this review. In this paper, the term “EBC” is broadly used to denote vision sensors capable of capturing dynamic event signals, including technologies referred to as DVS, DAVIS, EVS, and related event-based cameras. Meanwhile, the EB-PD task considered here primarily focuses on pedestrian-related perception for intelligent transportation and active safety, while also encompassing scenario-specific tasks such as posture identification, static pedestrian detection, and road-target segmentation.
Our methodology for locating relevant literature involved both direct searching and the snowballing technique. We utilized platforms such as IEEE Xplore digital library, Springer, ScienceDirect, ACM, and Google Scholar for direct search to include both scientific databases and open-access pre-prints.1 Table 1 summarizes the search terms used on each platform and the number of records retrieved. The search covered publications from 2014 to early 2026 and identified 353 relevant articles.
| Database | # Records | Query |
|---|---|---|
| IEEE | 72 | “All Metadata”: (DVS AND Pedestrian Detection) OR (EVS AND Pedestrian Detection) OR (Event based AND Pedestrian Detection) OR (Neuromorphic Vision AND Pedestrian Detection) |
| Elsevier | 19 | pub-date >2014 and (DVS AND Pedestrian Detection) OR (EVS AND Pedestrian Detection) OR (Event based AND Pedestrian Detection) OR (Neuromorphic Vision AND Pedestrian Detection) |
| Springer | 32 | (DVS AND Pedestrian Detection) OR (EVS AND Pedestrian Detection) OR (Event based AND Pedestrian Detection) OR (Neuromorphic Vision Pedestrian Detection) “within 2014 - 2026” |
| ACM | 27 | [[All Metadata: “DVS AND Pedestrian Detection”]] OR [[All Metadata: “EVS AND Pedestrian Detection”]] OR [[All Metadata: “Event based AND Pedestrian Detection”]] OR [[All Metadata: “Neuromorphic vision AND Pedestrian Detection”]] AND [E-Publication Date: (01/01/2014 TO 12/31/2024)] |
| Google Scholar | 193 | “DVS AND Pedestrian Detection” OR “Event based Pedestrian Detection” OR “Neuromorphic Vision AND Pedestrian Detection” OR “EVS AND Pedestrian Detection” custom range 2014 - present |
The primary objective of our inclusion and exclusion criteria was to identify recent and relevant peer-reviewed studies focused on the EB-PD task. Because search syntax differs across databases, the query expressions were adapted to each platform. The main requirement was that the search terms be explicitly mentioned within the articles’ titles, abstracts, or keywords. In the initial screening, we applied the following preliminary exclusion criteria:
F1. Exclude review and survey papers.
F2. Exclude publications that are neither conference nor journal articles, including master’s and doctoral theses.
F3. Exclude papers that mention EB-PD but do not focus on EB-PD tasks (such as image segmentation).
F4. Exclude short papers and duplicates.
F5. Exclude papers not written in English.
Following this preliminary screening, we identified 50 papers across various platforms for inclusion. The distribution of these papers from 2014 to 2026 is detailed in Fig. 2.
Upon reviewing the corpus, it was determined that seven papers explicitly contribute to dataset development, focusing on the construction of comprehensive datasets for EB-PD. A breakdown of the affiliations reveals that 31 papers were authored primarily by academics or researchers from institutions, while 19 originate from the industry, including notable contributions from corporations like Prophesee, Samsung, and Intel. These companies have advanced the field by sharing datasets derived from their state-of-the-art event cameras under diverse scenarios. Moreover, 46.9% of the included studies were published in recognized journals and conferences, including IEEE Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), IEEE Transactions on Pattern Analysis and Machine Intelligence (PAMI), IEEE Internet of Things Journal, IEEE Transactions on Information Forensics and Security, Neurorobotics, IEEE Access, Society of Photo-Optical Instrumentation Engineers (SPIE), and IEEE International Conference on Robotics and Automation (ICRA). These observations indicate sustained research interest in EB-PD and reflect the growing recognition of its technical challenges and application value.
This section categorizes the extensive body of research into comprehensive domains, reflecting recent advancements and aligning with current technological trends.
Research in EBC technology primarily focuses on the optimal utilization of the sparse and asynchronous data [87], [88] these sensors generate. Various methodologies for data capture are detailed, focusing on minimizing redundancy and improving the relevance of event signals through advanced filtering techniques [89]. Feature extraction from EBC data involves transforming event-triggered intensity changes at specific pixels into formats suitable for pedestrian detection algorithms [88]. This process often integrates temporal and spatial data to leverage the high temporal resolution of EBC, which is crucial for dynamic environments like urban traffic scenes [90].
Integrating EBC data with detection models poses unique challenges due to non-standard data formats. Existing EB-PD studies can generally be organized into three sensing and representation paradigms [91]–[93]. Converting event streams into frame-like representations and then processing them with conventional networks is common in both static- and dynamic-scene EB-PD. Meanwhile, researchers have also experimented with combining event streams and image frames as inputs [94], [95]. The potential for mining information in the event stream signals using mature frame-based deep learning models has been a major focus of recent research. Convolutional neural networks (CNNs) [91], [93], [96]–[101] have been adapted to handle the binary and sparse nature of EBC data, and modified to accommodate temporal dynamics absent in traditional video data. These models are designed to utilize the fine-grained temporal information from EBCs to improve detection accuracy and response times in pedestrian detection applications [102], [103].
The development of dedicated datasets for EB-PD is a pressing research need, aiming to provide benchmarks that accurately reflect the operational challenges and capabilities of event-based sensors. Open source datasets dedicated to EB-PD [104]–[107] remain limited, but some researchers have conducted relevant experiments on other datasets [105], [108]–[111] that can be used for EB-PD. These datasets, although abundant in pedestrian data, do not specifically address unique scenarios or areas of particular interest for pedestrian detection. Standardized dataset collection methods and clearly defined model evaluation metrics are imperative. In addition to these datasets, evaluation metrics have been compiled to assess the performance of pedestrian detection systems under various operating conditions. These metrics address the accuracy, reliability, and computational efficiency of systems designed to process high-density event data.
Current EB-PD methods encompass a broad spectrum of paradigms, ranging from handcrafted approaches to deep learning-based and hybrid frameworks. In general, the EB-PD pipeline can be decomposed into four core stages: sensing inputs, preprocessing, feature extraction and representation, and detection models. Fig. 3 illustrates the overall workflow of these four stages. It should be noted, however, that not every method strictly follows this canonical pipeline. For example, the work in [112] applies hierarchical clustering directly to event data, whereas [113] focuses on semantic segmentation based on event-frame signals. Nevertheless, from the perspective of methodological generalizability, the four-stage formulation remains a useful and sufficiently comprehensive abstraction for analyzing existing EB-PD methods.
The initial phase concerns the acquisition and organization of sensing inputs for downstream pedestrian detection. In existing EB-PD research, this stage can generally be categorized into three representative paradigms: direct event-stream processing, event-to-frame conversion, and joint event-frame fusion. The following subsections describe these three paradigms in detail.
Direct Event-Stream Processing in EB-PD systems fully exploits the fine-grained temporal information captured by event cameras [92]. Unlike traditional cameras, which produce image frames at fixed intervals, event cameras output an asynchronous stream of events \(\mathcal{E}\), where each event \(e_i \in \mathcal{E}\) is represented as a tuple \((t_i, x_i, y_i, p_i)\):
\(t_i\) denotes the timestamp at which a change in log intensity is detected;
\((x_i, y_i)\) denote the spatial coordinates of the pixel where the event occurs;
\(p_i \in \{-1, +1\}\) denotes the polarity of the intensity variation.
This data stream is inherently sparse and asynchronous, and therefore requires dedicated preprocessing operations before it can be effectively exploited by pedestrian detection models [92], [114]–[116].
Typical preprocessing for direct event streams includes noise filtering, temporal clustering, event deduplication, and density normalization. Noise filtering suppresses sensor artifacts and environmental interference: \[\mathcal{E}_{\mathrm{filtered}} = \{ e_i \in \mathcal{E} \mid \Phi(e_i) = 1 \},\] where \(\mathcal{E}\) denotes the raw event set, \(\mathcal{E}_{\mathrm{filtered}}\) denotes the filtered event set after noise suppression, \(e_i\) denotes the \(i\)-th event, and \(\Phi(\cdot)\) is a Boolean-valued validity function that returns 1 when an event is retained and 0 otherwise. Temporal clustering then groups events according to temporal proximity in order to identify meaningful motion patterns: \[\mathcal{C}_j = \bigcup_{i}\{e_i \mid t_{i+1} - t_i \leq \Delta T\},\] where \(\Delta T\) is a predefined temporal threshold. Event deduplication further merges redundant or duplicate events: \[\mathcal{E}_{\mathrm{dedup}} = \operatorname{merge}(\{e_i, e_{i+1} \mid e_i \approx e_{i+1}\}),\] while density normalization alleviates local response imbalance across the sensor plane: \[\mathcal{E}_{\mathrm{norm}} = \frac{\mathcal{E}_{\mathrm{filtered}} - \min(\mathcal{E}_{\mathrm{filtered}})}{\max(\mathcal{E}_{\mathrm{filtered}}) - \min(\mathcal{E}_{\mathrm{filtered}})}.\] Here, \(e_i \approx e_{i+1}\) indicates that two neighboring events are regarded as redundant according to predefined spatial, temporal, and polarity consistency criteria; \(\operatorname{merge}(\cdot)\) denotes the corresponding event-merging operation; and \(\min(\cdot)\) and \(\max(\cdot)\) are applied to the selected event-derived representation for density normalization.
Through the above operations, the direct event stream is transformed into a cleaner and more informative signal representation, thereby reducing sparsity-induced instability and enhancing the delineation of dynamic patterns relevant to pedestrian detection.
After preprocessing, the event stream can be embedded into deep learning architectures to construct a multi-dimensional representation of the scene. Such representations may include temporal event distributions, localized spatio-temporal volumes, polarity statistics, spatial depth cues, and information entropy, all of which jointly characterize pedestrian motion and scene dynamics. Formally, an example feature vector at time \(t\), derived from the normalized event stream \(\mathcal{E}_{\mathrm{norm}}\), can be expressed as:
\[\mathbf{F}_t = \operatorname{encode}(\mathcal{E}_{\mathrm{norm}}(t)).\]
This high-dimensional feature representation enables downstream models to capture complex pedestrian behaviors and transient motion patterns with high precision.
Unlike event-to-frame conversion or frame-event fusion, direct event stream processing preserves the original temporal granularity of asynchronous event data. This is particularly beneficial for modeling fast motion, rapid luminance changes, and short-lived visual structures. Spiking Neural Networks (SNNs) and related event-native architectures are especially suitable in this setting, as they can directly exploit event timing information for motion-sensitive detection. In practice, this processing route typically involves suppressing spurious events, stabilizing event density across the visual field, and encoding the refined stream into a representation that preserves temporal granularity while remaining suitable for neural inference.
The primary advantage of direct event stream processing lies in its ability to handle raw, high-frequency signals without collapsing them into frame-like structures. This makes it particularly suitable for capturing subtle and transient pedestrian dynamics under high-speed motion and challenging illumination, where conventional frame-based approaches often suffer from motion blur, temporal aliasing, or data redundancy.
Event-to-frame conversion transforms asynchronous event data into structured frame-like representations that can be directly processed by conventional computer vision models. This strategy preserves part of the temporal dynamics of event data while improving compatibility with mature frame-based detection architectures. At present, this remains the most widely adopted processing paradigm in EB-PD.
The predominant event-to-frame encoding strategies can be broadly categorized into three types [117], as illustrated in Fig. 4: frequency-based event encoding [118], active-event-surface-based encoding [119], [120], and Leaky Integrate-and-Fire (LIF)-based event accumulation [121]–[123]. Despite their differences, these methods share a common objective: to convert sparse asynchronous events into structured representations that are amenable to conventional feature extraction pipelines.
\(t_i\) denotes the timestamp of each event, which determines its temporal contribution to the generated frame;
\((x_i, y_i)\) are the spatial coordinates corresponding to the activated pixel location;
\(p_i \in \{-1, +1\}\) denotes the polarity of the intensity variation.
The event-to-frame conversion process usually includes four closely related operations: event integration, polarity fusion, spatial coherence enhancement, and inter-frame continuity modeling. In the following formulations, \(I_t(x,y)\), \(P_t(x,y)\), \(S_t(x,y)\), and \(F_t(x,y)\) denote the temporally integrated event image, polarity-fused map, spatially smoothed representation, and temporally stabilized event frame at time \(t\), respectively. Event integration accumulates asynchronous events within a temporal window to construct a frame-like signal: \[I_t(x, y) = \sum_{e_i \in \mathcal{W}_t(x, y)} p_i \cdot \exp\left(-\frac{t - t_i}{\tau}\right),\] where \(\mathcal{W}_t(x, y)\) denotes the set of events at pixel \((x,y)\) around time \(t\), and \(\tau > 0\) is the temporal decay constant controlling the attenuation of past events. On this basis, polarity fusion aggregates positive and negative events to emphasize meaningful motion-induced changes while suppressing noise: \[P_t(x, y) = \begin{cases} 1, & \text{if } \sum_{e_i \in \mathcal{W}_t(x, y)} p_i > \theta_p, \\ -1, & \text{if } \sum_{e_i \in \mathcal{W}_t(x, y)} p_i < -\theta_p, \\ 0, & \text{otherwise}, \end{cases}\] where \(\theta_p\) is a predefined polarity threshold. Spatial coherence enhancement is then applied to improve local consistency while preserving structural boundaries: \[S_t(x, y) = (G_{\sigma} * I_t)(x, y),\] where \(G_{\sigma}\) denotes a Gaussian smoothing kernel with standard deviation \(\sigma > 0\). Finally, inter-frame continuity is maintained through temporal filtering to improve frame-to-frame stability and reduce flickering artifacts: \[F_t(x, y) = \alpha \cdot S_t(x, y) + (1 - \alpha) \cdot F_{t-1}(x, y),\] where \(\alpha \in [0,1]\) controls the blending ratio between the current and previous frames.
While the above formulations summarize the canonical event-to-frame conversion pipeline, they are still predominantly based on handcrafted aggregation rules. Such strategies are computationally convenient and relatively easy to integrate into existing CNN-based detectors, but they often rely on fixed temporal windows and simple accumulation mechanisms, which may suppress fine-grained temporal order, distort local event statistics, and entangle motion-related and appearance-related information. To address these limitations, recent research has increasingly shifted toward more statistically grounded and task-adaptive event representations [124]–[128].
In particular, TORE (Time-Ordered Recent Event) volumes explicitly preserve the recency ordering of recent events, instead of merely aggregating them within a fixed window, thereby retaining richer temporal structure for fast motion, motion boundaries, and transient object contours [124]. MAD (Motion and Appearance Decoupling Representation) further argues that conventional event tensors often entangle motion saliency and appearance context, and therefore proposes a decoupled representation that separates motion and appearance components before feature interaction and fusion [125]. In addition, data-driven representation learning has recently become an important direction. EvRepSL constructs event representations from spatio-temporal statistics and refines them via self-supervised learning, reducing dependence on manually designed aggregation heuristics [126]. Similarly, recent recurrent self-supervised representations such as SSER aim to avoid hard temporal discretization by learning asynchronous encoders directly from event timestamps and polarities, while also exhibiting strong potential for low-latency hardware deployment [127]. Furthermore, EVA extends this line of research toward event-by-event asynchronous representation learning by combining asynchronous encoding with linear-attention-based self-supervised learning, suggesting a possible transition from synchronous tensorization to highly expressive asynchronous representations [128]. Collectively, these studies indicate that event representation is evolving from fixed accumulation toward order-preserving, decoupled, and learnable statistical encoding, which is particularly relevant for pedestrian detection under high-speed motion, cluttered backgrounds, and adverse illumination.
In addition to the three conventional categories, Wan et al. [117] introduced an innovative event representation termed Neighborhood Suppression Time Surface (NSTS). This method is inspired by time-surface generation and HATS-based processing [129], [130]. Its key novelty lies in suppressing local intensity responses around each pixel independently, thereby reducing the dominance of dense-event regions and improving the relative significance of sparse but informative event regions.
Moreover, [131] employed the Sparse Surface Model (SSM) to convert event streams into structured event-frame tensors and validated its effectiveness on the Gen1 [105] and 1Mpx [111] datasets. SSM accumulates events over fixed temporal intervals to construct dense tensor maps \(H_k\), allowing convolutional layers to extract spatial features from structured event data. To preserve temporal continuity, Convolutional Long Short-Term Memory (ConvLSTM) layers are introduced, with the hidden state updated as: \[h_k = F(H_k, h_{k-1}),\] where \(F\) denotes the ConvLSTM operation.
For detection, multi-scale feature extraction layers are used to predict bounding boxes \(B_k\) and \(B'_{k+1}\), which are aligned with the corresponding ground-truth boxes \(B^*_k\) and \(B^*_{k+1}\). The overall training objective can be formulated as: \[\mathcal{L} = \mathcal{L}_{\mathrm{cls}} + \mathcal{L}_{\mathrm{reg}} + \mathcal{L}_{\mathrm{aux}},\] where \(\mathcal{L}\) denotes the total training loss, \(\mathcal{L}_{\mathrm{cls}}\) denotes the classification loss, \(\mathcal{L}_{\mathrm{reg}}\) denotes the bounding-box regression loss, and \(\mathcal{L}_{\mathrm{aux}}\) denotes the auxiliary supervision loss.
SSM leverages the high temporal resolution of event cameras to capture rapid dynamics while reducing redundancy in data representation. In parallel, Spiking Neural Networks (SNNs) remain an important family of event-native models, as they mimic the biological spiking mechanism and can process events at fine temporal granularity. The output spike train of a neuron can be expressed as: \[y(t) = \sum_{i=1}^{N} w_i \cdot s(t - t_i),\] where \(y(t)\) denotes the continuous output response, \(w_i\) are synaptic weights, \(t_i\) are spike times, and \(s(\cdot)\) is the spike response function, typically modeled as a causal decaying kernel.
Overall, event-to-frame conversion establishes an effective bridge between the high temporal resolution of event streams and the rich ecosystem of mature frame-based vision architectures [132]–[134]. Although some temporal precision is inevitably sacrificed during conversion, this paradigm remains highly practical for EB-PD owing to its strong compatibility with advanced detection backbones and efficient deployment frameworks [135].
Joint event-frame fusion seeks to combine the high temporal sensitivity of event data with the rich spatial context provided by conventional image frames [136]–[138]. This multimodal strategy is particularly attractive in pedestrian detection, as it enables complementary modeling of motion dynamics and appearance information.
Recent multimodal representation learning has moved beyond conventional synchronized two-branch fusion. In particular, ACGR (Asynchronous Collaborative Graph Representation) proposes a unified graph-based paradigm for jointly representing frames and events, rather than first converting event data into image-like tensors tied to the frame rate [139]. By constructing sparse unimodal graphs for both modalities and aligning them through asynchronous collaborative modules, ACGR preserves the spatio-temporal sparsity of event data while alleviating inter-modal misalignment. This insight is especially relevant for EB-PD, as it suggests that parallel event-frame processing need not be restricted to frame-rate-limited fusion, but can instead evolve toward unified asynchronous representations with both high efficiency and strong detection potential.
Algorithm 5 [93], [99], [137] presents a representative example of simultaneous frame-event input, which has become a cornerstone of advanced EB-PD systems. The key idea is to harmonize frame-based spatial cues with event-based temporal cues so as to produce a richer and more robust data representation for pedestrian detection.
The fusion process begins with the independent preprocessing of event and frame streams. For the event stream, preprocessing focuses on suppressing noise and removing irrelevant signals: \[E_{\mathrm{filtered}} = \{e \in E \mid \operatorname{NoiseFilter}(e) \land \operatorname{SignalThreshold}(e)\},\] where \(E\) denotes the raw event stream, \(E_{\mathrm{filtered}}\) denotes the filtered event set after noise and irrelevant-signal suppression, and \(\operatorname{NoiseFilter}(\cdot)\) and \(\operatorname{SignalThreshold}(\cdot)\) denote the event validity tests used in preprocessing. In parallel, frame data undergo spatial enhancement to improve texture clarity and structural detail.
The central challenge is temporal synchronization, since event streams are asynchronous while frame streams are sampled at fixed intervals. A temporal alignment operation may be expressed as: \[\{E_{\mathrm{sync}}, F_{\mathrm{sync}}\} = \operatorname{TemporalAlign}(E_{\mathrm{filtered}}, F, \tau),\] where \(F\) denotes the frame sequence, \(E_{\mathrm{sync}}\) and \(F_{\mathrm{sync}}\) denote the temporally synchronized event and frame streams, respectively, and \(\tau\) denotes the synchronization tolerance window.
Once aligned, feature extraction can exploit both the dynamic information captured by events and the contextual appearance information available in frames: \[\mathrm{Features} = \operatorname{ExtractFeatures}(E_{\mathrm{sync}}, F_{\mathrm{sync}}),\] thereby constructing a unified spatio-temporal feature representation.
This integration strategy highlights the methodological principle underlying parallel event-frame processing: events contribute fine temporal sensitivity, while frames provide dense spatial context. Their complementary combination enables a richer and more discriminative representation of pedestrian behavior in complex and dynamic environments.
The preprocessing stage in the EB-PD pipeline aims to improve data integrity and representation quality for each processing paradigm, thereby facilitating more effective feature extraction and downstream recognition. More generally, in event-based vision, preprocessing should be regarded not merely as a task-specific auxiliary step for pedestrian detection, but as a fundamental upstream component of event-stream perception. Its primary goals include suppressing spurious events and background activity noise, preserving informative spatio-temporal correlations, stabilizing event density across space and time, and improving robustness under heterogeneous sensing conditions.
From a broader methodological perspective, generic event-stream preprocessing and denoising methods can be organized into four representative categories: statistical/probabilistic denoising, noise-specific lightweight filtering, learning-based spatio-temporal denoising, and task- or scenario-aware extensions. As summarized in Table 2, these categories exhibit different trade-offs in interpretability, adaptability, computational cost, and deployment suitability. In parallel, recent resources such as LED and E-MLB have enabled more systematic training and evaluation of event denoising methods across diverse scenes and noise levels [140], [141].
| Category | Representative methods | Advantages | Limitations | ||||
|---|---|---|---|---|---|---|---|
| training-free; | |||||||
| preserves local consistency | |||||||
| weaker adaptability to | |||||||
| complex noise | |||||||
| low latency; | |||||||
| hardware-friendly | |||||||
| limited generality | |||||||
| EDformer [142]; | |||||||
| EDmamba [143]; | |||||||
| ASTEDNet [144]; | |||||||
| DBRGNN [145]; | |||||||
| Point-cloud noise modeling [146] | |||||||
| robust to varied noise; | |||||||
| captures complex dependencies | |||||||
| data-dependent; | |||||||
| some methods need | |||||||
| broader validation | |||||||
| AWTS [147]; | |||||||
| Controlled noise injection [148]; | |||||||
| SPIE denoising method [149] | |||||||
| useful in specialized settings | |||||||
| higher modeling complexity |
3pt
For direct event-stream processing, preprocessing is primarily centered on temporal fidelity. Adaptive noise filtering purifies the event stream [100], while event deduplication and event aggregation strategies expose meaningful temporal structures [150]. Event density normalization further improves the consistency of temporal feature extraction [151]. In this paradigm, preprocessing must operate directly on sparse asynchronous events, which makes the pipeline particularly sensitive to isolated spurious activations and background activity noise. Statistical and probabilistic methods are advantageous when interpretability and local event-structure preservation are emphasized, whereas lightweight background-activity suppression is preferable when low latency and deployment efficiency are the dominant constraints [152], [153]. More recent direct-stream denoisers, including asynchronous spatio-temporal neural denoising, residual graph neural networks, and unsupervised spatio-temporal point-cloud noise modeling, further show that denoising can be made more adaptive while remaining closer to the native asynchronous structure of event streams [144]–[146].
For event-to-frame conversion, preprocessing transforms asynchronous events into structured frame-like signals [154], thereby ensuring temporal continuity at the representation level. Through event accumulation and polarity encoding, motion saliency is enhanced. Spatial filtering [151], [155], [156] and related enhancement operations are subsequently applied to improve frame quality and highlight pedestrian-related structures. Regularization also plays an important role in ensuring compatibility with mature image-based processing pipelines [157], [158]. Here, however, the quality of the generated representation depends critically on whether noisy events are sufficiently suppressed before accumulation. In this sense, event-to-frame preprocessing is no longer limited to fixed-window accumulation and heuristic denoising, but is gradually evolving toward noise-adaptive and statistically informed representation refinement. Learning-based denoisers such as EDformer are particularly attractive in this setting because they improve robustness under heterogeneous noise conditions, while more efficient state-space formulations such as EDmamba are promising when computational efficiency must also be considered [142], [143]. Scenario-specific strategies, such as attribute-weighted time-surface denoising, further indicate that local event reliability estimation remains useful when refining event frames under challenging sensing conditions [147].
For joint event-frame fusion, preprocessing aims to achieve a synergistic integration of temporally dense event streams and spatially informative image frames [99]. This requires precise temporal synchronization to align the two modalities [159], [160]. Feature standardization is then employed to ensure a consistent cross-modal representation [93]. Such careful alignment is essential for the extraction of integrated spatio-temporal features [161], which are critical for capturing pedestrian motion under complex real-world conditions. Compared with single-branch pipelines, parallel event-frame systems can partially tolerate moderate event noise because the frame stream provides complementary appearance cues. Nevertheless, they are also vulnerable to misalignment errors introduced by unstable event preprocessing. Accordingly, in this setting preprocessing should emphasize not only synchronization, but also stable denoising and representation consistency prior to fusion. In addition, recent studies suggest that preprocessing may be further strengthened by coupling denoising with motion estimation or robustness-oriented noise modeling, rather than treating denoising as a fully isolated step [148], [162].
Feature engineering is a key stage in the EB-PD pipeline, in which preprocessed signals are transformed into discriminative and informative descriptors suitable for pedestrian detection [163]–[165].
For direct event-stream processing, feature extraction primarily leverages the temporal distribution of events to characterize dynamic scene variations. Localized spatio-temporal volumes capture micro-movements in confined regions [116], thereby revealing subtle pedestrian behaviors. Polarity statistics are useful for modeling edges and motion direction [92], [164], while spatial depth-related cues may further improve scene understanding in three-dimensional space.
For event-to-frame conversion, feature extraction benefits from the temporal continuity embedded in the generated frame sequence. Motion formation cues [97], [101], derived from polarity changes and accumulated event responses, reveal pedestrian trajectories and object motion patterns. The transformation process also recovers spatial textures and structural contours [166], enabling conventional detectors to model pedestrian appearance more effectively. Multi-scale analysis further improves robustness by capturing fine and coarse features simultaneously. More recent representation designs improve feature quality at the source: order-preserving encodings retain recency cues valuable for motion boundaries and rapid limb movement, decoupled representations separate motion saliency from appearance context, statistical/self-supervised encoders provide more stable descriptors under background clutter and illumination changes, and asynchronous representations reduce the loss of high-frequency temporal structure caused by forced discretization or frame-rate synchronization [124]–[128], [139].
For joint event-frame fusion, feature extraction and representation aim to integrate detailed temporal descriptors from event streams with dense spatial descriptors from image frames [99]. Spatio-temporal fusion techniques combine motion and appearance into a unified representation. By jointly modeling polarity, texture, and contextual cues from both modalities [93], such methods build more comprehensive feature descriptors for pedestrian perception. In addition, features for dynamic background adaptation help maintain model sensitivity under challenging environmental variations.
The final stage of the EB-PD pipeline concerns the design of detection models capable of exploiting the feature representations generated in the preceding stages. It should be emphasized that the effectiveness of these networks depends not only on the backbone architecture itself, but also on the quality of the underlying event representation, particularly whether temporal order, local statistics, and motion cues are adequately preserved during encoding [124]–[126].
CNNs have been widely adopted in EB-PD tasks, with representative architectures including YOLOv7 [91], YOLOv5 [96], [97], YOLOv3 [93], [100], [167], and YOLOv3-Tiny [98], [167]. These architectures are attractive because they provide a strong trade-off between detection accuracy, inference speed, and deployment efficiency. For example, YOLOv7 [168] and YOLOv5 [169] are known for high accuracy and real-time performance, making them suitable for latency-sensitive pedestrian detection scenarios. YOLOv3 [170] and YOLOv3-Tiny [171], while earlier architectures, still provide an effective balance between detection performance and computational cost, especially on resource-constrained platforms.
Algorithm 6 outlines the general detection logic of the YOLO family when adapted to EB-PD. In such models, the input image is divided into an \(S \times S\) grid, and each grid cell predicts \(B\) bounding boxes together with class probabilities. This grid-based formulation is well suited to event-driven detection systems that require efficient inference and rapid localization.
In addition to mainstream CNN detectors, alternative models have also been explored to address the unique characteristics of EB-PD. For instance, the Genetic Algorithm-Back Propagation (GA-BP) neural network [92] combines the global optimization ability of genetic algorithms with the learning efficiency of backpropagation, thereby improving robustness in complex environments. Likewise, the Spatial Attention Model (SAM) [172] improves detection performance by emphasizing informative spatial features through attention-based reweighting.
To facilitate a structured comparison, Table ¿tbl:method? summarizes representative EB-PD methods in terms of processing paradigm, model family, time window, dataset, and reported performance. However, these methods should not be interpreted merely as isolated implementations. Rather, they occupy different operating points on a shared trade-off surface spanning temporal fidelity, detection accuracy, computational efficiency, and deployment complexity. A critical review of EB-PD therefore requires not only method listing, but also cross-paradigm comparison.
To clarify the organization of this section, we first provide a cross-paradigm comparison of representative EB-PD methods from the perspectives of temporal fidelity, detection accuracy, computational efficiency, and deployment complexity. This comparison establishes the general trade-off principles among direct event-stream processing, event-to-frame conversion, and joint event-frame fusion. The subsequent subsections then apply these principles to two major application settings: dynamic traffic scenes and static surveillance scenes. In this way, the cross-paradigm analysis serves as the conceptual basis for the following scene-oriented discussions rather than as an isolated subsection.
A first major trade-off concerns the representation pathway itself. Direct event-stream methods preserve the native asynchronous nature of EVS and therefore best retain temporal fidelity, motion sharpness, and low-latency responsiveness. This makes them conceptually well aligned with the sensing principle of EVS, particularly in high-speed or rapidly changing scenes. However, raw event streams are sparse, irregular, and difficult to process with conventional vision backbones, which often leads to weaker architectural maturity and less stable detection performance in practice. By contrast, event-to-frame methods sacrifice part of the asynchronous structure in exchange for compatibility with mature frame-based detectors. Their primary advantage is practical accuracy and engineering accessibility: once events are converted into frame-like representations, a large ecosystem of CNN detectors can be reused effectively. The cost of this convenience is that temporal sparsity, microsecond timing, and part of the latency advantage of EVS are weakened during aggregation. Event-frame fusion methods occupy an intermediate position. They exploit the temporal sensitivity of events together with the dense spatial context of conventional frames, and thus often provide stronger robustness under nighttime, low-light, or cluttered conditions. However, this gain comes at the expense of higher system complexity, synchronization overhead, and potentially heavier deployment requirements.
A second and more fundamental trade-off arises between SNN-based and CNN-based processing. SNNs are, in principle, more faithful to event cameras because they operate naturally on sparse asynchronous spikes, avoid forced discretization, and are better aligned with neuromorphic hardware. In theory, this makes them appealing for low-power, event-native, and latency-sensitive perception. Yet the current EB-PD literature also suggests clear limitations: training remains more difficult, large-scale benchmark evidence is still limited, and practical accuracy often trails strong CNN-based event-frame pipelines. In contrast, CNNs currently benefit from substantially greater architectural maturity, stronger optimization stability, and a richer ecosystem of reusable detection backbones. This explains why many of the best-performing EB-PD methods still rely on event-to-frame conversion or event-frame fusion followed by CNN-style detectors. Therefore, the current state of the field does not support the simplistic conclusion that SNNs are universally superior. Rather, SNNs are more faithful to the sensing principle of EVS, whereas CNN-based pipelines currently benefit from stronger practical accuracy and engineering maturity.
A third critical trade-off, especially for event-to-frame approaches, concerns the temporal accumulation window \(\Delta T\). Increasing \(\Delta T\) generally improves event density, signal-to-noise ratio, and representation stability, which is beneficial for frame-based detectors. However, a larger window also increases effective delay, compresses temporal ordering, and may introduce motion smearing or temporal aliasing in fast scenes. Conversely, reducing \(\Delta T\) preserves temporal responsiveness and maintains more of the low-latency advantage of EVS, but often produces sparse and noisy frame representations that are harder to detect robustly. Therefore, event-to-frame EB-PD is fundamentally a trade-off between temporal fidelity and detector compatibility. A simplified interpretation is that the effective latency of an event-to-frame detector can be approximated as \[L_{\mathrm{eff}} \approx \frac{\Delta T}{2} + L_{\mathrm{inf}},\] where \(L_{\mathrm{eff}}\) denotes the approximate effective latency of the event-to-frame detector, \(\Delta T\) denotes the temporal accumulation window, and \(L_{\mathrm{inf}}\) denotes the network inference time. Although this is not a strict universal formula, it provides an intuitive explanation of why temporal aggregation may erode the intrinsic low-latency benefit of EVS if the accumulation window is chosen too conservatively.
Taken together, existing EB-PD methods should not be viewed as merely different implementations, but rather as different operating points on a shared design surface defined by temporal fidelity, detection accuracy, computational efficiency, and deployment complexity. From this perspective, SNN-based direct-event methods are attractive when event nativeness, energy efficiency, and latency are prioritized; CNN-based event-frame methods remain appealing when accuracy, stability, and reuse of mature detectors are the main objectives; and fusion-based methods are especially valuable when robustness across difficult illumination and dynamic backgrounds is more important than architectural simplicity. These general trade-off principles also provide the basis for the following scene-oriented analysis. In dynamic traffic scenes, latency, robustness to motion, and illumination adaptability are usually dominant concerns, whereas in static surveillance scenes, stable long-duration observation, compact feature extraction, and reliable scene understanding become more important. Therefore, the following two subsections discuss how the above cross-paradigm trade-offs are reflected in dynamic and static EB-PD applications, respectively.
Dynamic-scene EB-PD, especially in autonomous driving, is fundamentally governed by a latency–robustness–accuracy trade-off. In such scenarios, the detector must respond quickly to abrupt scene changes, while remaining reliable under motion blur, severe illumination transitions, and environmental noise. As a result, existing methods can be interpreted less as isolated algorithmic proposals than as different strategies for balancing temporal responsiveness against representational stability.
A dominant strategy in this setting is to convert event streams into frame-like representations and then exploit mature CNN detectors. This design is attractive because it allows EB-PD to inherit the optimization maturity, architectural stability, and practical accuracy of frame-based detection pipelines. For example, [117] introduces an asynchronous feature extraction framework on top of event-frame conversion, achieving approximately 26 FPS and 87.43% AP on a real-world dataset. Likewise, [100] argues that direct application of raw event streams to conventional object detectors is often impractical, and therefore proposes an improved event-to-frame conversion strategy together with intermediate-feature reuse, again reaching 26 FPS and 87.43% accuracy. These methods illustrate a recurring pattern in dynamic-scene EB-PD: sacrificing part of the native asynchrony of EVS in order to obtain stronger detector maturity and more stable engineering performance. Their practical value is evident, but so is their limitation—as the accumulation process becomes heavier, the low-latency sensing advantage of EVS is progressively weakened.
A second line of work emphasizes robustness under real-world dynamics through motion modeling or multimodal fusion. In [101], event-based noise filtering is combined with CTRV motion modeling, three-dimensional enhanced K-means clustering, and SCDEKF-based motion estimation, with validation under actual road-testing conditions. The contribution of this work lies less in detector novelty and more in showing that real-road EB-PD benefits from explicitly modeling motion uncertainty and target association. Similarly, [99] integrates RGB and EVS data for semantic segmentation and depth estimation across multiple viewpoints, improving adaptability under nighttime and low-light conditions. Such methods indicate that, in dynamic scenes, robustness is often achieved not merely by strengthening the backbone, but by incorporating motion priors, multimodal complementarity, or scene-level constraints. The advantage of this paradigm is improved resilience under adverse environments; its drawback is greater system complexity, synchronization cost, and dependence on multiple sensing channels.
A third line of work seeks to preserve the event-native advantage of EVS in extreme dynamic conditions. [113] demonstrates that purely event-based inputs can already support semantic segmentation, especially under extreme lighting transitions and for highly dynamic targets, while [173] combines EVS data with human kinematic analysis to improve pedestrian detection and tracking. These studies are important because they highlight that EVS is not merely a drop-in replacement for RGB sensing, but a distinct modality whose motion selectivity can itself be algorithmically valuable. However, they also reveal a present limitation of the field: event-native methods are often more specialized, less standardized, and less supported by mature detection ecosystems than event-frame CNN pipelines.
Overall, dynamic-scene EB-PD currently favors methods that give up part of native event fidelity in exchange for stronger robustness and deployability. Event-to-frame CNN pipelines remain dominant because they provide the most mature balance of accuracy and engineering convenience, whereas event-native and fusion-based methods are most advantageous when rapid motion, illumination extremes, or adverse weather make the intrinsic sensing properties of EVS indispensable.
In static-scene EB-PD, the dominant design priorities differ from those in autonomous driving. While low latency remains beneficial, the central requirement is often stable scene understanding for surveillance, target monitoring, posture-related analysis, or behavior recognition. Consequently, methods in static environments more frequently prioritize compact feature extraction, signal interpretability, and robustness to long-duration observation, rather than preserving every aspect of the event camera’s microsecond temporal advantage.
One important category consists of direct event-stream analysis and lightweight decision models. For example, [150] integrates an event filtering module with a binary-neural-network-based detection module, showing that event-native preprocessing can reduce noise and simplify downstream classification. Similarly, [92] combines genetic optimization with backpropagation for road-target classification, while [112] performs hierarchical clustering directly on raw event streams for object tracking. These methods reflect an efficiency-oriented design philosophy: instead of relying on deep and heavy backbones, they attempt to exploit the sparse structure of event data through compact models or direct signal-level operations. Their main strength is low computational overhead and relatively clear interpretability; their main weakness is that performance often depends strongly on handcrafted assumptions, scenario tuning, or carefully controlled acquisition conditions.
A second category relies on event-derived descriptors or shallow learned representations. For instance, [174] extracts five target characteristics from 2D event point clouds and uses an SVM classifier for road-target recognition, reporting very high accuracy in a specific acquisition setup. Such methods demonstrate that event streams contain sufficiently rich geometric and motion cues to support discrimination even without large deep backbones. However, they also expose a recurring limitation of static-scene EB-PD: excellent results on self-collected or constrained scenarios do not necessarily imply strong generalization across broader benchmarks. In other words, descriptor-based success often reflects close alignment between feature design and acquisition setting, rather than universally transferable robustness.
A third category introduces fusion-based compensation for the incompleteness of event-only sensing. [167] observes that EVS data offer high dynamic range, low latency, and sparse motion sensitivity, but lack absolute brightness information; by combining APS frames with EVS signals through confidence-map fusion, the method improves consistency and accuracy over standard frame-based solutions. Likewise, [98] employs multiple event encoding schemes together with both channel-level and decision-level fusion to enrich representation quality. The key insight of these methods is that, in static or surveillance-like settings, fusion is used less to recover raw temporal responsiveness and more to compensate for the semantic incompleteness of event-only data. The resulting gain is improved recognition robustness; the corresponding cost is increased model complexity, synchronization overhead, and dependence on carefully balanced modality interaction.
Finally, works such as [116] and [175] show that static-scene EB-PD can benefit from event-specific sensing properties beyond conventional detection. [116] uses active event cuboids and EMST descriptors for efficient abnormal-event analysis, whereas [175] introduces a multimode neuromorphic sensor with illumination measurement capability and reports data transmission rates far exceeding those of frame-based cameras. These studies reinforce an important point: event cameras are not only alternative detector inputs, but also sensing devices that may enable different formulations of scene understanding, especially when sparse motion saliency is more informative than dense appearance.
Overall, static-scene EB-PD tends to favor compact feature extraction, direct signal exploitation, and multimodal compensation rather than purely latency-driven design. Compared with dynamic-scene methods, the key trade-off here is less about preserving every microsecond of temporal precision and more about whether event-native sparsity can be converted into reliable scene understanding without over-relying on handcrafted assumptions or narrowly tailored acquisition setups.
While datasets specifically designed for event-based pedestrian detection remain limited, this survey includes both dedicated pedestrian-detection datasets and broader event-based benchmarks that are relevant to EB-PD. Since acquisition conditions alone are insufficient for assessing the practical usefulness of a dataset in this field, we review both the sensing settings and the annotation resources available for downstream perception tasks. We then summarize the evaluation metrics most commonly used in EB-PD.
Table ¿tbl:tab:dataset-comparison? summarizes the acquisition settings and scenario characteristics of the datasets reviewed in this survey, including sensing modality, temporal resolution, environmental conditions, and image resolution. Table ¿tbl:tab:dataset-scale? further reports annotation-related statistics, including annotation targets, category labels, number of classes, annotation quantity, and overall dataset scale. It should be noted that not all surveyed datasets are standard 2D pedestrian-detection benchmarks; several remain highly relevant to EB-PD through tasks such as pose estimation, geometry-aware perception, activity recognition, and multimodal event-frame learning. Accordingly, the field “#BBoxes / annotations” is used in a broad sense to cover bounding boxes, pose labels, activity labels, semantic annotations, or aligned frames when appropriate.
max width=
max width=
PEDRo2 [104] is a comprehensive dataset targeting human subjects in various environments and lighting conditions. The dataset contains manually annotated event data and grayscale imagery depicting varied actions by different individuals. A handheld camera was utilized during data collection, which, despite the integration of a buffering mechanism, introduced additional noise into the recordings. A distinctive characteristic of PEDRo is therefore its realistic viewpoint variation and motion-induced perturbation, making it more challenging than purely static-camera benchmarks. In addition, the dataset includes individuals aged between 20 and 70 and covers a broad spectrum of non-extreme weather conditions and diurnal variations, which makes it particularly valuable for studying person detection under diverse real-world sensing conditions.
GEN13 [105], the preeminent event-based dataset available for pedestrian detection in vehicular environments, employs GEN1 cameras mounted behind the windscreen of the car, alongside conventional grayscale cameras. Data collection spanned various French locales—from bustling urban centers and quiet towns to highways, rural stretches, and suburbs—across different diurnal and seasonal times and under varying meteorological conditions. More importantly, GEN1 has become one of the most influential automotive event-based benchmarks for object detection, since its acquisition setting closely matches realistic autonomous-driving perception.
Henri’s4 [176] dataset depicts a busy cityscape as seen from the front windscreen of a car in Zurich. The GEN3 camera and the rigid settings of a HUAWEI P20 smartphone were used to capture the footage, with both the event camera and frame camera recording at a resolution of 640\(\times\)480. While the data are not organized as a standard pedestrian-detection benchmark, the synchronized event and frame recordings document urban driving scenes under varying weather conditions and times of day, with a frequent focus on pedestrians and road users. Additionally, the dataset includes local weather data from Zurich, road condition information, and close-up lighting transitions, such as entering and exiting tunnels, thereby enriching the analysis of driving-related visual perception.
PAFBenchmark5 [106] covers three scenarios: pedestrian detection, motion detection, and fall detection, with the latter two documented within an open office setting. The pedestrian-detection component of the dataset encapsulates diverse scenarios including pedestrian overlap, occlusion, collision, and other common situations in traffic monitoring tasks. The action detection subset records 15 subjects performing 10 different actions, while the fall-detection subset similarly encompasses recordings of 15 subjects executing predefined motions such as falling, bending, tripping, and tying shoes. The dataset was captured using a fixed camera tripod and a laptop computer, and is therefore useful not only for pedestrian detection but also for broader human-centered event-based understanding.
FJUPD6 [108] is an extension of the PAFBenchmark [106] dataset for more detailed pedestrian-detection experiments, with a particular focus on outdoor scenes and low-light conditions. The data are divided into “simple” and “complex” categories based on luminosity and shadow dynamics, where complex backgrounds are characterized by pronounced light fluctuations and moving shadows. Compared with the original benchmark, its main value lies in providing more challenging lighting and background conditions, making it particularly useful for evaluating the robustness of event-based pedestrian detection under difficult illumination and shadow interference.
DVS-OUTLAB7 [107] uses a static acquisition method to capture activity within a fixed square area across a 2,800 m2 amusement park monitored by three fixed sensors. Positioned approximately 6 meters high with a 25-degree tilt toward the ground, the entire system operates on a self-sufficient solar energy storage system. Beyond long-duration outdoor event capture, the dataset also provides labeled regions of interest, including both object-related labels and environmental interferences such as rain and shadows. This makes DVS-OUTLAB particularly valuable for studying long-term outdoor event sensing under realistic environmental disturbances.
DHP198 [109] is the first dataset to use EVS for 3D human pose estimation, capturing the 3D positions of human joints through streams of events from multiple synchronized EVS cameras. This recording utilized four DAVIS cameras and the Vicon motion capture system, which consists of ten infrared cameras surrounding a motorized treadmill, thus facilitating varied movements by subjects. Although it is not a standard pedestrian-detection benchmark, DHP19 is highly relevant to EB-PD because it reveals how event streams encode fine-grained human motion structure, which is valuable for posture-sensitive pedestrian perception and behavior analysis.
NU-AIR9 [110] consists of 70.75 minutes of event-camera footage captured by an EVS-equipped quadcopter, surveying diverse urban settings including crowds, various vehicles, and busy streetscapes. Annotations are provided at 30 Hz for pedestrians and vehicles, making it a valuable resource for aerial surveillance and event-based visual studies. Its main significance lies in extending EB-PD beyond ground-view driving scenarios to airborne perception settings.
1Mpx10 [111] includes a diverse set of driving environments such as city streets, highways, and rural areas. A key feature of this dataset is the inclusion of very large-scale high-frequency annotations covering cars, pedestrians, and two-wheelers. The dataset was created using a novel automated labeling protocol that combines data from an event camera and a standard RGB camera. This extensive labeling and high resolution make it particularly valuable for developing and testing event-based object-detection systems under realistic driving conditions.
MVSEC11 [177] is a large-scale multimodal dataset designed primarily for stereo depth estimation, visual odometry, and event-based 3D perception. It includes synchronized event streams, grayscale images, IMU measurements, LiDAR scans, and pose information collected from multiple platforms such as cars, motorcycles, handheld devices, and hexacopters. Rather than focusing on 2D pedestrian bounding boxes, MVSEC is especially valuable for geometry-aware event-based perception tasks. For EB-PD, its relevance lies in supporting pedestrian perception when integrated with depth, motion, and scene-structure understanding.
eTraM12 [178] offers fully event-based traffic footage recorded from a static overhead perspective using a Prophesee EVK4 HD event camera. Captured across diverse urban environments—including intersections, roadways, and local streets—the dataset features varying lighting and weather conditions and covers eight traffic participant classes such as pedestrians, vehicles, and micro-mobility. Its high resolution and overhead viewpoint make it a particularly valuable benchmark for event-based traffic monitoring and low-light object detection.
SEVD13 [179] presents a large-scale fully synthetic multimodal traffic dataset including event streams, RGB, depth, optical flow, semantic masks, and instance segmentation. Captured in both ego-vehicle and static-camera configurations across diverse road types, SEVD supports multiple traffic-perception tasks and provides large-scale 2D/3D annotations across major traffic-participant categories. Its key value lies in its synthetic yet richly annotated multimodal nature, which makes it particularly useful for benchmarking multimodal event-based perception, large-scale supervised learning, and synthetic-to-real transfer.
TUMTrafEvent14 [180] provides synchronized event-based and RGB frames captured from a fixed roadside gantry overlooking a busy urban intersection. The dataset combines high-resolution RGB imagery with dense asynchronous event streams and manually verified 2D bounding boxes across multiple road-user categories. With its precise alignment and realistic roadside viewpoint, it serves as an important benchmark for event-RGB fusion in intelligent transportation systems and offers a useful complement to ego-vehicle datasets.
HARDVS 2.015 [181] consists of paired RGB and event sequences recorded using a DAVIS346 sensor across indoor and outdoor scenes with diverse lighting and motion conditions. Captured at 30 FPS for RGB and microsecond resolution for events, the dataset includes 300 daily human activity categories under both static and dynamic camera setups. Although it is not a conventional pedestrian-detection dataset, HARDVS 2.0 is important because it expands event-based perception from object localization toward higher-level human-centered understanding, making it a comprehensive benchmark for multimodal human activity recognition in event-based vision.
DSEC16 [182] offers stereo event-based and frame-based driving data recorded in urban, suburban, and rural areas of Switzerland. It features two synchronized Prophesee Gen3.1 event cameras, stereo RGB cameras, LiDAR point clouds, and high-precision GPS. Unlike standard 2D detection datasets, DSEC primarily supports stereo depth estimation, optical flow, disparity, and related geometry-aware perception tasks, with dense pixel-level annotations and multiple driving sequences. Its relevance to EB-PD lies in the fact that pedestrian detection in autonomous driving is often only one component of a broader perception stack, and DSEC provides an important benchmark for understanding how event cameras contribute to geometry, motion, and scene-structure estimation.
The performance of EB-PD algorithms is commonly evaluated using three fundamental metrics: the log-average miss rate (MR), average precision (AP), and Intersection over Union (IoU) [183]. The MR quantifies detection failures by reflecting the proportion of ground-truth objects that are missed, whereas AP summarizes detection precision across different recall levels. IoU, in turn, measures the spatial overlap between predicted and ground-truth bounding boxes and serves as the basis for classifying predictions as true positives or false positives.
To determine true positives (TPs), false positives (FPs), and false negatives (FNs), a greedy matching strategy is typically adopted, as shown in Algorithm 7. Detections are first sorted according to confidence score, and each detection is then matched to the ground-truth instance with the highest IoU, provided that the overlap exceeds a predefined threshold. This matching procedure is central to the computation of both MR and AP.
Collectively, these metrics characterize different aspects of EB-PD performance. MR reflects detection sensitivity, AP captures the balance between precision and recall, and IoU measures spatial localization quality. In crowded scenes, where pedestrian overlap is frequent, IoU becomes particularly important because small localization errors can strongly affect whether a detection is counted as correct.
Intersection over Union (IoU) is analogous to the Jaccard index and measures the accuracy of spatial overlap between the predicted bounding box \(B_d\) and the ground-truth bounding box \(B_g\). It is defined as \[\mathrm{IoU}(B_d, B_g) = \frac{\mathrm{Area}(B_d \cap B_g)}{\mathrm{Area}(B_d \cup B_g)},\] where a higher IoU indicates more accurate localization. In most detection settings, a prediction is counted as a true positive only when the IoU exceeds a predefined threshold, commonly 0.5. The intersection area \(|B_d \cap B_g|\) can be written as \[\begin{align} & \max \bigl(0, \min(x_d^{\max}, x_g^{\max}) - \max(x_d^{\min}, x_g^{\min}) \bigr) \notag \\ & \times \max \bigl(0, \min(y_d^{\max}, y_g^{\max}) - \max(y_d^{\min}, y_g^{\min}) \bigr), \end{align}\] and the union area \(|B_d \cup B_g|\) is given by \[\begin{align} A_{B_d} + A_{B_g} - |B_d \cap B_g|, \end{align}\] where \(A_{B_d}\) and \(A_{B_g}\) denote the areas of \(B_d\) and \(B_g\), respectively.
Average Precision (AP) is a standard metric derived from the precision–recall curve. It reflects detection quality across different recall levels and is commonly computed as \[AP = \sum_{k=1}^{n} \bigl(R(k) - R(k-1)\bigr) P(k),\] where \(P(k)\) and \(R(k)\) denote the precision and recall at the \(k\)-th threshold, respectively. Precision is the ratio of true positive detections to all detections, whereas recall is the ratio of true positive detections to all actual positives.
Miss Rate (MR) is particularly useful for imbalanced detection tasks, where background samples substantially outnumber object instances. It is defined as \[MR = 1 - TPR = 1 - \frac{TP}{TP + FN},\] where \(TPR\) is the true positive rate, \(TP\) denotes the number of true positives, and \(FN\) denotes the number of false negatives. In pedestrian detection, the log-average miss rate is often computed over a predefined range of false positives per image on a logarithmic scale, typically from \(10^{-2}\) to \(10^{0}\).
Although EB-PD has advanced substantially in recent years, its large-scale real-world deployment remains constrained by several intertwined challenges. These challenges are not limited to algorithmic performance, but also involve hardware cost, dataset standardization, evaluation methodology, and the lack of sufficiently mature event-native perception models. Consequently, future progress in EB-PD depends not only on improving detection accuracy, but also on strengthening the practical foundations required for robust and scalable deployment.
A primary obstacle to the widespread adoption of event-based cameras is the current difficulty of industrial-scale deployment. Many event-based systems still rely heavily on Field-Programmable Gate Arrays (FPGAs) as their main control units [68], [69]. While such hardware offers flexibility for prototyping and sensor interfacing, it also significantly increases system cost, complexity, and integration difficulty. These factors hinder large-scale commercialization, especially in cost-sensitive sectors such as mainstream automotive systems, intelligent roadside infrastructure, and consumer-grade edge devices.
This hardware issue is closely related to the broader problem of deployment readiness. Despite the attractive theoretical advantages of EVS, including low latency, high dynamic range, and sparse data output, large-scale real-world studies remain limited. Current evidence remains insufficient to determine when and under what conditions event-based pedestrian detection can consistently outperform or reliably complement mature frame-based systems across diverse autonomous-driving environments. Therefore, broader real-world validation, long-term field trials, and hardware-software co-design remain essential before EB-PD can become a dependable industrial solution.
Another major challenge lies in the lack of standardized large-scale datasets and unified evaluation protocols specifically tailored to event-based pedestrian detection. Existing datasets vary substantially in sensor type, annotation quality, environmental diversity, class definitions, and data format, which complicates fair comparison across methods. In many cases, models are still validated on self-collected datasets or on event datasets originally developed for tasks beyond pedestrian detection. Although such datasets are valuable, they often limit reproducibility and weaken the comparability of published results.
In parallel, the evaluation of EB-PD still depends largely on metrics inherited from frame-based object detection. While standard metrics such as AP, MR, and IoU remain useful, they do not fully characterize key properties of event-based sensing, such as temporal fidelity, asynchronous responsiveness, and robustness under extreme dynamic illumination. As a result, there is still no complete evaluation framework specifically designed for EB-PD. Future work should therefore focus on both large-scale standardized benchmarks and more task-relevant evaluation criteria that better reflect the unique sensing characteristics of event cameras.
A further limitation of current EB-PD research is the insufficient development of models specifically designed for event data. Many existing approaches still process event streams through frame-like conversion or rely on conventional backbones originally developed for RGB imagery. While these strategies are practical and often effective, they do not fully exploit the fine-grained temporal structure and sparse asynchronous nature of event streams. This is particularly limiting from a three-dimensional and motion-aware perspective, where event data potentially contain richer cues than those captured by frame-based representations.
To address this gap, future research should move beyond adapting existing vision models and instead develop architectures that are intrinsically compatible with event-based sensing. This includes models that can more effectively mine spatio-temporal structure, exploit sparse asynchronous patterns, and jointly optimize sensing, preprocessing, and inference. In addition, progress will likely require stronger collaboration between academia and industry, not only for model design, but also for sensor engineering, standardization, and large-scale validation. In this sense, the future of EB-PD depends on a transition from proof-of-concept systems to integrated, benchmarked, and deployment-oriented event-native perception pipelines.
Looking forward, the future of pedestrian detection with event-based sensors is likely to be shaped by a convergence of advances in hardware, adaptive perception, multimodal learning, and responsible deployment. These directions are closely connected: improvements in sensing hardware affect algorithm design, while progress in representation learning and system integration will determine whether event cameras can transition from niche experimental devices to widely usable perception components.
One of the most promising research trends is the integration of event cameras with edge-intelligent and sense-compute convergence technologies. Because event-based cameras naturally produce sparse outputs and operate with low latency, they are well suited to scenarios requiring efficient on-device perception. In principle, tighter integration between sensing and computation could enable real-time processing with lower energy consumption and reduced communication overhead, which is particularly attractive for autonomous vehicles, wearable devices, drones, and IoT systems.
At the same time, this trend is strongly related to the need for cost reduction and industrial scalability. Future progress may depend partly on reducing reliance on expensive and relatively inflexible FPGA-dominated solutions and developing more integrated, application-specific hardware. Such developments could make event-based sensing more accessible and commercially viable, while also encouraging the design of algorithms that are more tightly matched to deployment constraints.
Another important direction is the development of event-camera systems with stronger environmental adaptability. In practical applications, the usefulness of event sensors depends not only on their temporal precision, but also on their ability to remain stable across changing illumination, cluttered backgrounds, and complex motion patterns. Continued progress in denoising, adaptive thresholding, bias tuning, sensor calibration, and self-adaptive representation learning may substantially improve the reliability of event-based perception under real-world variability.
This direction also naturally connects with multimodal learning. Event cameras alone provide motion-sensitive and temporally precise information, but often lack the dense appearance cues available in frame-based sensing. As a result, future systems are likely to rely increasingly on multimodal fusion, combining events with RGB images, depth, LiDAR, or inertial signals. Such integration may enable richer perception pipelines in which event streams contribute temporal sensitivity while complementary modalities provide spatial and semantic completeness.
Although autonomous driving remains one of the most important application domains for EB-PD, the scope of future deployment is much broader. Event-based pedestrian perception may also play an important role in smart-city infrastructure, roadside traffic monitoring, public safety, human behavior analysis, wearable assistance, and robotic perception in visually challenging environments. As event-based technology matures, its applications are likely to expand from specialized experimental settings toward more diverse human-centered perception tasks.
At the same time, broader deployment will inevitably raise regulatory and ethical questions. Since many applications involve surveillance, monitoring, or large-scale urban sensing, future research must address privacy, responsible data governance, and compliance with relevant legal frameworks. Therefore, the long-term development of EB-PD should not be viewed solely as a technical trajectory. Its success will also depend on whether advances in sensing and algorithms can be aligned with social acceptance, regulatory requirements, and practical deployment needs.
In summary, the future of EB-PD will be shaped not only by improvements in detection algorithms, but also by progress in sensor design, system integration, benchmark standardization, adaptive perception, and ethical deployment. Continued innovation across these dimensions will determine whether event-based pedestrian detection can evolve from a promising research topic into a robust and widely adopted real-world technology.
This review has systematically examined the development of event-based vision sensing for pedestrian detection, with particular emphasis on its relevance to autonomous driving, intelligent transportation, and intelligent surveillance. By reviewing the literature from sensing principles to downstream detection pipelines, we have outlined the technological evolution of EVS and highlighted its key advantages over conventional frame-based imaging, including low latency, high temporal resolution, sparse data output, and strong robustness under challenging illumination.
We further reviewed the major methodological components of EB-PD, including sensing and processing paradigms, preprocessing strategies, feature representations, detection models, datasets, and evaluation metrics. In particular, we emphasized that existing EB-PD methods should not be understood merely as separate implementations, but rather as different design choices along shared trade-offs involving temporal fidelity, detection accuracy, computational efficiency, and deployment complexity. From this perspective, direct event-stream methods, event-to-frame methods, and event-frame fusion approaches each occupy different but complementary positions within the current research landscape.
Despite recent progress, several important challenges remain. These include the scarcity of standardized large-scale benchmarks tailored to EB-PD, the continued reliance on frame-based evaluation conventions, the limited maturity of event-native detection models, and the difficulty of translating promising laboratory results into robust real-world systems. Future progress will likely depend on advances in event-specific representation learning, hardware-software co-design, multimodal fusion, adaptive perception, and deployment-oriented evaluation.
Overall, event-based vision offers a compelling direction for next-generation pedestrian perception. Although the field is still evolving, continued progress in sensing technology, model design, and benchmark development is likely to make EB-PD an increasingly important component of future intelligent perception systems.
This work was supported by the National Key R&D Program of China (Grant No. 2023YFB4503000), the Hefei Key Technology R&D “List and Command” Project (Grant No. 2023SGJ035), the HFIPS Director’s Fund (Grant No. YZJJKX202401), and the Hefei Municipal Natural Science Foundation (Grant No. HZR2414).
The compiled documentation supporting our research is publicly available at https://github.com/TristanWH/DVS4PD.↩︎
https://github.com/SSIGPRO/PEDRo-Event-Based-Dataset↩︎
https://www.prophesee.ai/2020/01/24/prophesee-gen1-automotive-detection-dataset/↩︎
http://rpg.ifi.uzh.ch/e2vid↩︎
https://github.com/CrystalMiaoshu/PAFBenchmark↩︎
https://github.com/fjcu-ee-islab/Spiking_Converted_YOLOv4↩︎
http://dnt.kr.hsnr.de/DVS-OUTLAB/↩︎
https://github.com/uzh-rpg/event-based_vision_resources/↩︎
https://bit.ly/nuair-data↩︎
https://www.prophesee.ai/category/dataset/↩︎
https://daniilidis-group.github.io/mvsec/↩︎
https://eventbasedvision.github.io/eTraM↩︎
https://eventbasedvision.github.io/SEVD/↩︎
https://innovation-mobility.com/en/project-providentia/a9-dataset/↩︎
https://github.com/Event-AHU/HARDVS/tree/HARDVSv2↩︎
http://rpg.ifi.uzh.ch/dsec.html↩︎