REAL-OW: REhearsAL-free Open World Object Detection with Low-Rank
Adaptation and Dual-Stage Objectness Modeling
Supplementary Material
July 03, 2026
Open-World Object Detection (OWOD) requires detectors to identify previously unseen objects as unknown and incrementally incorporate them into the set of known categories, while preserving previously acquired knowledge. Existing frameworks rely heavily on exemplar replay to mitigate catastrophic forgetting, but in some real applications, storing raw data conflicts with data access restrictions and leads to data exposure risks, while incurring significant memory overhead. In this paper, we propose REAL-OW, a novel rehearsal-free framework that decouples incremental knowledge through a collaborative adapter architecture based on Low-Rank Adaptation (LoRA). Specifically, we deploy General Adapters (GAs) in the backbone to enable the significance-aware refinement of cross-task universal representations, while Specific Adapters (SAs) in the decoder provide orthogonal storage for task-specific expertise. To resolve representation drift in objectness modeling under rehearsal-free constraints, we introduce Dual-Stage Objectness Modeling (DSOM), which alternates between feature aggregation and boundary consolidation to stabilize objectness distributions while maintaining the separation between known and unknown categories. Furthermore, DSOM is supported by a Calibrated Gaussian Negative Log-Likelihood (CG-NLL) distance tailored for the dispersed feature distributions inherent in rehearsal-free settings. Extensive evaluations demonstrate that REAL-OW achieves state-of-the-art performance, surpassing existing exemplar replay methods in both detection precision and unknown discovery. Our approach establishes a new baseline for rehearsal-free OWOD.
Traditional object detection assumes a closed-set environment[1], [2], and it can only recognize categories that were defined during the initial training phase. However, real-world settings are dynamic and unpredictable[3]. To bridge this gap, Open-World Object Detection (OWOD) [4] was introduced and it allows systems to operate in unconstrained settings[5]. An OWOD detector must detect unknown objects and take them as novel classes, incrementally learning these new categories without forgetting previous ones [6]–[9]. This ability is vital for computer vision scenarios like autonomous driving and open-ended robotics [10]–[12].
Despite its importance, current OWOD frameworks rely heavily on exemplar replay [4], [13], [14], which mitigates catastrophic forgetting by maintaining a buffer of representative samples from previously learned classes and replaying them during new task training to preserve old knowledge[15]–[18].
However, exemplar replay incurs several practical challenges in real-world applications. First, privacy-sensitive fields like medical imaging prevent models from retaining raw data after a task ends to avoid data exposure. This makes it prohibitive to store samples for exemplar replay[19]. Second, as the task sequence expands, the buffer must accumulate more images to cover every previous class, thus incurring growing storage costs [20]. Furthermore, from the perspective of the OWOD objective, which is the continuous exploration and discovery of unknown categories, the learning process should prioritize modeling novel classes rather than repeatedly revisiting previously learned ones[4], [10].
These concerns and limitations motivate us to explore a rehearsal-free approach that avoids storing raw exemplar buffers. Low-Rank Adaptation (LoRA) [21] emerges as a compelling solution for its underlying parameter-efficient fine-tuning mechanism [22]–[26]. By constraining each added task updates into low-rank subspaces, LoRA enables the model to acquire new category knowledge while minimizing interference with previously learned representations [26]–[30], and most importantly, it achieves this without relying on data replay.
In rehearsal-free OWOD, learning and preserving a continuous stream of knowledge without accessing old data is a fundamental challenge. A single shared representation cannot distinguish task-specific features, while completely independent modules hinder knowledge accumulation and transfer, leading to catastrophic forgetting[22], [24]. To address this, we decouple general and specific knowledge by adopting a collaborative adapter framework, which separates general representations from task-specific expertise. As illustrated in 1, our training process does not include any data replay. This avoids repeated data access, significantly enhancing data privacy protection and reducing storage requirements. To balance general knowledge and task-specific expertise, our architecture employs two collaborative adapters within the frozen pre-trained model. Specifically, General Adapters (GAs) are placed in the backbone to maintain cross-task universal representations, while Specific Adapters (SAs) are added to the transformer decoder incrementally for each new task to capture knowledge for newly introduced categories.
Beyond incremental learning, OWOD further requires the discovery of unknown objects, which are not annotated but need to be discovered during training[3], [4], [10]. Recent methods employ probabilistic objectness[13], [31]–[33] and energy-based modeling[34], [35] to discover unknown objects in unmatched queries and separate them from known classes and background, respectively. However, these strategies depend on exemplar replay to stay distribution stable[10]. In rehearsal-free settings, the model continuously updates its parameters with emerging new data without access to historical samples, leading to potential feature drift, which causes two major problems: (1) unknown decision bias, where emerging unknown class features are pulled toward current known class regions, causing them to inherit known class similarity bias, while unknown objects with significantly different textures from current known categories are often overly suppressed and misclassified as background; and (2) catastrophic forgetting.
To solve these issues, we develop a Dual-Stage Objectness Modeling (DSOM) strategy. As shown in 1, this approach uses learning phases to aggregate the known feature embeddings and consolidation phases to sharpen the boundary between the knowns and unknowns. Furthermore, we observe that under this rehearsal-free condition, the dual-stage update disperses the known object distribution, rendering the standard Mahalanobis distance[13], [35], [36] ineffective for uncertainty estimation due to its sensitivity to distribution compactness. To address this, we introduce a Calibrated Gaussian Negative Log-Likelihood (CG-NLL) distance, which accounts for the increased variance, better optimizes the expanded objectness density and stabilizes the energy landscape.
Experimental results demonstrate that our framework outperforms all compared exemplar replay methods in both precision and unknown discovery. Remarkably, our model achieves these results without replaying a single image. To the best of our knowledge, we are the first to propose a strictly rehearsal-free methodology for OWOD, thereby establishing a new baseline for the field. Our main contributions can be summarized as follows:
We design General and Specific Adapters to continuously refine shared representations while isolating task-specific knowledge, enabling rehearsal-free open-world object detection without storing historical samples.
We propose a Dual-Stage Objectness Modeling to avoid catastrophic forgetting caused by representation drift, together with a CG-NLL distance to calibrate objectness scores and stabilize the energy landscape.
Extensive evaluations demonstrate that our approach outperforms all compared exemplar replay based methods, establishing a new and efficient baseline for rehearsal-free open-world object detection.
Open-World Object Detection (OWOD) shifts the focus from static closed-set recognition to dynamic environment adaptation[3], [4]. In this paradigm, a detector must not only identify known categories but also detect previously unseen objects as unknown, subsequently incorporating them into the known set as new labels emerge[10], [12], [14]. Joseph et al.[4] formally introduced the OWOD framework by synthesizing concepts from Open Set Recognition[3] and Incremental Object Detection[14]. To alleviate catastrophic forgetting, they employed an exemplar replay (ER) mechanism, where a set of representative samples from earlier tasks is stored and replayed during incremental updates. Many follow-up approaches, such as OW-DETR[5], PROB[13], CAT[9], and ORTH[31], continue to rely on exemplar buffers to preserve performance on prior classes during knowledge expansion.
However, storing raw data raises privacy and storage concerns like we mention above. Since current methods have not fully eliminated this dependency, a truly rehearsal-free approach is necessary.
Parameter-Efficient Fine-Tuning (PEFT) adapts large-scale models by updating only a small subset of parameters[26], [37]. Among these techniques, Low-Rank Adaptation (LoRA)[21] is a prominent approach[24], [25], [28], [38]–[41]. Formally, LoRA freezes the pre-trained weight matrix \(W_0 \in \mathbb{R}^{d \times k}\) and introduces a low-rank update \(\Delta W = BA\), where \(B \in \mathbb{R}^{d \times r}\) and \(A \in \mathbb{R}^{r \times k}\) are trainable matrices with rank \(r \ll \min(d,k)\). This strategy significantly reduces the parameter space while preserving the generalization capability of the backbone. Recent research[26], [42]–[44] focuses on enhancing LoRA through structural constraints and decoupled modularity. In object detection, some works[43], [45], [46] use LoRA for domain-specific tasks, its potential for rehearsal-free Open-World Object Detection remains largely unexplored. The need to balance universal feature refinement with task-specific isolation in a strictly non-exemplar setting presents a unique challenge for current PEFT designs in detection domains.
In OWOD, class-agnostic objectness modeling decouples object existence from category identity by estimating whether a region corresponds to any valid object instance, thereby enabling the discovery of previously unseen categories[1], [35], [47]–[49]. Early methods such as PROB[13] and OrthogonalDet[32] estimate objectness through probabilistic modeling and orthogonal decomposition. More recent works, such as PASS[50], leverage attribute subsets to refine proposal scoring, while OWOBJ[35] assumes a dynamic Gaussian prior over objectness features and scores proposals using Mahalanobis distance. However, the optimization of these methods is typically restricted to current task data, relying heavily on exemplar replay to prevent forgetting. Without data rehearsal, continuous model updates cause representation drift, which destabilizes the objectness modeling and leads to a collapse in novelty discovery.
REAL-OW is a decoupled framework that synergizes General Adapters and Specific Adapters with Dual-Stage Objectness Modeling (DSOM) for rehearsal-free OWOD. Built on D-DETR-based detectors[2], [5], this architecture balances general and task-specific knowledge refinement for each incremental task with stable feature accumulation and robust unknown discovery. Together, these components enable incremental learning without relying on historical samples.
We adopt General Adapters to support the continuous refinement of shared representations within the Pyramid Vision Transformer (PVT) backbone. As illustrated in 2, we inject LoRA modules into the query \(q\) and value \(v\) projection matrices of the self-attention layers. During training, we keep the original backbone weights \(W_{0}\) frozen and update only the lightweight GAs parameters. The forward pass for each GA can be formulated as: \[f_{G}(x) = W_0 x + A_G B_G x,\] where x is input at any task \(t\). \(B_G \in \mathbb{R}^{r \times k}\) is a fixed, random orthogonal down-projection matrix, and \(A_G \in \mathbb{R}^{d \times r}\) is a trainable up-projection matrix initialized to zero. This asymmetric design is motivated by recent work[51] showing that tuning \(B\) is more impactful than tuning \(A\). To prevent standard regularization from indiscriminately suppressing all parameters through a uniform penalty, which leads to shrinkage bias[52], [53], we propose a novel Adaptive Gated Sparse Loss (\(\mathcal{L}_{AGS}\)) to optimize the GAs. This significance-aware mechanism allows the model to selectively accumulate impactful knowledge by dynamically modulating the regularization pressure. The loss function is formulated as: \[\mathcal{L}_{AGS} = \sum_{\ell=1}^{L} {\left( \frac{1}{1 + \gamma \cdot \left( \|\mathbf{A}_\ell\|_F + \|\mathbf{B}_\ell\|_F \right)} \right)} \cdot {\left( \|\mathbf{A}_\ell\|_F + \|\mathbf{B}_\ell\|_F \right)}, \label{eq:ga}\tag{1}\] where \(L\) is the number of backbone layers. The fractional term serves as an adaptive gate, which acts as a dynamic switch by sensing the Frobenius norm (\(\|\cdot\|_F\)) of each adapter layer. Specifically, the gate reduces regularization pressure for essential layers to preserve them as stable anchors for universal knowledge, and it applies higher sparse pressure to redundant layers, driving them to optimize, adapting and capturing shared features in later tasks of the sequence. To ensure high-fidelity consistency with the teacher model as well, we employ \(\mathcal{L}_{distill}\), an output-aligned distillation loss, formulated as: \[\mathcal{L}_{distill} = \sum_{i \in \mathcal{C}_t} s_{t-1,i} \log(s_{t,i}) \label{eq:distill},\tag{2}\] where \(s_t\) represents the normalized feature distribution from backbone over task \(t\) classes \(\mathcal{C}_t\). Combined with 1 and 2 , the whole GA loss formulated as: \[\mathcal{L}_{GA} = \mathcal{L}_{distill} + \mathcal{L}_{AGS}. \label{eq:fa}\tag{3}\]
While the GAs focuses on refining cross-task universal representations in the backbone, the feature interaction part requires specialized expertise to handle task-specific category features. To this end, we employ the Specific Adapters in the decoder. As shown in 2, the SAs employ the same LoRA placement in the \(q\) and \(v\) projections and the same asymmetric initialization strategy. But for each new task, a dedicated set of SA modules is added in parallel. This incremental expansion enables the decoder to store task-specific knowledge in isolated branches.
However, this multi-branch architecture is susceptible to feature entanglement. Because LoRA updates only a small fraction of the parameters relative to the pre-trained model, discriminating between unique task features remains difficult. Current incremental methods[22], [26] often address this by using scalar control gates but these mechanisms are content-agnostic, as they ignore actual parameter magnitudes, making them insufficient to fully prevent interference between tasks. To ensure that newly added adapters do not interfere with the specialized knowledge stored during preceding task stages, we optimize the SAs using the Layer-wise Interaction Orthogonal Loss (\(\mathcal{L}_{LIO}\)). Unlike standard geometric constraints, our approach is capacity-aware by coupling the gating signal with the actual parameter intensity. The loss is formulated as: \[\mathcal{L}_{LIO} = \sum_{i=1}^{t-1} \sum_{k=1}^{N} \left( \mu_t^k \cdot \text{sg} \left[ \mu_i^k \cdot \|\mathbf{A}_i^k\|_F \right] \right)^2, \label{eq:cga}\tag{4}\] where \(N\) represents the number of layers in decoder. \(\mu_t^k\) denotes the learnable weight for task \(t\) at the \(k\)-th block, and \(\|\mathbf{A}_i^k\|_F\) represents the Frobenius norm of the adapter parameters from previous task \(i\). The operator \(\text{sg}\left[ \cdot \right]\) indicates a stop-gradient operation to protect historical weights. The product \(\mu_i^k \cdot \|\mathbf{A}_i^k\|_F\) defines the effective strength of a layer, which quantifies both its activation level and stored feature intensity. \(\mathcal{L}_{LIO}\) dynamically guides the current task to utilize layers that are either structurally inactivated or parametrically sparse in history. This design ensures that each adapter group selects a non-overlapping set of significant layers, effectively isolating task-specific knowledge and preventing the corruption of prior feature representations.
Objectness probability and energy-based modeling are essential for identifying candidate proposals and distinguishing between known objects, unknown objects, and background. In a rehearsal-free setting where replay is unavailable, feature drift occurs, which destabilizes the modeling process. To address this, we propose Dual-Stage Objectness Modeling (DSOM), which partitions the optimization process into alternating Learning and Consolidation phases. To further improve objectness quantification, we introduce CG-NLL distance.
CG-NLL Distance. As shown in 3, we observe that known category distribution becomes more dispersed over time as semantic diversity grows without exemplar replay. Standard Mahalanobis distance fails to account for this expansion, leading to inaccurate objectness estimation. To bridge the gap between
geometric distance and probabilistic objectness, we formulate the Calibrated Gaussian Negative Log-Likelihood (CG-NLL) distance to incorporate distribution density into a scaled objectness modeling. Our CG-NLL distance is defined as: \[d_C(\mathbf{x}) = \frac{(\mathbf{x} - \boldsymbol{\mu})^\top \Sigma^{-1} (\mathbf{x} - \boldsymbol{\mu})}{1 + \beta \ln(1 + |\Sigma|)}, \label{eq:distance}\tag{5}\] where \(\mu\) and \(\Sigma\) represent the mean and covariance of the known category distribution. We utilize the term \(\tau(\Sigma) = 1 + \beta \ln(1 + |\Sigma|)\) as an adaptive scaling factor to normalize the distance based on the dispersion. \(\beta\) is a calibration hyperparameter. By adaptively rescaling the Mahalanobis Distance, CG-NLL distance maintains a strict manifold for compact distributions to suppress background noise, while providing a calibration for dispersed distributions to increase tolerance against catastrophic forgetting. Based on it, for unmatched queries \(\boldsymbol{q}_i \in \mathcal{Q}_{\text{unk}}\), the objectness score \[S_{obj}^i = \exp(-d_C(\mathbf{q}_i)), \label{eq:score}\tag{6}\] serves as the confidence score for the unknown class pseudo-labels.
Learning Phase. This phase focuses on aggregating the feature representations of currently known categories and recognizes unknown objects. \(\mathcal{L}_{Learning}\) is formulated as: \[\mathcal{L}_{Learning} = \mathcal{L}_{obj} + \mathcal{L}_{cls} + \mathcal{L}_{reg}, \label{eq:l95learning}\tag{7}\] where \(\mathcal{L}_{obj} = \sum_{{q}_i \in \mathcal{Q_{\text{known}}}} d_C(q_i)\) minimizes the CG-NLL distance for matched known queries to improve objectness modeling, thereby enforcing a compact clustering of known features around the estimated distribution center. \(\mathcal{L}_{cls} = \text{FL}(\hat{l}, l)\) is the sigmoid focal loss for classification. The predictions \(\hat{l} \in \mathbb{R}^{Z+1}\) are supervised by targets \(l\), which consist of discrete known class labels \(l_{1:Z} \in \{0, 1\}^Z\) and a continuous objectness score \(l_{Z+1} \in [0, 1]\). For unmatched queries, the \(Z+1\)-th dimension is assigned soft objectness pseudo-labels \({l}_i[Z+1] = S_{\text{obj}}^i\), given by 6 , allowing the model to regard high-scoring regions as potential unknown objects. \(\mathcal{L}_{reg}\) is a combination of the \(\ell_1\) loss and the GIoU loss.
Consolidation Phase. To maintain stability and sharpen the boundary between known and unknown energy distribution for preventing misclassification, the consolidation phase employs two optimization objectives: \[\mathcal{L}_{Consolidation} = \mathcal{L}_{margin} + \mathcal{L}_{KL}, \label{bnuevthz}\tag{8}\] where \(\mathcal{L}_{margin}\) stands for energy margin loss, separating known and unknown energy distributions. \(\mathcal{L}_{margin}\) is formulated as: \[\mathcal{L}_{margin} = \mathbb{E}_{q^+ \in Q_{\text{known}}} \mathbb{E}_{q^- \in Q_{\text{unk}}} \left[ \max(0, E(q^+) - E(q^-) + \delta) \right], \label{lfikrmhg}\tag{9}\] where \(E(q) = -\log \sum_{i=1}^{C} \exp(f^i(q))\) denotes the energy level based on the classifier logits \(f^i(q)\). This hinge loss pushes the energy of unknown candidates above a margin \(\delta\) relative to knowns, establishing a discriminative surface. \(\mathcal{L}_{KL}\) stands for KL divergence loss[35], acting as a variance regularizer to prevent latent space collapse and maintaining embedding diversity, thereby mitigating over-fitting bias. Through the alternation of these two phases, DSOM achieves a stable and rehearsal-free optimization of objectness knowledge across incremental tasks.
| Task IDs (\(\rightarrow\)) | Task 1 | Task 2 | Task 3 | Task 4 | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| (\(\uparrow\)) | Current known | (\(\uparrow\)) | Previously known | Current known | Both | (\(\uparrow\)) | Previously known | Current known | Both | Previously known | Current known | Both | |
| UC-OWOD[54] | 2.4 | 50.7 | 3.4 | 33.1 | 30.5 | 31.8 | 8.7 | 28.8 | 16.3 | 24.6 | 25.6 | 15.9 | 23.2 |
| ORE[4] | 4.9 | 56.0 | 2.9 | 52.7 | 26.0 | 39.4 | 3.9 | 38.2 | 12.7 | 29.7 | 29.6 | 12.4 | 25.3 |
| 2B-OCD[55] | 12.1 | 56.4 | 9.4 | 51.6 | 25.3 | 38.5 | 11.6 | 37.2 | 13.2 | 29.2 | 30.0 | 13.3 | 25.8 |
| OCPL[49] | 8.3 | 56.6 | 7.7 | 50.6 | 27.5 | 39.1 | 11.9 | 38.7 | 14.7 | 30.7 | 30.7 | 14.4 | 26.7 |
| OW-DETR[5] | 7.5 | 59.2 | 6.2 | 53.6 | 33.5 | 42.9 | 5.7 | 38.3 | 15.8 | 30.8 | 31.4 | 17.1 | 27.8 |
| CAT[9] | 23.7 | 60.0 | 19.1 | 55.5 | 32.7 | 44.1 | 24.4 | 42.8 | 18.7 | 34.8 | 34.4 | 16.6 | 29.9 |
| PROB[13] | 19.4 | 59.5 | 17.4 | 55.7 | 32.2 | 44.0 | 19.6 | 43.0 | 22.2 | 36.0 | 35.7 | 18.9 | 31.5 |
| OWOBJ[35] | 23.6 | 61.4 | 23.8 | 58.4 | 34.4 | 45.7 | 25.1 | 44.8 | 27.8 | 38.8 | 36.4 | 20.7 | 32.0 |
| Hyp-OW[33] | 23.5 | 59.4 | 20.6 | - | - | 44.0 | 26.3 | - | - | 36.8 | - | - | 33.6 |
| RandBox[7] | 10.6 | 61.8 | 6.3 | - | - | 45.3 | 7.8 | - | - | 39.4 | - | - | 35.4 |
| ORTH[32] | 61.3 | 55.5 | 38.5 | 47.0 | 46.7 | 30.6 | 41.3 | 42.4 | 24.3 | 37.9 | |||
| Ours:REAL-OW | 25.5 | 62.6 | 26.9 | 59.4 | 38.7 | 49.0 | 29.6 | 47.7 | 31.2 | 42.2 | 42.8 | 26.7 | 38.7 |
| ORE[4] | 1.5 | 61.4 | 3.9 | 56.5 | 26.1 | 40.6 | 3.6 | 38.7 | 23.7 | 33.7 | 33.6 | 26.3 | 31.8 |
| OW-DETR[5] | 5.7 | 71.5 | 6.2 | 62.8 | 27.5 | 43.8 | 6.9 | 45.2 | 24.9 | 38.5 | 38.2 | 28.1 | 33.1 |
| PROB[13] | 17.6 | 73.4 | 22.3 | 66.3 | 36.0 | 50.4 | 24.8 | 47.8 | 30.4 | 42.0 | 42.6 | 31.7 | 39.9 |
| CAT[9] | 24.0 | 74.2 | 23.0 | 67.6 | 35.5 | 50.7 | 24.6 | 51.2 | 32.6 | 45.0 | 45.4 | 35.1 | 42.8 |
| OWOBJ[35] | 22.3 | 76.2 | 69.8 | 41.0 | 54.8 | 30.9 | 50.6 | 35.7 | 46.8 | 46.7 | 36.9 | 43.2 | |
| Hyp-OW[33] | 23.9 | 72.7 | 23.3 | - | - | 50.6 | 25.4 | - | - | 46.2 | - | - | 44.8 |
| ORTH[32] | 71.6 | 27.9 | 64.0 | 39.9 | 51.3 | 52.1 | 42.2 | 48.8 | 48.7 | 38.8 | 46.2 | ||
| Ours:REAL-OW | 24.9 | 76.9 | 29.8 | 71.5 | 42.3 | 56.9 | 33.3 | 53.4 | 44.6 | 50.5 | 49.7 | 39.3 | 47.1 |
2pt
| Methods | Trainable Parameters |
|---|---|
| Transformer-based | 39.7M |
| CNN-based | 106.6M |
| Ours | 2.0M |
12pt
| old cls. | new cls. | final mAP | 15 + 5 setting | old cls. | new cls. | final mAP | 19 + 1 setting | old cls. | new cls. | final mAP | |
|---|---|---|---|---|---|---|---|---|---|---|---|
| ILOD[17] | 63.2 | 63.1 | 63.2 | ILOD[17] | 68.3 | 58.4 | 65.8 | ILOD[17] | 68.5 | 62.7 | 68.2 |
| ORE[4] | 60.4 | 68.8 | 64.5 | ORE[4] | 71.8 | 58.7 | 68.5 | ORE[4] | 69.4 | 60.1 | 68.8 |
| OW-DETR[5] | 63.5 | 67.9 | 65.7 | OW-DETR[5] | 72.2 | 59.8 | 69.4 | OW-DETR[5] | 70.7 | 62.0 | 70.2 |
| PROB[13] | 66.0 | 67.2 | 66.5 | PROB[13] | 73.2 | 60.8 | 70.1 | PROB[13] | 73.9 | 48.5 | 72.6 |
| CAT[9] | 67.9 | 67.4 | 67.7 | CAT[9] | 75.6 | 59.3 | 72.2 | CAT[9] | 74.5 | 61.1 | 73.8 |
| OWOBJ[35] | 70.5 | 72.0 | 69.9 | OWOBJ[35] | 76.5 | 63.7 | 73.3 | OWOBJ[35] | 76.6 | 53.8 | 75.8 |
| ORTH[32] | 74.5 | 70.2 | 72.3 | ORTH[32] | 78.3 | 71.8 | 74.7 | ORTH[32] | 75.6 | 74.9 | 75.6 |
| Ours:REAL-OW | 73.8 | 74.3 | 74.0 | Ours:REAL-OW | 81.2 | 68.1 | 75.9 | Ours:REAL-OW | 78.2 | 60.7 | 77.3 |
2.5pt
Datasets. We evaluate REAL-OW on M-OWODB [4] and S-OWODB [5], derived from MS-COCO [56] and PASCAL VOC [57]. S-OWODB groups classes by their super-categories like animals or vehicles, which is more challenging. Both benchmarks comprise four incremental
tasks (\(T_1\) to \(T_4\)), with 20 new classes introduced per task. During training, only labels for the current task are provided, while testing requires detecting all previously
encountered categories. Most importantly, our approach is strictly rehearsal-free and stores no historical samples.
Evaluation Metrics. We use mean Average Precision (mAP) to measure detection accuracy on known classes, and Unknown Recall (U-Recall) to evaluate the ability to discover unseen objects. Additionally, A-OSE and WI are employed to quantify
misclassifications between known and unknown categories.
Implementation Details. We implement our framework in PyTorch and conduct all experiments on 3 NVIDIA RTX A\(6000\) GPUs. Our model uses PVT as the backbone and Deformable DETR as the detection framework.
Following standard practice, we pre-train the model on ImageNet [58] using a self-supervised approach to prevent premature exposure to
future unknown classes [5]. For a fair comparison, this pre-training setup is applied to all evaluated methods. The decoder consists of 6 layers. We
set the ranks for the GAs and SAs to 16 and 10, respectively. Within the DSOM module, each iteration employs 2 learning cycles and 10 consolidation cycles. The hyperparameters \(\delta\), \(\gamma\), \(\beta\) are set to \(0.2\), \(0.01\), \(0.1\) respectively. To ensure the
reliability, all reported results represent the average performance calculated over five different random seeds.
For a fair comparison, we focus on models that generate unknown pseudo-labels internally, excluding approaches that leverage external large models, such as vision language models (VLMs), to provide pseudo-label assignments.
Quantitative. As summarized in 1, our REAL-OW framework achieves state-of-the-art performance on both M-OWODB and S-OWODB benchmarks. Notably, REAL-OW is the only strictly rehearsal-free method among all compared techniques. Despite this constraint, our model consistently achieves the top performance across all results, surpassing heavily parameterized CNN-based (e.g.RandBox, ORTH) and Transformer-based (e.g.CAT, PROB) methods. As shown in 2, these baseline methods require over 50 times and nearly 20 times more trainable parameters than our approach, respectively. Furthermore, REAL-OW reaches the highest precision in Task 4, proving its superior resistance to catastrophic forgetting without any data replay. The performance gap between REAL-OW and OWOBJ also suggests that our dual-stage modeling is more effective than single-stage with known-unknown separation. REAL-OW also achieves the best results in A-OSE and WI among all compared methods. Refer to 7 in supplementary material for more details.
We evaluate REAL-OW on three standard iOD settings to assess knowledge retention. As shown in 3, our framework achieves the highest final mAP across all settings, outperforming SOTA methods. These results validate our architecture: GAs maintain stable feature foundations, while SAs and DSOM ensure task isolation and energy stability. Collectively, these components enable robust rehearsal-free learning with superior knowledge retention and exceptional resistance to catastrophic forgetting. Per-class details are in supplementary material 6.
Qualitative. As illustrated in 4, REAL-OW demonstrates higher U-Recall and more precise localization. It successfully discovers additional novel objects missed by other methods, such as the curtain in column \(1\) and ladder in column \(3\), while accurately separating individual cups in column \(3\). These improvements show the CG-NLL distance effectively quantifies unknown objectness. Additionally, our model avoids the misclassification errors seen in CAT (tvmonitor, column \(2\)) and OWOBJ (pottedplant, column \(1\)). High confidence scores for these detections confirm that DSOM maintains a stable energy landscape for reliable boundary separation. For more visualization results on the strong resistance of REAL-OW to catastrophic forgetting and its superior capability for discovering unknown classes, please refer to supplementary material 8.
LoRA. We evaluate various LoRA rank configurations for the GAs and SAs. Following previous PEFT methods [29], [43], [53], we apply LoRA to \(q\) and \(v\) projections. Further validation for this placement provided in supplementary material 9. As shown in 4, we observe that performance does not scale linearly with the increase in trainable parameters. Instead, the configuration (\(R_{GA}=16, R_{SA}=8\)) achieves the optimal trade-off between performance and parameter capacity. Excessively high ranks lead to a decline in U-Recall and offer only marginal gains for mAP. This degradation is more pronounced as \(R_{GA}\) increases, suggesting that an over-parameterized general feature space induces an unknown decision bias.
| GAs \(\downarrow\) SAs \(\rightarrow\) (rank) | \(r = 1\) | \(r = 4\) | \(r = 8\) | \(r = 16\) | \(r = 32\) | \(r = 64\) | ||||||
| U-Recall\(\uparrow\) | mAP\(\uparrow\) | U-Recall\(\uparrow\) | mAP\(\uparrow\) | U-Recall\(\uparrow\) | mAP\(\uparrow\) | U-Recall\(\uparrow\) | mAP\(\uparrow\) | U-Recall\(\uparrow\) | mAP\(\uparrow\) | U-Recall\(\uparrow\) | mAP\(\uparrow\) | |
| \(r = 1\) | \(23.32\) | \(35.31\) | \(24.88\) | \(38.69\) | \(26.92\) | \(40.37\) | \(27.61\) | \(40.44\) | \(27.90\) | \(39.23\) | \(27.88\) | \(38.84\) |
| \(r = 4\) | \(24.18\) | \(36.28\) | \(26.14\) | \(39.60\) | \(27.50\) | \(41.86\) | \(28.47\) | \(41.75\) | \(28.83\) | \(41.02\) | \(28.90\) | \(40.54\) |
| \(r = 8\) | \(24.98\) | \(36.90\) | \(27.03\) | \(39.72\) | \(28.82\) | \(42.40\) | \(28.90\) | \(42.33\) | \(28.85\) | \(42.18\) | \(28.87\) | \(41.63\) |
| \(r = 16\) | \(25.32\) | \(37.78\) | \(28.11\) | \(40.80\) | \(\mathbf{29.64}\) | \(\mathbf{42.23}\) | \(29.88\) | \(42.11\) | \(29.65\) | \(41.20\) | \(29.70\) | \(40.75\) |
| \(r = 32\) | \(24.64\) | \(38.57\) | \(28.46\) | \(41.75\) | \(28.70\) | \(42.33\) | \(29.42\) | \(42.08\) | \(29.48\) | \(41.86\) | \(29.60\) | \(39.50\) |
| \(r = 64\) | \(22.76\) | \(38.61\) | \(26.21\) | \(41.81\) | \(27.54\) | \(42.21\) | \(28.72\) | \(41.97\) | \(28.32\) | \(40.22\) | \(28.38\) | \(38.43\) |
3pt
| Task IDs (\(\rightarrow\)) | Task 1 | Task 2 | Task 3 | Task 4 | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| U-Recall | mAP (\(\uparrow\)) | U-Recall | mAP (\(\uparrow\)) | U-Recall | mAP (\(\uparrow\)) | mAP (\(\uparrow\)) | |||||||
| (\(\uparrow\)) | Current known | (\(\uparrow\)) | Previously known | Current known | Both | (\(\uparrow\)) | Previously known | Current known | Both | Previously known | Current known | Both | |
| w/o \(\mathcal{L}_{AGS}\) | 26.7 | 63.2 | 25.4 | 58.6 | 38.9 | 48.7 | 28.0 | 46.5 | 28.5 | 40.5 | 40.2 | 25.8 | 36.6 |
| w/ \(\mathcal{L}_{AGS}\) | 25.5 | 62.6 | 26.9 | 59.4 | 38.7 | 49.0 | 29.6 | 47.7 | 31.2 | 42.2 | 42.8 | 26.7 | 38.7 |
2pt
Adaptive Gated Sparse Loss. 5 compares \(\mathcal{L}_{AGS}\) with standard distillation \(\mathcal{L}_{distill}\) that relies only on output alignment. As expected, during the initial task, \(\mathcal{L}_{AGS}\) lags slightly behind the baseline due to its sparsity constraint. However, by suppressing redundant layers, this constraint ensures that critical layers carry general features applicable to the encountered known categories, while other layers remain plastic. The performances on later tasks show the superior refinement of \(\mathcal{L}_{AGS}\). 5 further confirms this targeted update pattern that critical layers preserve foundational representations with \(\mathcal{L}_{AGS}\), while standard distillation leads to a uniform and undifferentiated weight distribution.
| Methods | Task 1 | Task 2 | ||||
| U-Recall | mAP | U-Recall | Prev. | Curr. | Both | |
| D-DETR-upper limit | \(32.9\) | \(63.8\) | \(40.5\) | \(62.9\) | \(41.2\) | \(52.1\) |
| D-DETR-ER | \(-\) | \(60.3\) | \(-\) | \(54.5\) | \(34.4\) | \(44.7\) |
| D-DETR-RF | \(-\) | \(61.9\) | \(-\) | \(57.9\) | \(38.5\) | \(48.2\) |
| REAL-OW w/o CG-NLL | \(23.3\) | \(58.3\) | \(25.1\) | \(57.5\) | \(37.9\) | \(47.7\) |
| REAL-OW w/o Learning | \(25.4\) | \(60.5\) | \(26.0\) | \(47.8\) | \(37.5\) | \(42.7\) |
| REAL-OW w/o Consolidation | \(20.1\) | \(63.1\) | \(21.0\) | \(58.9\) | \(39.5\) | \(49.2\) |
| REAL-OW - complete | \(\mathbf{25.5}\) | \(\mathbf{62.6}\) | \(\mathbf{26.9}\) | \(\mathbf{59.4}\) | \(\mathbf{38.7}\) | \(\mathbf{49.0}\) |
0pt 0pt
6pt
Dual-Stage Objectness Modeling. 6 validates the core components of our DSOM. (1) CG-NLL Distance. Replacing CG-NLL distance with the standard Mahalanobis distance reduces U-Recall, demonstrating that standard metrics fail to calibrate energy scores as features disperse. This result is further explained in 6. By accounting for distribution dispersion, CG-NLL facilitates more accurate objectness modeling. This optimization enables the model to establish a significantly more distinct energy boundary between known and unknown classes.
(2) The Learning and Consolidation Phases. Removing the learning phase causes previous mAP to drop from 59.4% to 47.8%, proving that feature aggregation is essential to prevent new knowledge from overwriting old representations. Conversely, skipping the consolidation phase slightly inflates current precision but collapses U-Recall from 26.9% to 21.0%. This indicates that without boundary separation, unknown objects are incorrectly biased toward known distributions and causes misclassification of the known and unknown.
(3) Cycle Ratios. 7 illustrates the performance after Task \(2\) under different cycle ratios of two phases. While mAP increases with more learning cycles, U-Recall peaks at a 2:10 ratio before declining. This confirms that while the learning phase maintains precision, allocating too many cycles to it creates a bias that hinders unknown discovery. We select 2:10 as the optimal balance to ensure both high stability and robust novelty detection.
In this paper, we propose REAL-OW, a novel rehearsal-free framework for Open-World Object Detection. REAL-OW decouples incremental knowledge through a collaborative adapter architecture, utilizing General Adapters for significance-aware feature refinement and Specific Adapters for orthogonal task isolation. To counteract representation drift, we introduce Dual-Stage Objectness Modeling (DSOM) supported by CG-NLL distance, which ensures robust objectness modeling and unknown discovery without data replay. Extensive experiments validate that REAL-OW outperforms all compared SOTA methods while maintaining a strictly rehearsal-free constraint. The continuous stability of DSOM provides a reliable foundation for the future integration of external large models like vision-language models (VLMs), enabling the generation of high-quality pseudo-labels for future open-world research.
This work was supported by the National Natural Science Foundation of China (Grant No. 62176163), the Shenzhen Higher Education Stable Support Program General Project (Grant No. 20231120175215001), the Scientific Foundation for Youth Scholars of Shenzhen University, the National Natural Science Foundation of China (Grant No. 62576218), the Guangdong Provincial Key Laboratory (Grant No. 2023B1212060076), and the Intelligent Computing Center of Shenzhen University.
This supplementary material provides additional details and experimental evidence to support our main paper. We first present more detailed results for incremental object detection (iOD) and evaluate the model performance using the Wilderness Index (WI) and Absolute Open-Set Error (A-OSE). We also provide visualizations to demonstrate the forgetting resistance and discovery capabilities of REAL-OW. Finally, we include ablation studies to validate the placement of LoRA within attention projections.
Previous works[5], [13] utilize incremental object detection (iOD) to evaluate model stability and knowledge retention. We assess the performance of REAL-OW on these iOD tasks, and the results are detailed in 7. We compare our framework against current state-of-the-art methods, all of which rely on exemplar replay. In each tested setting, the model is first trained on a base subset of categories (10, 15, or 19 classes), followed by incremental learning on the additional classes (10, 5, or 1 class). As shown in 7, our model achieves the highest mAP of 74.0, 75.9, and 77.3 across the three scenarios. Remarkably, REAL-OW outperforms these SOTA methods while remaining strictly rehearsal-free. These results demonstrate the exceptional robustness of our approach in incremental learning without any access to historical data. This performance validates the effectiveness of our collaborative adapter architecture. In a rehearsal-free setting, General Adapters (GAs) update universal knowledge while Specific Adapters (SAs) preserve task-specific expertise. This collaborative design significantly mitigates catastrophic forgetting. Furthermore, the Dual-Stage Objectness Modeling (DSOM) phase establishes a stable and discriminative energy distribution, which allows the model to clearly separate known categories from unknown objects throughout the learning process.
1 in the main paper presents the performance of REAL-OW in mAP and U-Recall. To further analyze the detection quality, 8 provides a comparison using Wilderness Index (WI) and Absolute Open-Set Error (A-OSE). Specifically, WI measures the drop in precision for known categories when unknown objects are introduced into the evaluation. A-OSE counts the number of unknown instances that the model misclassifies as known classes. Together, these metrics quantify how often a model confuses unknown instances with background or known categories. Notably, REAL-OW achieves the highest U-Recall while consistently maintaining the lowest WI and A-OSE across all task stages. These results indicate that our model effectively discovers novel objects without compromising the detection precision of known categories. Furthermore, it significantly reduces the misclassification of unknown instances as known classes. Such performance validates that the DSOM phase successfully establishes discriminative energy boundaries. Remarkably, REAL-OW maintains this robustness under strict rehearsal-free constraints, whereas most exemplar-replay-based methods struggle with error accumulation.
| 10 + 10 setting | aero | cycle | bird | boat | bottle | bus | car | cat | chair | cow | table | dog | horse | bike | person | plant | sheep | sofa | train | tv | mAP |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| ILOD[17] | 69.9 | 70.4 | 69.4 | 54.3 | 48.0 | 68.7 | 78.9 | 68.4 | 45.5 | 58.1 | 59.7 | 72.7 | 73.5 | 73.2 | 66.3 | 29.5 | 63.4 | 61.6 | 69.3 | 62.2 | 63.2 |
| ORE[4] | 63.5 | 70.9 | 58.9 | 42.9 | 34.1 | 76.2 | 80.7 | 76.3 | 34.1 | 66.1 | 56.1 | 70.4 | 80.2 | 72.3 | 81.8 | 42.7 | 71.6 | 68.1 | 77.0 | 67.7 | 64.5 |
| OW-DETR[5] | 61.8 | 69.1 | 67.8 | 45.8 | 47.3 | 78.3 | 78.4 | 78.6 | 36.2 | 71.5 | 57.5 | 75.3 | 76.2 | 77.4 | 79.5 | 40.1 | 66.8 | 66.3 | 75.6 | 64.1 | 65.7 |
| PROB[13] | 70.4 | 75.4 | 67.3 | 48.1 | 55.9 | 73.5 | 78.5 | 75.4 | 42.8 | 72.2 | 64.2 | 73.8 | 76.0 | 74.8 | 75.3 | 40.2 | 66.2 | 73.3 | 64.4 | 64.0 | 66.5 |
| CAT[9] | 76.5 | 75.7 | 67.0 | 51.0 | 62.4 | 73.2 | 82.3 | 83.7 | 42.7 | 64.4 | 56.8 | 74.1 | 75.8 | 79.2 | 78.1 | 39.9 | 65.1 | 59.6 | 78.4 | 67.4 | 67.7 |
| OWOBJ[35] | 75.9 | 80.7 | 73.3 | 52.1 | 58.8 | 77.7 | 81.9 | 79.1 | 47.9 | 77.8 | 70.4 | 78.9 | 80.3 | 79.3 | 80.0 | 45.1 | 70.2 | 78.4 | 68.5 | 68.4 | 69.9 |
| ORTH[32] | 82.4 | 77.3 | 78.2 | 59.7 | 61.2 | 84.3 | 90.1 | 80.2 | 49.8 | 81.7 | 58.2 | 74.0 | 82.9 | 81.0 | 81.2 | 38.3 | 70.8 | 68.0 | 77.4 | 70.2 | 72.3 |
| Ours: REAL-OW | 78.3 | 79.5 | 74.4 | 55.5 | 62.7 | 88.8 | 87.3 | 81.0 | 51.3 | 79.2 | 75.4 | 77.5 | 82.4 | 81.5 | 80.6 | 48.7 | 71.0 | 79.2 | 75.5 | 70.8 | 74.0 |
| 15 + 5 setting | aero | cycle | bird | boat | bottle | bus | car | cat | chair | cow | table | dog | horse | bike | person | plant | sheep | sofa | train | tv | mAP |
| ILOD[17] | 70.5 | 79.2 | 68.8 | 59.1 | 53.2 | 75.4 | 79.4 | 78.8 | 46.6 | 59.4 | 59.0 | 75.8 | 71.8 | 78.6 | 69.6 | 33.7 | 61.5 | 63.1 | 71.7 | 62.2 | 65.8 |
| ORE[4] | 75.4 | 81.0 | 67.1 | 51.9 | 55.7 | 77.2 | 85.6 | 81.7 | 46.1 | 76.2 | 55.4 | 76.7 | 86.2 | 78.5 | 82.1 | 32.8 | 63.6 | 54.7 | 77.7 | 64.6 | 68.5 |
| OW-DETR[5] | 77.1 | 76.5 | 69.2 | 51.3 | 61.3 | 79.8 | 84.2 | 81.0 | 49.7 | 79.6 | 58.1 | 79.0 | 83.1 | 67.8 | 85.4 | 33.2 | 65.1 | 62.0 | 73.9 | 65.0 | 69.4 |
| PROB[13] | 77.9 | 77.0 | 77.5 | 56.7 | 63.9 | 75.0 | 85.5 | 82.3 | 50.0 | 78.5 | 63.1 | 75.8 | 80.0 | 78.3 | 77.2 | 38.4 | 69.8 | 57.1 | 73.7 | 64.9 | 70.1 |
| CAT[9] | 75.3 | 81.0 | 84.4 | 64.5 | 56.6 | 74.4 | 84.1 | 86.6 | 53.0 | 70.1 | 72.4 | 83.4 | 85.5 | 81.6 | 81.0 | 32.0 | 58.6 | 60.7 | 81.6 | 63.5 | 72.2 |
| OWOBJ[35] | 82.8 | 80.0 | 82.4 | 60.1 | 68.0 | 79.9 | 90.0 | 86.4 | 54.4 | 83.1 | 64.2 | 77.3 | 85.1 | 80.3 | 80.1 | 42.1 | 73.2 | 61.8 | 77.9 | 68.7 | 73.3 |
| ORTH[32] | 82.7 | 80.4 | 78.5 | 55.3 | 65.5 | 81.0 | 89.8 | 85.9 | 52.6 | 84.6 | 62.3 | 78.4 | 82.7 | 81.1 | 84.2 | 46.5 | 71.6 | 79.0 | 82.5 | 79.2 | 74.7 |
| Ours: REAL-OW | 82.2 | 81.5 | 84.6 | 57.8 | 67.5 | 83.3 | 92.3 | 88.8 | 56.5 | 85.0 | 63.1 | 78.3 | 86.2 | 83.2 | 87.8 | 45.0 | 72.2 | 65.9 | 85.5 | 72.1 | 75.9 |
| 19 + 1 setting | aero | cycle | bird | boat | bottle | bus | car | cat | chair | cow | table | dog | horse | bike | person | plant | sheep | sofa | train | tv | mAP |
| ILOD[17] | 69.4 | 79.3 | 69.5 | 57.4 | 45.4 | 78.4 | 79.1 | 80.5 | 45.7 | 76.3 | 64.8 | 77.2 | 80.8 | 77.5 | 70.1 | 42.3 | 67.5 | 64.4 | 76.7 | 62.7 | 68.2 |
| ORE[4] | 67.3 | 76.8 | 60.0 | 48.4 | 58.8 | 81.1 | 86.5 | 75.8 | 41.5 | 79.6 | 54.6 | 72.8 | 85.9 | 81.7 | 82.4 | 44.8 | 75.8 | 68.2 | 75.7 | 60.1 | 68.8 |
| OW-DETR[5] | 70.5 | 77.2 | 73.8 | 54.0 | 55.6 | 79.0 | 80.8 | 80.6 | 43.2 | 80.4 | 53.5 | 77.5 | 89.5 | 82.0 | 74.7 | 43.3 | 71.9 | 66.6 | 79.4 | 62.0 | 70.2 |
| PROB[13] | 80.3 | 78.9 | 77.6 | 59.7 | 63.7 | 75.2 | 86.0 | 83.9 | 53.7 | 82.8 | 66.5 | 82.7 | 80.6 | 83.8 | 77.9 | 48.9 | 74.5 | 69.9 | 77.6 | 48.5 | 72.6 |
| CAT[9] | 86.0 | 85.8 | 78.8 | 65.3 | 61.3 | 71.4 | 84.8 | 84.8 | 52.9 | 78.4 | 71.6 | 82.7 | 83.8 | 81.2 | 80.7 | 43.7 | 75.9 | 58.5 | 85.2 | 61.1 | 73.8 |
| OWOBJ[35] | 86.1 | 83.9 | 83.4 | 62.9 | 65.9 | 79.9 | 90.6 | 87.3 | 56.9 | 86.5 | 70.3 | 85.9 | 84.7 | 86.9 | 81.6 | 51.9 | 78.8 | 73.5 | 80.7 | 53.8 | 75.8 |
| ORTH[32] | 83.8 | 84.7 | 77.0 | 62.9 | 60.8 | 80.9 | 88.6 | 85.8 | 51.1 | 81.4 | 67.2 | 86.7 | 86.3 | 83.4 | 83.4 | 44.7 | 74.5 | 73.1 | 81.1 | 74.9 | 75.6 |
| Ours: REAL-OW | 85.5 | 82.1 | 84.0 | 63.7 | 66.6 | 80.4 | 90.4 | 86.9 | 58.2 | 88.3 | 70.9 | 85.9 | 80.5 | 85.7 | 83.4 | 53.2 | 79.0 | 72.9 | 81.5 | 60.7 | 77.3 |
| Task IDs (\(\rightarrow\)) | Task 1 | Task 2 | Task 3 | ||||||
|---|---|---|---|---|---|---|---|---|---|
| U-Recall | WI | A-OSE | U-Recall | WI | A-OSE | U-Recall | WI | A-OSE | |
| (\(\uparrow\)) | (\(\downarrow\)) | (\(\downarrow\)) | (\(\uparrow\)) | (\(\downarrow\)) | (\(\downarrow\)) | (\(\uparrow\)) | (\(\downarrow\)) | (\(\downarrow\)) | |
| ORE[4] | 4.9 | 0.0621 | 10459 | 2.9 | 0.282 | 10445 | 3.9 | 0.0211 | 7990 |
| 2B-OCD[55] | 12.1 | 0.0481 | - | 9.4 | 0.160 | - | 11.6 | 0.0137 | - |
| OW-DETR[5] | 7.5 | 0.0571 | 10240 | 6.2 | 0.0278 | 8441 | 5.7 | 0.0156 | 6803 |
| OCPL[49] | 8.3 | 0.0413 | 5670 | 7.6 | 0.0220 | 5690 | 11.9 | 0.0162 | 5166 |
| PROB[13] | 19.4 | 0.0569 | 5195 | 17.4 | 0.0344 | 6452 | 19.6 | 0.0151 | 2641 |
| OWOBJ[35] | 23.6 | 0.0395 | 3102 | 23.8 | 0.0215 | 4032 | 25.1 | 821 | |
| REAL-OW | 25.5 | 0.0307 | 2885 | 26.9 | 0.0198 | 3289 | 29.6 | 0.0079 | 711 |
5pt
| \(R_{GA}\):\(R_{SA}\)\(\to\) Attention Layers\(\downarrow\) | 4 : 1 | 4 : 4 | 8 : 4 | 8 : 8 | 16 : 8 | 32 : 16 | 64 : 32 | |||||||
| U-Recall\(\uparrow\) | mAP\(\uparrow\) | U-Recall\(\uparrow\) | mAP\(\uparrow\) | U-Recall\(\uparrow\) | mAP\(\uparrow\) | U-Recall\(\uparrow\) | mAP\(\uparrow\) | U-Recall\(\uparrow\) | mAP\(\uparrow\) | U-Recall\(\uparrow\) | mAP\(\uparrow\) | U-Recall\(\uparrow\) | mAP\(\uparrow\) | |
| \(\{ v \}\) | 23.24 | 32.64 | 23.75 | 35.21 | 25.07 | 37.16 | 26.71 | 38.81 | 27.34 | 39.47 | 27.40 | 39.22 | 27.35 | 38.45 |
| \(\{k,v\}\) | 24.37 | 36.20 | 26.02 | 39.11 | 26.52 | 39.83 | 28.35 | 41.30 | 28.97 | 41.83 | 29.38 | 41.66 | 28.10 | 40.03 |
| \(\{q,v\}\) | 24.18 | 36.28 | 26.14 | 39.60 | 27.03 | 39.72 | 28.82 | 42.40 | 29.64 | 42.23 | 29.42 | 42.08 | 28.32 | 40.22 |
| \(\{q,k,v\}\) | 24.23 | 36.65 | 26.26 | 38.27 | 27.33 | 38.10 | 29.22 | 40.45 | 29.50 | 41.07 | 29.17 | 41.48 | 26.84 | 38.67 |
4pt
We further illustrate the performance of REAL-OW in 8, specifically focusing on its robustness in resisting catastrophic forgetting and its effectiveness in transitioning unknown objects into known categories in the task sequence.
Resistance to Catastrophic Forgetting. Columns 1 and 2 illustrate the performance of REAL-OW on long-term knowledge retention. For the initial classes identified in Task 1, competing methods show clear signs of forgetting by Task 4. Specifically, CAT fails to detect the bottle, and OWOBJ loses the bird. In contrast, REAL-OW persistently identifies these categories across incremental stages. This demonstrates significant resistance to catastrophic forgetting and maintains high localization precision without the need for data rehearsal.
Learning Unknown Categories. Columns 3 and 4 visualize the process of discovering novel objects and learning them incrementally. In Task 1, REAL-OW successfully identifies the book, keyboard, and mouse as unknown candidates. By Task 4, the model learns these objects as newly introduced known categories. Conversely, CAT fails to detect certain mouse and keyboard regions as unknowns in Task 1. Meanwhile, OWOBJ forgets its Task 1 unknown labels by the final stage, failing to recognize them as known classes in Task 4. These results confirm that our framework effectively discovers novel information and learns to recognize it in subsequent tasks.
Building on the analysis of 4 in the main paper, we observe that optimal performance is concentrated in the asymmetric rank regime where \(R_{GA} > R_{SA}\). This suggests that the backbone necessitates a higher parameter capacity to effectively refine and store universal representations, yielding a superior trade-off between accuracy and efficiency. Consequently, we fix the ranks within this optimal range to further investigate the impact of adapter placement within the Transformer architecture, as illustrated in 9. According to these results, \(\{q, v\}\) configuration remains the most robust across nearly all rank settings, validating its general superiority. We find that adding more adapter positions does not necessarily improve performance. Specifically, as the rank increases, using all three projections (\(\{q, k, v\}\)) yields lower accuracy than the \(\{q, v\}\) or \(\{k, v\}\) alternatives. This suggests that redundant projections introduce excessive task-specific noise and cause overfitting, which compromises stability and accelerates forgetting. Thus, the \(\{q, v\}\) placement provides the most stable and parameter-efficient performance for our framework.
| Methods | Task 1 | Task 2 | ||||
| U-Recall | mAP | U-Recall | Prev. | Curr. | Both | |
| w/o \(S_{obj}\) | 10.8 | 55.7 | 11.2 | 53.1 | 33.9 | 43.5 |
| w/o \(\mathcal{L}_{AGS}\) | 26.7 | 63.2 | 25.4 | 58.6 | 38.9 | 48.7 |
| w/o \(\mathcal{L}_{distill}\) | 23.9 | 61.2 | 25.8 | 57.3 | 37.8 | 47.5 |
| w/o \(\mathcal{L}_{LIO}\) | 24.5 | 62.2 | 23.6 | 55.8 | 37.6 | 46.7 |
| w/o \(\mathcal{L}_{KL}\) | 20.4 | 61.9 | 21.8 | 57.5 | 38.0 | 47.7 |
| w/o \(\mathcal{L}_{margin}\) | 20.8 | 62.3 | 22.3 | 58.0 | 37.3 | 47.6 |
| REAL-OW (Ours) | 25.5 | 62.6 | 26.9 | 59.4 | 38.7 | 49.0 |
3.0pt
10 presents a comprehensive ablation of REAL-OW on M-OWODB. Removing \(S_{obj}\) causes Task 1 U-Recall to drop from 25.5% to 10.8%, underscoring its fundamental role in reliable unknown pseudo-labeling. Without \(\mathcal{L}_{AGS}\), Task 2 performance degrades despite a marginal initial gain, confirming its importance in preserving universal representations for long-term stability. The exclusion of \(\mathcal{L}_{LIO}\) results in a sharp Task 2 mAP drop from 49.0% to 46.7% due to severe feature entanglement, highlighting its critical function in enforcing task-specific isolation within the decoder. Finally, \(\mathcal{L}_{KL}\) and \(\mathcal{L}_{margin}\) jointly stabilize the latent objectness manifold and establish a discriminative energy margin to effectively separate known categories from unknowns.