October 03, 2022
Unsupervised Domain Adaptation (UDA) aims at classifying unlabeled target images leveraging source labeled ones. In this work, we consider the Partial Domain Adaptation (PDA) variant, where we have extra source classes not present in the target domain. Most successful algorithms use model selection strategies that rely on target labels to find the best hyper-parameters and/or models along training. However, these strategies violate the main assumption in PDA: only unlabeled target domain samples are available. Moreover, there are also inconsistencies in the experimental settings - architecture, hyper-parameter tuning, number of runs - yielding unfair comparisons. The main goal of this work is to provide a realistic evaluation of PDA methods with the different model selection strategies under a consistent evaluation protocol. We evaluate 7 representative PDA algorithms on 2 different real-world datasets using 7 different model selection strategies. Our two main findings are: (i) without target labels for model selection, the accuracy of the methods decreases up to 30 percentage points; (ii) only one method and model selection pair performs well on both datasets. Experiments were performed with our PyTorch framework, BenchmarkPDA, which we open source.
Domain adaptation. Deep neural networks are highly successful in image recognition for in-distribution samples [1] with this success being intrinsically tied to the large number of labeled training data. However, they tend to not generalize as well on images with different background or colors not seen during training. Such shift in the samples is referred to as domain shift in the literature. Unfortunately, enriching the training set with new samples from different domains is challenging as labeling data is both an expensive and time-consuming task. Thus, researchers have focused on unsupervised domain adaptation (UDA) where we have access to unlabelled samples from a different domain, known as the target domain. The purpose of UDA is to classify these unlabeled samples by leveraging the knowledge given by the labeled samples from the source domain [2], [3]. In the standard UDA problem, the source and target domains are assumed to share the same classes. In this paper, we consider a more challenging variant of the problem called partial domain adaptation (PDA): the classes in the target domain \(\mathcal{Y}_t\) form a subset of the classes in the source domain \(\mathcal{Y}_s\) [4], i.e., \(\mathcal{Y}_t \subset \mathcal{Y}_s\). The number of target classes is unknown as we do not have access to the labels. The extra source classes, not present in the target domain, make the PDA problem more difficult: simply aligning the source and target domains forces a negative transfer where target samples are matched to outlier source-only labels.
Realistic evaluations. Most recent PDA methods report an increase of the target accuracy up to 15 percentage points on average when compared to the baseline approach that uses only source domain samples. While these successes constitute important breakthroughs in the DA research literature, target labels are used for model selection, violating the main UDA assumption. In their absence, the effectiveness of PDA methods remains unclear and model selection constitutes a yet to be solved problem as we show in this work. Moreover, the hyper-parameter tuning is either unknown or lacks details and sometimes requires labeled target data, which makes it challenging to apply PDA methods to new datasets (see Table 2 for a full summary). Recent work has highlighted the importance of model selection in the presence of domain shift. [5] showed that when evaluating domain generalization (DG) algorithms, whose goal is to generalize to a completely unseen domain, in a consistent and realistic setting no method outperforms the baseline ERM method by more than 1 percentage point. They argue that DG methods without a model selection strategy remain incomplete and should therefore be specified as part of the method. [6] make the same recommendation in the context of UDA and PDA.
Model selection strategies have been designed in a research environment but they have not been tested extensively with a realistic and fair experimental protocol. In this work, we study their applicability to real-world settings and evaluate 7 different PDA methods with 7 different model selection strategies on 2 different datasets (see Table 3 for a summary). We reimplemented all our considered PDA methods with the same architecture, optimizer and learning rate schedule for comparison purposes. We list below our major findings:
The accuracy attained by models selected without target labels can decrease up to 30 percentage points compared to the one reported using target labels (See Table 1 for a summary of the results).
Out of the 49 model selection strategies and PDA methods pairs considered, only one gave consistent results over all tasks on both datasets.
Random seed plays an important role in the selection of hyper-parameters. Selected parameters are not stable across different seeds and the standard deviation between accuracies on the same task can be up to \(8.4\%\) even when relying on target labels for model selection.
Under a more realistic scenario where some target labels are available, 100 random samples is enough to see only a drop of 1 percentage point in accuracy (when compared to using all target samples). However, the extreme case of using only one labeled target sample per class leads to significant drop in performance.
Outline. In Section 2, we provide an overview of the different model selection strategies considered in this work. Then in Section 3, we discuss the PDA methods that we consider. In Section 4 we describe the training procedures, hyper-parameter tuning and evaluation protocols used to evaluate all methods fairly. In Section 5, we discuss the results of the different benchmarked methods and the performance of the different model selection strategies. Finally in Section 6, we give some recommendations for future work in partial domain adaptation.
| Dataset | Model Selection | s. only | pada | safn | ba3us | ar | jumbot | mpot |
|---|---|---|---|---|---|---|---|---|
| office- | Worst (w/o target labels) | 59.55 (-2.31) | 52.72 (-11.00) | 61.37 (-1.93) | 62.25 (-13.73) | 64.32 (-8.42) | 61.28 (-15.87) | 46.92 (-30.38) |
| Best (w/o target labels) | 60.73 (-1.14) | 63.08 (-0.64) | 62.59 (-0.71) | 75.37 (-0.61) | 70.58 (-2.16) | 74.61 (-2.54) | 66.24 (-11.07) | |
| oracle | 61.87 | 63.72 | 63.30 | 75.98 | 72.73 | 77.15 | 77.31 | |
| visda | Worst (w/o target labels) | 55.02 (-4.46) | 32.32 (-22.26) | 42.83 (-19.81) | 51.07 (-16.60) | 55.69 (-18.15) | 59.86 (-24.15) | 61.62 (-25.33) |
| Best (w/o target labels) | 55.24 (-4.24) | 56.83 (2.26) | 58.62 (-4.02) | 65.58 (-2.09) | 67.20 (-6.65) | 77.69 (-6.31) | 78.40 (-8.54) | |
| oracle | 59.48 | 54.57 | 62.64 | 67.67 | 73.85 | 84.01 | 86.95 |
Model selection (choosing hyper-parameters, training checkpoints, neural network architectures) is a crucial part of training neural networks. In the supervised learning setting, a validation set is used to estimate the model’s accuracy. However, in UDA such approach is not possible as we have unlabeled target samples. Several strategies have been designed to address this issue. Below, we discuss the ones used in this work.
Source Accuracy (s-acc). [7] used the accuracy estimated on a small validation set from the source domain to perform the model selection. While the source and target accuracies are related, there are no theoretical guarantees. [8] showed that when the domain gap is large this approach fails to select competitive models.
Deep Embedded Validation (dev). [9] and [10] perform model selection through Importance-Weighted Cross-Validation (IWCV). Under the assumption that the source and target domain follow a covariate shift, the target risk can be estimated from the source risk through importance weights that give increased importance to source samples that are closer to target samples. These importance weights correspond to the ratio of the target and source densities and are estimated using Gaussian kernels. Recently, [8] proposed an improved variant, Deep Embedded Validation (dev), that controls the variance of the estimator and estimates the importance weights with a discriminative model that distinguish source samples from target samples leading to a more stable and effective method.
Entropy (ent). While minimizing the entropy of the target samples has been used in domain adaptation to improve accuracy by promoting tighter clusters, [11] showed that it can also be used for model selection. The intuition is that a lower entropy model corresponds to a high confident model with discriminative target features and therefore reliable predictions.
Soft Neighborhood Density (snd). [6] argue that a good UDA model will have a neighborhood structure where nearby target samples are in the same class. They point out that entropy is not able to capture this property and propose the Soft Neighborhood Density (snd) score to address it.
Target Accuracy (oracle). We consider as well the target accuracy on all target samples. While we emphasize once again its use is not realistic in unsupervised domain adaptation (hence why we will refer to it as oracle), it has nonetheless been used to report the best accuracy achieved by the model along training in several previous works [4], [12]–[15]. Here, we use it as an upper bound for all the other model selection strategies and to check the reproducibility of previous works.
Small Labeled Target Set (1-shot and 100-rnd). For real-world applications in an industry setting, it is unlikely that a model will be deployed without the very least of an estimate of its performance for which target labels are required. Therefore, one can imagine a situation where a PDA method is used and a small set of target samples is available. Thus, we will compute the target accuracy with 1 labeled sample per class (1-shot) and 100 random labeled target samples (100-rnd) as model selection strategies. One could argue that the 100 random samples could have been used in the training with semi-supervised domain adaptation methods. However, note that we do not know how many classes we have on the target domain so it is hard to form a split when we have uncertainty of classes. For instance, 100 random represents possibly less than 2 samples per class for one of our real-world dataset, as we do not know the number of classes, making a potential split between a train and validation target sets not possible.
| Method | Architecture | Runs | Model Selection | |
| (bottleneck) | per task | Hyper-Parameters | Along Training | |
| pada | Linear | 1 | IWCV (lacks details) | oracle |
| safn | Non-Linear | 3 | Unknown | oracle |
| ba3us | Linear | 3 | Unknown | oracle |
| ar | Non-Linear | 1 | IWCV (lacks details) | oracle |
| jumbot | Linear | 1 | oracle | final |
| mpot | Linear | 3 | Unknown | oracle |
In this section, we give a brief description of the PDA methods considered in our study. They can be grouped into two families: adversarial training and divergence minimization.
Adversarial training. To solve the UDA problem, [16] aligned the source and target domains with the help of a domain discriminator trained adversarially to be able to distinguish the samples from the two domains. However, when applied to the PDA problem this strategy leads to negative transfer and the model performs worse than a model trained only on source data. [4] proposed pada that introduces a PDA specific solution to adversarial domain adaptation: the contribution of the source-only class samples to the training of both the source classifier and the domain adversarial network is decreased. This is achieved through class weights that are calculated by simply averaging the classifier prediction on all target samples. As the source-only classes should not be predicted in the target domain, they should have lower weights. More recently, [13] proposed ba3us which augments the target mini-batch with source samples to transform the PDA problem in a vanilla DA problem. In addition, an adaptive weighted complement entropy objective is used to encourage incorrect classes to have uniform and low prediction scores.
Divergence minimization. Another standard direction to align the source and target distributions in the feature space of a neural network is to minimize a given divergence between distributions of domains. [12] empirically found than target samples have low feature norm compared to source samples. Based on this insight, they proposed safn which progressively adapts the feature norms of the two domains by minimizing the Maximum Mean Feature Norm Discrepancy. Other approaches are based on optimal transport (OT) [17]. For the PDA problem in specific, [18] developed jumbot, a mini-batch unbalanced optimal transport that learns a joint distribution of the embedded samples and labels. The use of unbalanced OT is critical for the PDA problem as it allows to transport only a portion of the mass limiting the negative transfer between distributions. Based on this work, [15] investigated the partial OT variant [19], a particular case of unbalanced OT, proposing m-pot. Finally, another line of work is to use the Kantorovich-Rubenstein duality of optimal transport to perform the alignment similarly to WGAN[20]. This is precisely the work of [14] that proposed, ar. In addition, source samples are reweighted in order to reduce the negative transfer from the source-only class samples. The Kantorovich-Rubenstein duality relies on a one Lipschitz function which is approximated using adversarial training like the PDA methods described above.
| PDA Methods | pada, safn, ba3us ar, jumbot, mpot |
|---|---|
| Model Selection Strategies | s-acc, ent, dev, snd, 1-shot, 100-rnd, oracle |
| Architecture | ResNet50 backbone \(\oplus\) linear bottleneck \(\oplus\) linear classification head |
| Experimental protocol | 3 seeds on the 12 tasks of office-home and 2 tasks of VisDA |
In this section, we discuss our choices regarding the training details, datasets and neural network architecture. We then discuss the hyper-parameter tuning used in this work. We summarize the PDA methods, model selection strategies and experimental protocol used in this work in Table 3. The main differences in the experimental protocol of the different published state-of-the-art (SOTA) methods is summarized in Table 2. To perform our experiments we developed a PyTorch [21] framework: BenchmarkPDA. We make it available for other researchers to use and contribute with new algorithms and model selection strategies:
It is the standard in the literature when proposing a new method to report directly the results of its competitors from the original papers [4], [12]–[15]. As a result some methods differ for instance in the neural network architecture implementation (ar [14], safn [12]) or evaluation protocol jumbot [18] with other methods. These changes often contribute to an increased performance of the newly proposed method leaving previous methods at a disadvantage. Therefore we chose to implement all methods with the same commonly used neural network architecture, optimizer, learning rate schedule and evaluation protocol. We discuss the details below.
c@|c?@||c?@|c?@|c?@||c?@|c?@|c?@||c?@|c?@|c?@||c?@|c?@|c? & & & & &
& & ent & dev & snd & ent & dev & snd
& ent & dev & snd & ent & dev & snd
& Naive & 52.60 & 63.10 & 44.48 & 52.30 & 26.75 & 17.67 & 49.01 & 16.72 & 30.63 & 32.12 & 49.67 & 5.01
& Heuristic & 58.45 & 63.10 & 60.96 & 56.24 & 45.79 & 55.16 & 49.01 & 45.61 &
30.63 & 46.27 & 49.67 & 49.67
& Naive & 39.06 & 36.99 & 1.14 & 35.89 & 54.53 & 11.99 & 75.04 & 55.33 & 36.11 & 52.82 & 53.26 &
0.83
& Heuristic & 67.50 & 34.94 & 38.76 & 47.23 & 54.53 & 66.42 & 75.04 & 55.33 &
85.36 & 52.82 & 53.26 & 52.82
Methods. We implemented 7 PDA methods by adapting the code from the Official GitHub repositories of each method: Source Only, pada [4], safn [12], ba3us [13], ar [14], jumbot [18], mpot [15]. We provide the links to the different official repositories in Appendix 7.1.
Datasets. We consider two standard real-world datasets used in DA. Our first dataset is office-home [22]. It is a difficult dataset for unsupervised domain adaptation (UDA), it has 15,500 images from four different domains: Art (A), Clipart (C), Product (P) and Real-World (R). For each domain, the dataset contains images of 65 object categories that are common in office and home scenarios. For the partial office-home setting, we follow [4] and select the first 25 categories (in alphabetic order) in each domain as a partial target domain. We evaluate all methods in all 12 adaptation scenarios. visda [23] is a large-scale dataset for UDA. It has 152,397 synthetic images as source domain and 55,388 real-world images as target domain, where 12 object categories are shared by these two domains. For the partial VisDA setting, we follow [4] and select the first 6 categories, taken in alphabetic order, in each domain as a partial target domain. We evaluate the models in the two possible scenarios. We highlight that we are the first to investigate the performance of jumbot and mpot on partial visda.
Model Selection Strategies We consider the 7 different strategies for model selection described in Section 2: s-acc, dev, ent, snd, oracle, 1-shot, 100-rnd. We use them both for hyper-parameter tuning as well selecting the best model along training. Since s-acc, dev and snd require a source validation set, we divide the source samples into a training subset (80%) and validation subset (20%). Regardless of the model selection strategy used, all methods are trained using the source training subset. This is in contrast with previous work that uses all source samples, but necessary to ensure a fair comparison of the model selection strategies. We refer to Appendix 7.2 for additional details.
Architecture. Our network is composed of a feature extractor with a linear classification layer on top of it. The feature extractor is a ResNet50 [1], pre-trained on ImageNet [24], with its last linear layer removed and replaced by a linear bottleneck layer of dimension 256.
Optimizer. We use the SGD [25] algorithm with momentum of 0.9, a weight decay of \(5e^{-4}\) and Nesterov acceleration. As the bottleneck and classifer layers are randomly initialized, we set their learning rates to be 10 times that of the pre-trained ResNet50 backbone. We schedule the learning rate with a strategy similar to the one in [16]: \(\chi_p = \frac{\chi_0}{(1+\mu i)^{-\nu}}\), where \(i\) is the current iteration, \(\chi_0 = 0.001\), \(\gamma = 0.001\), \(\nu = 0.75\). While this schedule is slightly different than the one reported in previous work, it is the one implemented in the different official code implementations. We elaborate in the Appendix 7.3 on the differences and provide additional details. Finally, as for the mini-batch size, jumbot and m-pot were designed with a stratified sampling, i.e., a balanced source mini-batch with the same number of samples per class. This allows to reduce the negative transfer between domains and is crucial to their success. On the other hand, it was shown that for some methods (e.g. BA3US) using a larger mini-batch, than what was reported, leads to a decreased performance [18]. As a result, we used the default mini-batch strategies for each method. jumbot and m-pot use stratefied mini-batches of size 65 for office-home and 36 for visda. All other methods use a standard random uniform sampling strategy with a mini-batch size of 36.
Evaluation Protocol. For the hyper-parameters chosen with each model selection strategy, we run the methods for each task 3 times, each with a different seed (2020, 2021, 2022). We tried to control for the randomness across methods by setting the seeds at the beginning of training. Interestingly, as we discuss in more detail in Section 5, some methods demonstrated a non-negligible variance across the different seeds showing that some hyper-parameters and methods are not robust to randomness.w
| Method | AC | AP | AR | CA | CP | CR | PA | PC | PR | RA | RC | RP | Avg |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| s. only\(^\dagger\) | 46.33 | 67.51 | 75.87 | 59.14 | 59.94 | 62.73 | 58.22 | 41.79 | 74.88 | 67.40 | 48.18 | 74.17 | 61.35 |
| s. only (Ours) | 45.43 | 68.91 | 79.53 | 55.59 | 57.42 | 65.23 | 59.32 | 40.80 | 75.80 | 69.88 | 47.20 | 77.31 | 61.87 |
| pada\(^\dagger\) | 51.95 | 67.00 | 78.74 | 52.16 | 53.78 | 59.03 | 52.61 | 43.22 | 78.79 | 73.73 | 56.60 | 77.09 | 62.06 |
| pada (Ours) | 50.53 | 67.45 | 80.14 | 57.30 | 54.47 | 64.55 | 61.07 | 40.94 | 79.55 | 73.09 | 54.63 | 80.93 | 63.72 |
| safn\(^{\dagger *}\) | 58.93 | 76.25 | 81.42 | 70.43 | 72.97 | 77.78 | 72.36 | 55.34 | 80.40 | 75.81 | 60.42 | 79.92 | 71.84 |
| safn* (Ours) | 59.98 | 79.85 | 85.18 | 72.02 | 73.73 | 78.54 | 76.09 | 59.32 | 83.25 | 80.04 | 64.20 | 84.44 | 74.72 |
| safn (Ours) | 49.57 | 68.55 | 78.26 | 57.91 | 59.29 | 66.81 | 59.87 | 45.29 | 75.98 | 69.08 | 51.68 | 77.29 | 63.30 |
| ba3us\(^\dagger\) | 60.62 | 83.16 | 88.39 | 71.75 | 72.79 | 83.40 | 75.45 | 61.59 | 86.53 | 79.25 | 62.80 | 86.05 | 75.98 |
| ba3us (Ours) | 63.26 | 82.75 | 89.16 | 69.91 | 71.93 | 77.58 | 75.73 | 59.94 | 86.89 | 80.93 | 66.77 | 86.93 | 75.98 |
| ar\(^{\dagger *}\) | 62.13 | 79.22 | 89.12 | 73.92 | 75.57 | 84.37 | 78.42 | 61.91 | 87.85 | 82.19 | 65.37 | 85.27 | 77.11 |
| ar* (Ours) | 62.75 | 81.55 | 89.07 | 71.63 | 73.41 | 82.94 | 75.88 | 61.03 | 85.70 | 79.86 | 62.93 | 85.30 | 76.00 |
| ar (Ours) | 57.33 | 79.61 | 86.31 | 69.45 | 71.88 | 79.94 | 70.28 | 53.57 | 83.78 | 77.26 | 59.68 | 83.72 | 72.73 |
| jumbot\(^\dagger\) | 62.70 | 77.50 | 84.40 | 76.00 | 73.30 | 80.50 | 74.70 | 60.80 | 85.10 | 80.20 | 66.50 | 83.90 | 75.47 |
| jumbot (Ours) | 61.87 | 78.19 | 88.11 | 77.69 | 76.75 | 84.15 | 76.83 | 63.72 | 84.80 | 81.79 | 64.70 | 87.17 | 77.15 |
| mpot\(^\dagger\) | 64.60 | 80.62 | 87.17 | 76.43 | 77.61 | 83.58 | 77.07 | 63.74 | 87.63 | 81.42 | 68.50 | 87.38 | 77.98 |
| mpot (Ours) | 64.48 | 80.88 | 86.78 | 76.22 | 77.95 | 82.59 | 75.18 | 64.60 | 84.87 | 80.59 | 67.04 | 86.52 | 77.31 |
Previous works [5], [26], [27] perform random searches with the same number of runs for each method. In contrast, we perform hyper-parameter grid searches for each method. As a result, the hyper-parameter tuning budgets differs across the methods depending on the number of hyper-parameters and the chosen grid. While one can argue this leads to an unfair comparison of the methods, in practice in most real-world applications one will be interested in using the best method and our approach will capture precisely that.
The hyper-parameter tuning needs to be performed for each task of each dataset, but that would require a significant computational resources without a clear added benefit. Instead for each dataset, we perform the hyper-parameter tuning on a single task: AC for office-home and SR for visda. This same strategy was adopted in [18] and the hyper-parameters were found to generalize to the remaining tasks in the dataset. We conjecture that this may be due to the fact that information regarding the number of target only classes is implicitly hidden in the hyper-parameters. See Appendix 7.4 for more details regarding the hyper-parameters and the grids chosen for each method.
Several runs in our hyper-parameter search for jumbot, m-pot and ba3us were unsuccessful with the optimization reaching its end without the model being trained at all. This poses a challenge to dev, snd and ent and its one of the failures modes accounted for in [6]. Following their recommendations, for jumbot, m-pot and ba3us we discard runs with low source accuracy. Our threshold was 69.01% and 89.83% for the AC task on office-home and TV task on visda, respectively. It corresponds to 90% of the source accuracy attained by the Source-Only model on each task. We consider 90\(\%\) because the ablation study of some methods showed that doing the adaptation decreased slightly the performance on the source domain [17]. See Table [table:model95selection95heuristic] that shows that this heuristic leads to improved results.
Lastly, when choosing the hyper-parameters, we only consider the model at the end of training, discarding the intermediate checkpoint models in order to select hyper-parameters which do not lead to overfitting at the end of training and better generalize to the other tasks. Following the above protocol, for each dataset we trained 468 models in total in order to find the best hyper-parameters. Then, to obtain the results with our neural network architecture on all tasks of each dataset, we trained an additional 1224 models for office-home and 156 models for visda. We additionally trained 231 models with the different neural network architectures for ar and safn. In total, 2547 models were trained to make this study and we present the different results in the next section.
| Dataset :=========================: office-home | Method :====================: s. only | s-acc :===================: 60.38\(\pm\)0.5 | ent :=================: 60.73\(\pm\)0.2 | dev :=================: 60.22\(\pm\)0.3 | snd :=================: 59.55\(\pm\)0.3 | 1-shot :====================: 58.92\(\pm\)0.4 | 100-rnd :=====================: 60.34\(\pm\)0.4 | oracle :====================: 61.87\(\pm\)0.3 |
| pada | 63.08\(\pm\)0.3 | 59.74\(\pm\)0.5 | 52.72\(\pm\)2.8 | 62.36\(\pm\)0.4 | 62.00\(\pm\)0.5 | 63.22\(\pm\)0.1 | 63.72\(\pm\)0.3 | |
| safn | 62.09\(\pm\)0.2 | 61.37\(\pm\)0.3 | 62.03\(\pm\)0.4 | 62.59\(\pm\)0.1 | 49.30\(\pm\)0.7 | 62.36\(\pm\)0.2 | 63.30\(\pm\)0.2 | |
| ba3us | 68.32\(\pm\)1.1 | 73.36\(\pm\)0.6 | 62.25\(\pm\)7.1 | 75.37\(\pm\)0.8 | 65.56\(\pm\)7.6 | 75.19\(\pm\)0.4 | 75.98\(\pm\)0.3 | |
| ar | 65.68\(\pm\)0.3 | 70.58\(\pm\)0.4 | 64.32\(\pm\)0.9 | 70.25\(\pm\)0.2 | 70.56\(\pm\)0.7 | 70.34\(\pm\)0.2 | 72.73\(\pm\)0.3 | |
| jumbot | 62.89\(\pm\)0.2 | 74.61\(\pm\)0.8 | 61.28\(\pm\)0.1 | 72.29\(\pm\)0.2 | 74.95\(\pm\)0.1 | 75.74\(\pm\)0.3 | 77.15\(\pm\)0.4 | |
| mpot | 66.24\(\pm\)0.1 | 64.46\(\pm\)0.1 | 61.37\(\pm\)0.2 | 46.92\(\pm\)0.4 | 68.28\(\pm\)0.2 | 73.06\(\pm\)0.3 | 77.31\(\pm\)0.5 | |
| visda | s. only | 55.15\(\pm\)2.4 | 55.24\(\pm\)3.2 | 55.07\(\pm\)1.2 | 55.02\(\pm\)2.9 | 55.72\(\pm\)2.2 | 58.16\(\pm\)0.6 | 59.48\(\pm\)0.4 |
| pada | 47.48\(\pm\)4.8 | 32.32\(\pm\)4.9 | 43.43\(\pm\)5.3 | 56.83\(\pm\)1.0 | 53.15\(\pm\)2.9 | 54.38\(\pm\)2.7 | 54.57\(\pm\)2.6 | |
| safn | 58.20\(\pm\)1.7 | 42.83\(\pm\)6.3 | 58.62\(\pm\)1.3 | 44.82\(\pm\)8.8 | 56.89\(\pm\)2.1 | 59.09\(\pm\)2.8 | 62.64\(\pm\)1.5 | |
| ba3us | 55.10\(\pm\)3.7 | 65.58\(\pm\)1.4 | 58.40\(\pm\)1.4 | 51.07\(\pm\)4.3 | 64.77\(\pm\)1.4 | 67.44\(\pm\)1.2 | 67.67\(\pm\)1.3 | |
| ar | 66.68\(\pm\)1.0 | 64.27\(\pm\)3.6 | 67.20\(\pm\)1.5 | 55.69\(\pm\)0.9 | 70.29\(\pm\)1.7 | 72.60\(\pm\)0.8 | 73.85\(\pm\)0.9 | |
| jumbot | 60.63\(\pm\)0.7 | 62.42\(\pm\)2.4 | 59.86\(\pm\)0.6 | 77.69\(\pm\)4.2 | 78.34\(\pm\)1.9 | 83.49\(\pm\)1.9 | 84.01\(\pm\)1.9 | |
| mpot | 70.02\(\pm\)2.0 | 74.64\(\pm\)4.4 | 61.62\(\pm\)1.3 | 78.40\(\pm\)3.9 | 70.96\(\pm\)3.7 | 86.69\(\pm\)5.1 | 86.95\(\pm\)5.0 |
We start the results section by discussing the differences between our reproduced results and the published results from the different PDA methods. Then, we compare the performance of the different model selection strategies. Finally, we discuss the sensitivity of methods to the random seed.
We start by ensuring that our reimplementation of PDA methods were done correctly by comparing our reproduced results and the reported results in Table 4. On office-home, both pada and jumbot achieved higher average task accuracy (1.6 and 1.7 percentage points, respectively) in our reimplementation, while for ba3us and mpot we recover the reported accuracy in their respective papers. However, we saw a decrease in performance for both safn and ar of roughly 8 and 5 percentage points respectively. This is to be expected due to the differences in the neural network architectures. While we use a linear bottleneck layer, safn uses a nonlinear bottleneck layer. As for ar, they make two significant changes: the linear classification head is replaced by a spherical logistic regression (SLR) layer [28] and the features are normalized (the 2-norm is set to a dataset dependent value, another hyper-parameter that requires tuning) before feeding them to the classification head. While we account for the first change by comparing to AR (w/ linear) results reported in [14], in our neural network architecture we do not normalize the features. These changes, nonlinear bottleneck layer for safn and feature normalization for ar, significantly boost the performance of both methods. When now comparing our reimplementation with the same neural network architectures, our SAFN reimplementation achieves a higher average task accuracy by 3 percentage points, while our AR reimplementation is now only 1 percentage points below. The fact that AR reported results are from only one run, while ours are averaged across 3 distinct seeds, justifies the small remaining gap. Moreover, we report higher accuracy or on par on 4 tasks of the 12 tasks. Given all the above and further discussion of the visda dataset results in Appendix 8, our reimplementations are trustworthy and give validity to the results we discuss in the next sections.
| Task :==================: SR | Method :====================: s. only | s-acc :=====================: 46.96\(\:\pm\:\)1.5 | ent :=====================: 48.17\(\:\pm\:\)3.9 | dev :=====================: 49.00\(\:\pm\:\)0.9 | snd :=====================: 48.17\(\:\pm\:\)3.9 | 1-shot :=====================: 49.43\(\:\pm\:\)0.8 | 100-rnd :=====================: 50.01\(\:\pm\:\)1.6 | oracle :=====================: 51.86\(\:\pm\:\)1.4 |
| pada | 44.56\(\:\pm\:\)5.9 | 40.83\(\:\pm\:\)11.3 | 41.04\(\:\pm\:\)4.3 | 56.14\(\:\pm\:\)9.7 | 52.94\(\:\pm\:\)4.3 | 49.34\(\:\pm\:\)8.4 | 49.34\(\:\pm\:\)8.4 | |
| safn | 52.04\(\:\pm\:\)3.5 | 29.86\(\:\pm\:\)16.7 | 52.42\(\:\pm\:\)2.9 | 28.46\(\:\pm\:\)16.5 | 49.97\(\:\pm\:\)3.3 | 47.83\(\:\pm\:\)0.6 | 56.88\(\:\pm\:\)2.1 | |
| ba3us | 44.21\(\:\pm\:\)3.0 | 71.17\(\:\pm\:\)1.9 | 48.78\(\:\pm\:\)1.9 | 46.12\(\:\pm\:\)7.8 | 66.79\(\:\pm\:\)1.5 | 71.45\(\:\pm\:\)0.8 | 71.77\(\:\pm\:\)1.1 | |
| ar | 68.39\(\:\pm\:\)1.3 | 75.28\(\:\pm\:\)2.9 | 68.54\(\:\pm\:\)1.3 | 57.61\(\:\pm\:\)0.4 | 70.11\(\:\pm\:\)1.4 | 75.09\(\:\pm\:\)5.2 | 76.33\(\:\pm\:\)4.5 | |
| jumbot | 55.23\(\:\pm\:\)2.3 | 56.25\(\:\pm\:\)2.1 | 54.35\(\:\pm\:\)2.0 | 75.23\(\:\pm\:\)8.4 | 81.27\(\:\pm\:\)6.9 | 89.94\(\:\pm\:\)1.1 | 90.55\(\:\pm\:\)0.5 | |
| mpot | 64.57\(\:\pm\:\)2.9 | 82.10\(\:\pm\:\)2.0 | 57.02\(\:\pm\:\)1.5 | 84.45\(\:\pm\:\)0.4 | 71.33\(\:\pm\:\)4.4 | 87.20\(\:\pm\:\)2.3 | 87.23\(\:\pm\:\)2.3 | |
| RS | s. only | 63.34\(\:\pm\:\)3.4 | 62.32\(\:\pm\:\)2.7 | 61.13\(\:\pm\:\)3.3 | 61.88\(\:\pm\:\)2.3 | 62.00\(\:\pm\:\)3.9 | 66.30\(\:\pm\:\)2.0 | 67.11\(\:\pm\:\)2.1 |
| pada | 50.39\(\:\pm\:\)3.8 | 23.80\(\:\pm\:\)1.6 | 45.82\(\:\pm\:\)9.2 | 57.53\(\:\pm\:\)10.3 | 53.36\(\:\pm\:\)1.7 | 59.43\(\:\pm\:\)5.8 | 59.81\(\:\pm\:\)6.2 | |
| safn | 64.37\(\:\pm\:\)0.7 | 55.80\(\:\pm\:\)5.2 | 64.82\(\:\pm\:\)0.5 | 61.19\(\:\pm\:\)3.3 | 63.82\(\:\pm\:\)1.0 | 70.34\(\:\pm\:\)5.8 | 68.40\(\:\pm\:\)1.2 | |
| ba3us | 65.99\(\:\pm\:\)4.6 | 59.99\(\:\pm\:\)1.3 | 68.01\(\:\pm\:\)1.9 | 56.01\(\:\pm\:\)2.9 | 62.75\(\:\pm\:\)2.6 | 63.44\(\:\pm\:\)1.9 | 63.56\(\:\pm\:\)1.8 | |
| ar | 64.97\(\:\pm\:\)0.8 | 53.26\(\:\pm\:\)9.7 | 65.86\(\:\pm\:\)3.5 | 53.78\(\:\pm\:\)2.1 | 70.46\(\:\pm\:\)4.7 | 70.11\(\:\pm\:\)5.0 | 71.36\(\:\pm\:\)5.5 | |
| jumbot | 66.04\(\:\pm\:\)1.0 | 68.59\(\:\pm\:\)4.6 | 65.36\(\:\pm\:\)0.8 | 80.16\(\:\pm\:\)1.1 | 75.42\(\:\pm\:\)4.8 | 77.03\(\:\pm\:\)2.7 | 77.46\(\:\pm\:\)3.3 | |
| mpot | 75.47\(\:\pm\:\)3.8 | 67.18\(\:\pm\:\)9.1 | 66.21\(\:\pm\:\)1.2 | 72.36\(\:\pm\:\)7.4 | 70.58\(\:\pm\:\)3.1 | 86.18\(\:\pm\:\)8.1 | 86.67\(\:\pm\:\)7.8 | |
| Avg | s. only | 55.15\(\:\pm\:\)2.4 | 55.24\(\:\pm\:\)3.2 | 55.07\(\:\pm\:\)1.2 | 55.02\(\:\pm\:\)2.9 | 55.72\(\:\pm\:\)2.2 | 58.16\(\:\pm\:\)0.6 | 59.48\(\:\pm\:\)0.4 |
| pada | 47.48\(\:\pm\:\)4.8 | 32.32\(\:\pm\:\)4.9 | 43.43\(\:\pm\:\)5.3 | 56.83\(\:\pm\:\)1.0 | 53.15\(\:\pm\:\)2.9 | 54.38\(\:\pm\:\)2.7 | 54.57\(\:\pm\:\)2.6 | |
| safn | 58.20\(\:\pm\:\)1.7 | 42.83\(\:\pm\:\)6.3 | 58.62\(\:\pm\:\)1.3 | 44.82\(\:\pm\:\)8.8 | 56.89\(\:\pm\:\)2.1 | 59.09\(\:\pm\:\)2.8 | 62.64\(\:\pm\:\)1.5 | |
| ba3us | 55.10\(\:\pm\:\)3.7 | 65.58\(\:\pm\:\)1.4 | 58.40\(\:\pm\:\)1.4 | 51.07\(\:\pm\:\)4.3 | 64.77\(\:\pm\:\)1.4 | 67.44\(\:\pm\:\)1.2 | 67.67\(\:\pm\:\)1.3 | |
| ar | 66.68\(\:\pm\:\)1.0 | 64.27\(\:\pm\:\)3.6 | 67.20\(\:\pm\:\)1.5 | 55.69\(\:\pm\:\)0.9 | 70.29\(\:\pm\:\)1.7 | 72.60\(\:\pm\:\)0.8 | 73.85\(\:\pm\:\)0.9 | |
| jumbot | 60.63\(\:\pm\:\)0.7 | 62.42\(\:\pm\:\)2.4 | 59.86\(\:\pm\:\)0.6 | 77.69\(\:\pm\:\)4.2 | 78.34\(\:\pm\:\)1.9 | 83.49\(\:\pm\:\)1.9 | 84.01\(\:\pm\:\)1.9 | |
| mpot | 70.02\(\:\pm\:\)2.0 | 74.64\(\:\pm\:\)4.4 | 61.62\(\:\pm\:\)1.3 | 78.40\(\:\pm\:\)3.9 | 70.96\(\:\pm\:\)3.7 | 86.69\(\:\pm\:\)5.1 | 86.95\(\:\pm\:\)5.0 |
All average accuracies on the office-home and visda datasets can be found in Table 5. For all methods on office-home, we can see that the results for model selections strategies which do not use target labels are below the results given by oracle. For some pairs, the drop of performance can be significant, leading some methods to perform on par with the s. only method. That is the case on office-home when dev is paired with either ba3us, jumbot and mpot. Even worse is mpot with snd as the average accuracy is more than 10 percentage points below that of s. only with any model selection strategy. Overall on office-home, except for mpot, all methods when paired with either ent or snd give results that are at most 2 percentage points below compared to when paired with oracle.
A similar situation can be seen over the visda dataset where the accuracy without target labels can be down to 25 percentage points. Yet again, some model selection strategies can lead to scores even worse than s. only. That is the case for pada, safn and ba3us. Contrary to office-home, all model selection strategies without target labels lead to at least one method with results on par or worse in comparison to the s. only method. Overall, no model selection strategy without target labels can lead to score on par to the oracle model selection strategy. Finally, pada performs worse than s. only for most model selection strategies, including the ones which use target labels. However, when combined with snd it performs better than with oracle on average, although still within the standard deviation. This is a consequence of the random seed dependence mentioned before on visda: as the hyper-parameters were chosen by performing just one run, we were simply “unlucky”. In general, all of this confirms the standard assumption in the literature regarding the difficulty of the visda dataset.
We recall that the oracle model selection strategy uses all the target samples to compute the accuracy while 1-shot and 100-rnd use only subsets: 1-shot has only one sample per class for a total of 25 and 6 on office-home and visda, respetively, while 100-rnd has 100 random target samples. Our results show that using only 100 random target labeled samples is enough to reasonably approximate the target accuracy leading to only a small accuracy drop (one percentage point in almost all cases) for both datasets. Not surprisingly, the gap between the 1-shot and oracle model selection strategies is even bigger, leading in some instances to worse results than with a model selection strategy that uses no target labels. This poor performance of the 1-shot model selection strategy also highlights that semi-supervised domain adaptation (SSDA) methods are not a straightforward alternative to the 100-rnd model selection strategy. While one could argue that the target labels could be leveraged during training like in SSDA methods, one still needs labeled target data to perform model selection. However our results suggest that we would need at least 3 samples per class for SSDA methods. In addition, knowing that we have a certain number of labeled samples per class provides information regarding which classes are target only, one of the main assumptions in PDA. In that case, PDA methods could be tweaked. This warrants further study that we leave as future work. Finally, we have also investigated a smaller labeled target set of 50 random samples (50-rnd) instead of 100 random samples. The accuracies of methods using 50-rnd were not as good as when using 100-rnd. All results of pairs of methods and 50-rnd can be found in Appendix 8. The smaller performance show that the size of the labeled target set is an important element and we suggest to use at least 100 random samples.
Overall, only the jumbot and snd pair performed reasonably well with respect to the jumbot and oracle pair on both datasets. All other pairs failed in either one of the datasets. Our experiments show that there is no model selection strategy which performs well for all methods. That is why to deploy models in a real-world scenario, we advise to test selected models on a small labeled target set (our 100-rnd model selection strategy) to assess the performance of the models as current model selection without target labels can perform very poorly.
Our conclusion is that the model selection for PDA methods is still an open problem. We conjecture that it is also the case for all domain adaptation as the considered metrics were developed first for this setting. For future proposed methods, researchers should specify not only which model selection strategy should be used, but also which hyper-parameter search grid should be considered, in order to deploy them in a real-world scenario.
Ideally, PDA methods should be robust to the choice of random seed. This is of particular importance when performing hyper-parameter tuning since typically only one run per set of hyper-parameters is done (that was the case in our work as well). We investigate this robustness by averaging all the results presented over three different seeds (2020, 2021 and 2022) and reporting the standard deviations. This is in contrast with previous work where only a single run is reported [14], [18]. Other works [4], [12], [13] that report standard deviations do not specify if the random seed is different across runs. Results for all tasks on visda dataset are in Table 6 and on office-home in Appendix 8 due to space constraints.
Our experiments show that some methods express a non-negligible instabilities over randomness with respect to any model selection methods. This is particularly true for ba3us when paired with dev and 1-shot as model selection strategies: there are several tasks where the standard deviation is above 10%. While in this case this instability may stem from the poor performance of the model selection strategies, it is also visible when oracle is the model selection strategy used. For instance, the m-pot has a standard deviation of 3.3% on the AP task of office-home which corresponds to a variance of 11%. On visda this instability and seed dependence is even larger.
In this paper, we investigated how model selection strategies affect the performance of PDA methods. We performed a quantitative study with seven PDA methods and seven model selection strategies on two real-word datasets. Based on our findings, we provide the following recommendations:
i) Target label samples should be used to test models before using them in real-world scenario. While this breaks the main PDA assumption, it is impossible to confidently deploy PDA models selected without the use of target labels. Indeed, model selection strategies without target labels lead to a significant drop in performance in most cases in comparison to using a small validation set. We argue that the cost of labeling it outweighs the uncertainty in current model selection strategies.
ii) The robustness of new PDA method to randomness should be tested over at least three different seeds. We suggest to use the seeds (2020, 2021, 2022) to allow for a fair comparison with our results.
iii) An ablation study should be considered when a novel architecture is proposed to quantify the associated increase of performance.
As our work focus on a quantitative study of model selection methods and reproducibility of state-of-the-art partial domain adaptation methods, we do not see any potential ethical concern. Future work will investigate new model selection strategies which can achieve similar results as model selection strategies which use label target samples.
This work was partially supported by NSERC Discovery grant (RGPIN-2019-06512) and a Samsung grant. Thanks also to CIFAR for their support through the CIFAR AI Chairs program. Authors thank Christos Tsirigotis and Chen Sun for early comments on the manuscript.
Outline. The supplementary material of this paper is organized as follows:
In Section 7, we give more details on our experimental protocol.
In Section 8, we provide additional results from our experiments.
In order to reimplement the different PDA methods, we adapted the code from the official repository associated with each of the paper. We list them in Table 7.
| Method | Code Repository |
|---|---|
| pada | https://github.com/thuml/PADA/blob/master/pytorch/src/ |
| safn | https://github.com/jihanyang/AFN/blob/master/partial/OfficeHome/SAFN/code/ |
| ba3us | https://github.com/tim-learn/BA3US/blob/master/ |
| ar | https://github.com/XJTU-XGU/Adversarial-Reweighting-for-Partial-Domain-Adaptation |
| jumbot | https://github.com/kilianFatras/JUMBOT |
| m-pot | https://github.com/UT-Austin-Data-Science-Group/Mini-batch-OT/tree/master/PartialDA |
One of our main claims regarding previous work is the use of target labels to choose the best model along training. This can be easily verified by inspecting the code. For pada it can be seen on line 240 of the script “train_pada.py”, for ba3us in line 116 for the script “run_partial.py”, for m-pot it can be seen line 164 of the file “run_mOT.py”, for safn it can be seen in the “eval.py” file and finally for ar in line 149 of the script “train.py”.
dev requires learning a discriminative model to distinguish source samples from target samples. Its neural network architecture must be specified as well the training details. [8] (dev) use a multilayer perceptron, while [6] (snd) use a Support Vector Machine in their reimplementation of dev. We empirically observed the latter to yield more stable weights and so that was the one we used. In order to train the SVM discriminator, following [6], we take 3000 feature embeddings from source samples used in training and 3000 random feature embeddings from target samples, both chosen randomly. We do a 80/20 split into training and test data. The SVM is trained with a linear kernel for a maximum of 4000 iterations. Of 5 different SVM models trained with decay values spaced evenly on log space between \(10^{-2}\) and \(10^4\) the one that leads to the highest accuracy (in distinguishing source from target features) on the test data split is the chosen one.
As for snd, it also requires specifying a temperature for temperature scaling component of the strategy. We used the default value of 0.05 that is suggested in [6].
Finally, we mention that the samples used for 100-rnd were randomly selected and their list is made available together with the code. As for the samples used for 1-shot, they are the same as the ones used in semi-supervised domain adaptation.
In general, all methods claim to adopt Nesterov’s acceleration method as the optimization method with a momentum of 0.9 and setting the weight decay set to \(5\times 10^{-4}\). The learning rate follows the annealing strategy as in [16]: \[\mu_p = \mu_0 (1 + \alpha · p)^{-\beta},\] where \(p\) is the training progress linearly changing from 0 to 1, \(\mu_0 = 0.01\) and \(\alpha = 10\) and \(\beta = 0.75\).
However, inspecting the Official code repo for each PDA method, the actual learning schedule is given by \[\mu_i = \mu_0 (1 + \alpha · i)^{-\beta},\] where \(i\) is the iteration number in the training procedure, \(\mu_0 = 0.01\) and \(\alpha = 0.001\) and \(\beta = 0.75\). Only when the total number of iterations is 10000 do the learning rate schedules match. In this work, we followed the latter since it is the one indeed used. For office-home, all methods are trained for 5000 iterations, while for visda they are trained for 10000 iterations, with the exception of the s. only which is trained for 1000 iterations on office-home and 5000 iterations on visda.
In Table 8, we report the values used for each hyper-parameter in our grid search. We report in Table 9 the hyper-parameters chosen by each model selection strategy for each method on both datasets. In addition, for the reproducibility of ar with the proposed architecture in [14],a feature normalization layer is added in the bottleneck which requires specifying \(r\), the value to which the 2-norm is set. This hyper-parameter is therefore included in the hyper-parameter grid search with the possible values of \([5, 10, 20]\) which are the different values used in the experiments in [14].
| Method | HP | Values |
|---|---|---|
| pada | \(\lambda\) | \([0.1, 0.5, 1.0, 5.0, 10.0]\) |
| ba3us | \(\lambda_{wce}\) | \([0.1, 0.5, 1, 5, 10]\) |
| \(\lambda_{ent}\) | \([0.01, 0.05, 0.1, 0.5, 1]\) | |
| safn | \(\lambda\) | \([0.005, 0.01, 0.05, 0.1, 0.5]\) |
| \(\Delta_r\) | \([0.01, 0.1, 1.0]\) | |
| ar | \(\rho_0\) | \([2.5, 5.0, 7.5, 10.0]\) |
| \(A_{up}\) | \([5.0, 10.0]\) | |
| \(A_{low}\) | \(-A_{up}\) | |
| \(\lambda_{ent}\) | \([0.01, 0.1, 1.0]\) | |
| jumbot | \(\tau\) | \([0.001, 0.01, 0.1]\) |
| \(\eta_1\) | \([0.00001, 0.0001, 0.001, 0.01, 0.1]\) | |
| \(\eta_2\) | \([0.1, 0.5, 1.]\) | |
| \(\eta_3\) | \([5, 10, 20]\) | |
| mpot | \(\epsilon\) | \([0.5, 1.0, 1.5]\) |
| \(\eta_1\) | \([0.0001, 0.001, 0.01, 0.1, 1.0]\) | |
| \(\eta_2\) | \([0.1, 1.0, 5.0, 10.0]\) | |
| \(m\) | \([0.1, 0.2, 0.3, 0.4]\) |
| Method | Dataset | HP | oracle | 1-shot | 50-rnd | 100-rnd | s-acc | ent | dev | snd |
|---|---|---|---|---|---|---|---|---|---|---|
| pada | office-home | \(\lambda\) | 0.5 | 0.1 | 0.1 | 0.5 | 0.1 | 1.0 | 5.0 | 0.5 |
| visda | \(\lambda\) | 0.5 | 1.0 | 10.0 | 0.5 | 1.0 | 0.5 | 5.0 | 0.1 | |
| safn | office-home | \(\lambda\) | 0.005 | 0.1 | 0.005 | 0.01 | 0.005 | 0.01 | 0.005 | 0.005 |
| \(\Delta r\) | 0.1 | 0.01 | 0.01 | 0.01 | 0.01 | 0.1 | 0.1 | 0.1 | ||
| visda | \(\lambda\) | 0.005 | 0.005 | 0.05 | 0.05 | 0.005 | 0.05 | 0.005 | 0.05 | |
| \(\Delta r\) | 0.1 | 0.01 | 0.01 | 0.01 | 0.01 | 0.01 | 0.01 | 0.01 | ||
| ba3us | office-home | \(\lambda_{wce}\) | 5.0 | 10.0 | 5.0 | 5.0 | 5.0 | 0.1 | 10.0 | 1.0 |
| \(\lambda_{ent}\) | 0.05 | 0.05 | 0.01 | 0.05 | 0.01 | 0.1 | 0.05 | 0.01 | ||
| visda | \(\lambda_{wce}\) | 1.0 | 1.0 | 0.1 | 1.0 | 5.0 | 1.0 | 5.0 | 5.0 | |
| \(\lambda_{ent}\) | 0.5 | 0.5 | 0.5 | 0.5 | 0.05 | 0.5 | 0.05 | 1.0 | ||
| ar | office-home | \(\rho_0\) | 2.5 | 2.5 | 5.0 | 5.0 | 2.5 | 5.0 | 7.5 | 10.0 |
| \(A_{up}\) | 5.0 | 5.0 | 10.0 | 5.0 | 5.0 | 10.0 | 10.0 | 10.0 | ||
| \(A_{low}\) | -5.0 | -5.0 | -10.0 | -5.0 | -5.0 | -10.0 | -10.0 | -10.0 | ||
| \(\lambda_{ent}\) | 0.1 | 0.1 | 1.0 | 1.0 | 0.01 | 1.0 | 0.01 | 1.0 | ||
| visda | \(\rho_0\) | 2.5 | 2.5 | 2.5 | 2.5 | 2.5 | 7.5 | 2.5 | 10.0 | |
| \(A_{up}\) | 10.0 | 10.0 | 10.0 | 10.0 | 5.0 | 10.0 | 10.0 | 10.0 | ||
| \(A_{low}\) | -10.0 | -10.0 | -10.0 | -10.0 | -5.0 | -10.0 | -10.0 | -10.0 | ||
| \(\lambda_{ent}\) | 0.1 | 0.1 | 0.1 | 0.1 | 0.01 | 0.1 | 0.01 | 0.01 | ||
| jumbot | office-home | \(\tau\) | 0.01 | 0.01 | 0.01 | 0.001 | 0.1 | 0.01 | 0.01 | 0.001 |
| \(\eta_1\) | 0.0001 | 0.0001 | 0.001 | 0.0001 | 0.01 | 1e-05 | 0.01 | 1e-05 | ||
| \(\eta_2\) | 0.5 | 1.0 | 0.5 | 0.1 | 0.1 | 0.5 | 1.0 | 1.0 | ||
| \(\eta_3\) | 10.0 | 5.0 | 5.0 | 5.0 | 5.0 | 20.0 | 10.0 | 5.0 | ||
| visda | \(\tau\) | 0.01 | 0.01 | 0.01 | 0.01 | 0.001 | 0.01 | 0.001 | 0.01 | |
| \(\eta_1\) | 0.001 | 0.001 | 0.001 | 0.001 | 0.01 | 1e-05 | 0.01 | 0.0001 | ||
| \(\eta_2\) | 1.0 | 1.0 | 0.5 | 1.0 | 0.1 | 0.5 | 1.0 | 1.0 | ||
| \(\eta_3\) | 5.0 | 5.0 | 5.0 | 5.0 | 10.0 | 5.0 | 20.0 | 5.0 | ||
| mpot | office-home | \(\epsilon\) | 0.5 | 0.5 | 1.0 | 0.5 | 1.0 | 1.5 | 1.0 | 1.5 |
| \(\eta_1\) | 0.01 | 0.01 | 0.01 | 0.01 | 0.001 | 0.0001 | 1.0 | 0.01 | ||
| \(\eta_2\) | 10.0 | 1.0 | 1.0 | 1.0 | 1.0 | 10.0 | 0.1 | 1.0 | ||
| \(m\) | 0.3 | 0.1 | 0.1 | 0.2 | 0.3 | 0.4 | 0.2 | 0.4 | ||
| visda | \(\epsilon\) | 0.5 | 0.5 | 0.5 | 0.5 | 1.0 | 1.0 | 1.0 | 0.5 | |
| \(\eta_1\) | 0.01 | 0.001 | 0.01 | 0.01 | 0.001 | 0.0001 | 0.0001 | 0.01 | ||
| \(\eta_2\) | 1.0 | 1.0 | 1.0 | 1.0 | 1.0 | 10.0 | 1.0 | 10.0 | ||
| \(m\) | 0.3 | 0.1 | 0.3 | 0.3 | 0.2 | 0.4 | 0.2 | 0.3 |
| Metric :=====================: s-acc | Method :====================: s. only | AC :=====================: 44.50\(\:\pm\:\)1.7 | AP :=====================: 67.71\(\:\pm\:\)2.4 | AR :=====================: 78.37\(\:\pm\:\)0.3 | CA :=====================: 52.56\(\:\pm\:\)0.9 | CP :=====================: 54.81\(\:\pm\:\)0.1 | CR :=====================: 62.88\(\:\pm\:\)0.9 | PA :=====================: 58.77\(\:\pm\:\)0.5 | PC :=====================: 39.28\(\:\pm\:\)0.8 | PR :=====================: 75.08\(\:\pm\:\)0.5 | RA :=====================: 68.90\(\:\pm\:\)0.6 | RC :=====================: 45.33\(\:\pm\:\)1.0 | RP :=====================: 76.34\(\:\pm\:\)0.7 | Avg :=====================: 60.38\(\:\pm\:\)0.5 |
| pada | 50.15\(\:\pm\:\)2.8 | 66.93\(\:\pm\:\)1.2 | 76.73\(\:\pm\:\)1.7 | 58.00\(\:\pm\:\)1.4 | 56.13\(\:\pm\:\)1.4 | 66.45\(\:\pm\:\)0.8 | 60.33\(\:\pm\:\)2.1 | 43.50\(\:\pm\:\)1.2 | 76.70\(\:\pm\:\)0.4 | 69.27\(\:\pm\:\)3.5 | 53.93\(\:\pm\:\)1.3 | 78.88\(\:\pm\:\)0.8 | 63.08\(\:\pm\:\)0.3 | |
| safn | 47.36\(\:\pm\:\)0.1 | 66.82\(\:\pm\:\)1.9 | 77.62\(\:\pm\:\)0.2 | 57.85\(\:\pm\:\)0.6 | 57.89\(\:\pm\:\)0.7 | 66.92\(\:\pm\:\)0.9 | 58.80\(\:\pm\:\)0.7 | 42.49\(\:\pm\:\)0.6 | 75.46\(\:\pm\:\)0.4 | 67.92\(\:\pm\:\)0.0 | 49.73\(\:\pm\:\)0.1 | 76.23\(\:\pm\:\)0.8 | 62.09\(\:\pm\:\)0.2 | |
| ba3us | 54.89\(\:\pm\:\)4.7 | 71.34\(\:\pm\:\)0.8 | 81.91\(\:\pm\:\)3.9 | 61.68\(\:\pm\:\)5.2 | 67.13\(\:\pm\:\)3.9 | 72.96\(\:\pm\:\)1.0 | 68.90\(\:\pm\:\)5.0 | 55.92\(\:\pm\:\)1.3 | 79.13\(\:\pm\:\)4.7 | 72.27\(\:\pm\:\)3.5 | 51.84\(\:\pm\:\)0.5 | 81.85\(\:\pm\:\)4.1 | 68.32\(\:\pm\:\)1.1 | |
| ar | 51.12\(\:\pm\:\)1.2 | 72.79\(\:\pm\:\)0.7 | 77.91\(\:\pm\:\)0.2 | 63.21\(\:\pm\:\)1.5 | 60.54\(\:\pm\:\)4.0 | 72.76\(\:\pm\:\)0.9 | 63.39\(\:\pm\:\)3.1 | 48.36\(\:\pm\:\)1.7 | 78.02\(\:\pm\:\)1.7 | 70.00\(\:\pm\:\)1.1 | 52.52\(\:\pm\:\)1.0 | 77.55\(\:\pm\:\)2.6 | 65.68\(\:\pm\:\)0.3 | |
| jumbot | 49.07\(\:\pm\:\)0.2 | 65.45\(\:\pm\:\)0.4 | 77.14\(\:\pm\:\)0.3 | 60.09\(\:\pm\:\)0.1 | 59.59\(\:\pm\:\)1.3 | 66.67\(\:\pm\:\)1.3 | 60.24\(\:\pm\:\)1.0 | 43.60\(\:\pm\:\)0.0 | 74.43\(\:\pm\:\)0.9 | 70.19\(\:\pm\:\)0.5 | 51.12\(\:\pm\:\)1.1 | 77.12\(\:\pm\:\)1.3 | 62.89\(\:\pm\:\)0.2 | |
| mpot | 53.07\(\:\pm\:\)0.3 | 72.61\(\:\pm\:\)1.2 | 78.50\(\:\pm\:\)0.7 | 61.92\(\:\pm\:\)0.5 | 64.16\(\:\pm\:\)1.8 | 70.22\(\:\pm\:\)0.2 | 64.13\(\:\pm\:\)0.9 | 50.87\(\:\pm\:\)1.1 | 77.40\(\:\pm\:\)0.1 | 70.40\(\:\pm\:\)0.6 | 53.99\(\:\pm\:\)1.5 | 77.61\(\:\pm\:\)0.3 | 66.24\(\:\pm\:\)0.1 | |
| ent | s. only | 45.27\(\:\pm\:\)1.1 | 68.91\(\:\pm\:\)1.4 | 79.26\(\:\pm\:\)0.7 | 54.21\(\:\pm\:\)2.1 | 55.52\(\:\pm\:\)0.6 | 63.19\(\:\pm\:\)0.3 | 56.96\(\:\pm\:\)1.5 | 38.75\(\:\pm\:\)0.6 | 75.65\(\:\pm\:\)1.3 | 69.24\(\:\pm\:\)1.0 | 45.31\(\:\pm\:\)1.0 | 76.47\(\:\pm\:\)0.8 | 60.73\(\:\pm\:\)0.2 |
| pada | 46.03\(\:\pm\:\)2.9 | 62.09\(\:\pm\:\)2.8 | 76.05\(\:\pm\:\)1.4 | 55.07\(\:\pm\:\)2.7 | 47.28\(\:\pm\:\)0.1 | 60.92\(\:\pm\:\)2.4 | 56.69\(\:\pm\:\)2.8 | 38.43\(\:\pm\:\)3.0 | 77.08\(\:\pm\:\)0.2 | 69.48\(\:\pm\:\)1.3 | 49.73\(\:\pm\:\)3.5 | 78.00\(\:\pm\:\)1.7 | 59.74\(\:\pm\:\)0.5 | |
| safn | 47.08\(\:\pm\:\)2.0 | 66.83\(\:\pm\:\)0.5 | 77.73\(\:\pm\:\)0.2 | 56.54\(\:\pm\:\)2.2 | 59.07\(\:\pm\:\)0.7 | 66.22\(\:\pm\:\)0.5 | 56.75\(\:\pm\:\)2.1 | 39.58\(\:\pm\:\)2.0 | 73.90\(\:\pm\:\)0.9 | 67.80\(\:\pm\:\)0.2 | 48.76\(\:\pm\:\)0.1 | 76.23\(\:\pm\:\)0.7 | 61.37\(\:\pm\:\)0.3 | |
| ba3us | 59.26\(\:\pm\:\)0.9 | 76.38\(\:\pm\:\)1.5 | 86.03\(\:\pm\:\)0.6 | 68.96\(\:\pm\:\)1.8 | 71.07\(\:\pm\:\)0.8 | 76.22\(\:\pm\:\)1.2 | 73.16\(\:\pm\:\)0.6 | 57.91\(\:\pm\:\)2.5 | 85.59\(\:\pm\:\)1.2 | 78.11\(\:\pm\:\)1.4 | 62.85\(\:\pm\:\)2.7 | 84.84\(\:\pm\:\)0.6 | 73.36\(\:\pm\:\)0.6 | |
| ar | 54.91\(\:\pm\:\)1.8 | 78.45\(\:\pm\:\)1.8 | 84.23\(\:\pm\:\)0.9 | 64.86\(\:\pm\:\)2.3 | 68.16\(\:\pm\:\)3.5 | 80.45\(\:\pm\:\)0.8 | 67.58\(\:\pm\:\)0.4 | 52.34\(\:\pm\:\)1.0 | 82.48\(\:\pm\:\)1.9 | 74.75\(\:\pm\:\)2.1 | 55.64\(\:\pm\:\)1.2 | 83.06\(\:\pm\:\)1.2 | 70.58\(\:\pm\:\)0.4 | |
| jumbot | 57.69\(\:\pm\:\)5.6 | 75.44\(\:\pm\:\)1.4 | 85.24\(\:\pm\:\)2.7 | 75.97\(\:\pm\:\)1.4 | 74.85\(\:\pm\:\)3.3 | 79.75\(\:\pm\:\)1.2 | 72.85\(\:\pm\:\)2.4 | 60.18\(\:\pm\:\)0.9 | 83.21\(\:\pm\:\)1.1 | 81.97\(\:\pm\:\)1.0 | 61.81\(\:\pm\:\)4.6 | 86.33\(\:\pm\:\)1.6 | 74.61\(\:\pm\:\)0.8 | |
| mpot | 52.94\(\:\pm\:\)2.0 | 68.94\(\:\pm\:\)1.2 | 75.98\(\:\pm\:\)0.6 | 60.58\(\:\pm\:\)0.8 | 65.99\(\:\pm\:\)2.2 | 71.51\(\:\pm\:\)0.8 | 58.28\(\:\pm\:\)0.9 | 49.87\(\:\pm\:\)2.6 | 73.77\(\:\pm\:\)1.3 | 64.98\(\:\pm\:\)0.4 | 57.53\(\:\pm\:\)0.6 | 73.17\(\:\pm\:\)2.7 | 64.46\(\:\pm\:\)0.1 | |
| dev | s. only | 43.74\(\:\pm\:\)1.8 | 67.81\(\:\pm\:\)1.2 | 78.28\(\:\pm\:\)0.7 | 51.42\(\:\pm\:\)2.7 | 54.55\(\:\pm\:\)1.2 | 63.94\(\:\pm\:\)1.7 | 57.94\(\:\pm\:\)0.9 | 39.40\(\:\pm\:\)0.9 | 74.91\(\:\pm\:\)0.6 | 69.27\(\:\pm\:\)1.0 | 45.33\(\:\pm\:\)1.0 | 75.99\(\:\pm\:\)1.3 | 60.22\(\:\pm\:\)0.3 |
| pada | 44.70\(\:\pm\:\)1.3 | 61.61\(\:\pm\:\)5.4 | 68.99\(\:\pm\:\)11.3 | 35.08\(\:\pm\:\)13.1 | 24.24\(\:\pm\:\)20.5 | 61.66\(\:\pm\:\)2.4 | 57.91\(\:\pm\:\)1.7 | 38.03\(\:\pm\:\)0.6 | 73.11\(\:\pm\:\)3.4 | 66.33\(\:\pm\:\)0.7 | 29.97\(\:\pm\:\)21.0 | 71.07\(\:\pm\:\)11.3 | 52.72\(\:\pm\:\)2.8 | |
| safn | 48.12\(\:\pm\:\)0.4 | 67.30\(\:\pm\:\)0.5 | 77.43\(\:\pm\:\)0.5 | 56.75\(\:\pm\:\)0.3 | 58.17\(\:\pm\:\)1.2 | 65.64\(\:\pm\:\)1.3 | 59.08\(\:\pm\:\)0.5 | 43.00\(\:\pm\:\)1.1 | 74.64\(\:\pm\:\)0.4 | 68.11\(\:\pm\:\)0.9 | 50.53\(\:\pm\:\)0.6 | 75.65\(\:\pm\:\)0.5 | 62.03\(\:\pm\:\)0.4 | |
| ba3us | 41.67\(\:\pm\:\)18.9 | 50.05\(\:\pm\:\)28.7 | 63.74\(\:\pm\:\)26.1 | 60.70\(\:\pm\:\)2.2 | 59.08\(\:\pm\:\)10.9 | 67.88\(\:\pm\:\)0.9 | 64.62\(\:\pm\:\)1.6 | 56.74\(\:\pm\:\)1.3 | 75.21\(\:\pm\:\)0.6 | 70.92\(\:\pm\:\)2.0 | 58.39\(\:\pm\:\)2.3 | 78.06\(\:\pm\:\)1.3 | 62.25\(\:\pm\:\)7.1 | |
| ar | 49.25\(\:\pm\:\)2.8 | 70.20\(\:\pm\:\)1.7 | 79.73\(\:\pm\:\)2.5 | 62.72\(\:\pm\:\)1.0 | 61.85\(\:\pm\:\)4.6 | 70.86\(\:\pm\:\)5.6 | 61.65\(\:\pm\:\)1.0 | 43.72\(\:\pm\:\)0.7 | 76.29\(\:\pm\:\)0.7 | 70.31\(\:\pm\:\)1.7 | 49.61\(\:\pm\:\)0.8 | 75.61\(\:\pm\:\)0.4 | 64.32\(\:\pm\:\)0.9 | |
| jumbot | 46.11\(\:\pm\:\)0.1 | 66.33\(\:\pm\:\)0.6 | 76.42\(\:\pm\:\)0.3 | 56.81\(\:\pm\:\)0.1 | 56.36\(\:\pm\:\)0.5 | 66.70\(\:\pm\:\)0.8 | 58.03\(\:\pm\:\)1.1 | 41.99\(\:\pm\:\)0.8 | 74.97\(\:\pm\:\)0.5 | 67.43\(\:\pm\:\)0.3 | 48.12\(\:\pm\:\)0.5 | 76.04\(\:\pm\:\)0.1 | 61.28\(\:\pm\:\)0.1 | |
| mpot | 46.07\(\:\pm\:\)0.7 | 65.43\(\:\pm\:\)0.8 | 76.46\(\:\pm\:\)0.4 | 56.44\(\:\pm\:\)1.0 | 57.95\(\:\pm\:\)1.0 | 66.35\(\:\pm\:\)1.0 | 57.64\(\:\pm\:\)0.8 | 43.60\(\:\pm\:\)0.6 | 74.86\(\:\pm\:\)1.3 | 67.68\(\:\pm\:\)0.5 | 48.12\(\:\pm\:\)0.8 | 75.89\(\:\pm\:\)0.4 | 61.37\(\:\pm\:\)0.2 | |
| snd | s. only | 42.23\(\:\pm\:\)1.3 | 68.91\(\:\pm\:\)1.4 | 79.35\(\:\pm\:\)0.6 | 51.76\(\:\pm\:\)3.7 | 53.48\(\:\pm\:\)2.1 | 63.94\(\:\pm\:\)1.7 | 55.37\(\:\pm\:\)0.6 | 37.35\(\:\pm\:\)1.0 | 74.10\(\:\pm\:\)2.8 | 68.53\(\:\pm\:\)1.4 | 43.78\(\:\pm\:\)0.6 | 75.84\(\:\pm\:\)1.6 | 59.55\(\:\pm\:\)0.3 |
| pada | 50.43\(\:\pm\:\)0.8 | 66.72\(\:\pm\:\)1.5 | 79.72\(\:\pm\:\)1.8 | 57.30\(\:\pm\:\)1.9 | 52.10\(\:\pm\:\)1.7 | 63.11\(\:\pm\:\)1.9 | 60.82\(\:\pm\:\)3.0 | 39.26\(\:\pm\:\)2.0 | 79.33\(\:\pm\:\)1.3 | 73.09\(\:\pm\:\)1.5 | 45.77\(\:\pm\:\)1.6 | 80.62\(\:\pm\:\)0.4 | 62.36\(\:\pm\:\)0.4 | |
| safn | 49.57\(\:\pm\:\)0.3 | 68.18\(\:\pm\:\)1.3 | 77.86\(\:\pm\:\)0.5 | 57.91\(\:\pm\:\)0.3 | 58.17\(\:\pm\:\)1.2 | 66.13\(\:\pm\:\)1.0 | 59.14\(\:\pm\:\)0.8 | 43.90\(\:\pm\:\)0.5 | 75.81\(\:\pm\:\)0.7 | 68.17\(\:\pm\:\)1.6 | 49.59\(\:\pm\:\)1.6 | 76.64\(\:\pm\:\)0.5 | 62.59\(\:\pm\:\)0.1 | |
| ba3us | 62.21\(\:\pm\:\)0.9 | 83.29\(\:\pm\:\)0.4 | 88.50\(\:\pm\:\)0.6 | 68.50\(\:\pm\:\)0.9 | 71.45\(\:\pm\:\)3.6 | 76.96\(\:\pm\:\)0.6 | 76.19\(\:\pm\:\)1.2 | 59.94\(\:\pm\:\)1.7 | 86.31\(\:\pm\:\)1.4 | 79.46\(\:\pm\:\)1.4 | 65.35\(\:\pm\:\)1.9 | 86.35\(\:\pm\:\)0.9 | 75.37\(\:\pm\:\)0.8 | |
| ar | 54.37\(\:\pm\:\)1.6 | 79.01\(\:\pm\:\)2.2 | 84.54\(\:\pm\:\)0.8 | 64.52\(\:\pm\:\)1.6 | 68.05\(\:\pm\:\)3.2 | 79.16\(\:\pm\:\)2.8 | 65.60\(\:\pm\:\)1.7 | 51.28\(\:\pm\:\)1.6 | 83.05\(\:\pm\:\)1.1 | 75.02\(\:\pm\:\)1.6 | 55.02\(\:\pm\:\)1.8 | 83.40\(\:\pm\:\)0.9 | 70.25\(\:\pm\:\)0.2 | |
| jumbot | 56.60\(\:\pm\:\)2.8 | 68.48\(\:\pm\:\)1.5 | 84.70\(\:\pm\:\)2.1 | 71.81\(\:\pm\:\)1.8 | 71.84\(\:\pm\:\)1.6 | 80.91\(\:\pm\:\)0.9 | 70.28\(\:\pm\:\)0.8 | 50.69\(\:\pm\:\)4.9 | 83.89\(\:\pm\:\)1.5 | 81.21\(\:\pm\:\)0.6 | 58.85\(\:\pm\:\)1.7 | 88.18\(\:\pm\:\)0.4 | 72.29\(\:\pm\:\)0.2 | |
| mpot | 32.96\(\:\pm\:\)0.4 | 49.73\(\:\pm\:\)1.1 | 57.39\(\:\pm\:\)1.4 | 44.11\(\:\pm\:\)2.4 | 38.66\(\:\pm\:\)1.2 | 50.06\(\:\pm\:\)1.0 | 43.74\(\:\pm\:\)4.3 | 28.66\(\:\pm\:\)2.6 | 58.40\(\:\pm\:\)1.9 | 56.90\(\:\pm\:\)1.9 | 39.34\(\:\pm\:\)1.2 | 63.14\(\:\pm\:\)0.6 | 46.92\(\:\pm\:\)0.4 | |
| 1-shot | s. only | 43.84\(\:\pm\:\)1.7 | 66.52\(\:\pm\:\)3.1 | 77.38\(\:\pm\:\)0.9 | 50.47\(\:\pm\:\)2.4 | 53.24\(\:\pm\:\)2.0 | 61.77\(\:\pm\:\)1.1 | 56.11\(\:\pm\:\)1.7 | 37.35\(\:\pm\:\)1.0 | 71.97\(\:\pm\:\)1.8 | 68.96\(\:\pm\:\)0.5 | 46.13\(\:\pm\:\)2.0 | 73.33\(\:\pm\:\)2.2 | 58.92\(\:\pm\:\)0.4 |
| pada | 52.98\(\:\pm\:\)0.2 | 63.03\(\:\pm\:\)1.6 | 78.06\(\:\pm\:\)2.6 | 51.67\(\:\pm\:\)5.0 | 56.28\(\:\pm\:\)0.4 | 64.00\(\:\pm\:\)1.4 | 58.92\(\:\pm\:\)3.3 | 43.62\(\:\pm\:\)1.0 | 74.27\(\:\pm\:\)4.1 | 68.26\(\:\pm\:\)3.1 | 54.25\(\:\pm\:\)1.6 | 78.62\(\:\pm\:\)0.4 | 62.00\(\:\pm\:\)0.5 | |
| safn | 31.40\(\:\pm\:\)3.7 | 49.73\(\:\pm\:\)4.3 | 62.82\(\:\pm\:\)2.0 | 48.88\(\:\pm\:\)2.4 | 45.27\(\:\pm\:\)0.7 | 57.26\(\:\pm\:\)2.2 | 42.33\(\:\pm\:\)1.6 | 29.77\(\:\pm\:\)2.6 | 63.52\(\:\pm\:\)3.2 | 56.11\(\:\pm\:\)3.2 | 37.55\(\:\pm\:\)0.8 | 67.00\(\:\pm\:\)1.4 | 49.30\(\:\pm\:\)0.7 | |
| ba3us | 44.60\(\:\pm\:\)21.0 | 51.39\(\:\pm\:\)29.8 | 65.47\(\:\pm\:\)27.2 | 65.63\(\:\pm\:\)1.4 | 59.78\(\:\pm\:\)15.3 | 68.49\(\:\pm\:\)1.3 | 68.38\(\:\pm\:\)1.7 | 57.83\(\:\pm\:\)1.3 | 82.05\(\:\pm\:\)1.0 | 80.78\(\:\pm\:\)1.1 | 63.10\(\:\pm\:\)0.8 | 79.20\(\:\pm\:\)1.1 | 65.56\(\:\pm\:\)7.6 | |
| ar | 56.00\(\:\pm\:\)2.3 | 78.58\(\:\pm\:\)1.9 | 82.77\(\:\pm\:\)2.0 | 68.99\(\:\pm\:\)0.2 | 68.35\(\:\pm\:\)1.9 | 77.25\(\:\pm\:\)1.4 | 69.67\(\:\pm\:\)1.5 | 51.98\(\:\pm\:\)1.8 | 78.72\(\:\pm\:\)1.0 | 76.19\(\:\pm\:\)0.7 | 55.48\(\:\pm\:\)2.1 | 82.73\(\:\pm\:\)1.0 | 70.56\(\:\pm\:\)0.7 | |
| jumbot | 61.59\(\:\pm\:\)1.7 | 76.86\(\:\pm\:\)3.4 | 86.45\(\:\pm\:\)2.1 | 74.20\(\:\pm\:\)0.9 | 73.43\(\:\pm\:\)3.3 | 79.85\(\:\pm\:\)0.3 | 74.96\(\:\pm\:\)3.4 | 62.87\(\:\pm\:\)0.6 | 81.83\(\:\pm\:\)0.9 | 78.48\(\:\pm\:\)2.0 | 61.59\(\:\pm\:\)2.2 | 87.34\(\:\pm\:\)0.2 | 74.95\(\:\pm\:\)0.1 | |
| mpot | 53.97\(\:\pm\:\)1.3 | 68.78\(\:\pm\:\)1.7 | 78.04\(\:\pm\:\)2.1 | 69.24\(\:\pm\:\)0.4 | 65.88\(\:\pm\:\)0.5 | 71.42\(\:\pm\:\)0.7 | 70.31\(\:\pm\:\)1.0 | 53.03\(\:\pm\:\)0.7 | 76.88\(\:\pm\:\)1.3 | 76.52\(\:\pm\:\)0.4 | 57.39\(\:\pm\:\)1.7 | 77.95\(\:\pm\:\)1.4 | 68.28\(\:\pm\:\)0.2 | |
| 100-rnd | s. only | 43.28\(\:\pm\:\)1.6 | 68.76\(\:\pm\:\)1.6 | 77.97\(\:\pm\:\)1.2 | 53.75\(\:\pm\:\)1.1 | 55.57\(\:\pm\:\)2.2 | 63.94\(\:\pm\:\)0.4 | 58.37\(\:\pm\:\)0.4 | 39.12\(\:\pm\:\)0.4 | 75.56\(\:\pm\:\)1.3 | 69.02\(\:\pm\:\)0.5 | 43.46\(\:\pm\:\)0.2 | 75.28\(\:\pm\:\)2.3 | 60.34\(\:\pm\:\)0.4 |
| pada | 50.41\(\:\pm\:\)0.8 | 67.21\(\:\pm\:\)1.8 | 79.97\(\:\pm\:\)1.5 | 56.69\(\:\pm\:\)1.5 | 53.86\(\:\pm\:\)1.6 | 63.94\(\:\pm\:\)1.3 | 60.27\(\:\pm\:\)2.7 | 40.56\(\:\pm\:\)1.8 | 78.91\(\:\pm\:\)1.8 | 72.70\(\:\pm\:\)1.4 | 53.39\(\:\pm\:\)2.2 | 80.73\(\:\pm\:\)0.9 | 63.22\(\:\pm\:\)0.1 | |
| safn | 47.58\(\:\pm\:\)0.8 | 67.53\(\:\pm\:\)0.8 | 77.91\(\:\pm\:\)0.4 | 56.47\(\:\pm\:\)1.0 | 58.19\(\:\pm\:\)0.4 | 65.88\(\:\pm\:\)0.2 | 59.69\(\:\pm\:\)0.1 | 43.14\(\:\pm\:\)1.7 | 75.00\(\:\pm\:\)0.7 | 69.64\(\:\pm\:\)1.0 | 50.85\(\:\pm\:\)0.3 | 76.41\(\:\pm\:\)0.8 | 62.36\(\:\pm\:\)0.2 | |
| ba3us | 62.53\(\:\pm\:\)2.0 | 82.09\(\:\pm\:\)0.8 | 88.28\(\:\pm\:\)0.4 | 69.15\(\:\pm\:\)1.2 | 71.65\(\:\pm\:\)1.5 | 77.21\(\:\pm\:\)0.6 | 75.15\(\:\pm\:\)1.3 | 58.17\(\:\pm\:\)1.0 | 85.92\(\:\pm\:\)1.3 | 79.86\(\:\pm\:\)2.1 | 66.57\(\:\pm\:\)1.5 | 85.66\(\:\pm\:\)1.0 | 75.19\(\:\pm\:\)0.4 | |
| ar | 54.89\(\:\pm\:\)2.0 | 78.54\(\:\pm\:\)1.4 | 84.34\(\:\pm\:\)0.6 | 64.95\(\:\pm\:\)2.4 | 69.00\(\:\pm\:\)3.7 | 79.57\(\:\pm\:\)0.2 | 66.73\(\:\pm\:\)0.3 | 50.85\(\:\pm\:\)1.4 | 82.39\(\:\pm\:\)1.9 | 74.66\(\:\pm\:\)2.3 | 55.42\(\:\pm\:\)1.6 | 82.80\(\:\pm\:\)0.4 | 70.34\(\:\pm\:\)0.2 | |
| jumbot | 61.07\(\:\pm\:\)0.9 | 77.87\(\:\pm\:\)1.4 | 86.01\(\:\pm\:\)1.3 | 74.56\(\:\pm\:\)0.4 | 76.40\(\:\pm\:\)1.4 | 81.54\(\:\pm\:\)1.7 | 72.60\(\:\pm\:\)1.2 | 59.92\(\:\pm\:\)0.4 | 84.63\(\:\pm\:\)2.3 | 81.85\(\:\pm\:\)1.7 | 64.84\(\:\pm\:\)1.0 | 87.64\(\:\pm\:\)0.7 | 75.74\(\:\pm\:\)0.3 | |
| mpot | 61.59\(\:\pm\:\)1.2 | 75.56\(\:\pm\:\)1.7 | 82.59\(\:\pm\:\)0.6 | 72.48\(\:\pm\:\)1.0 | 69.77\(\:\pm\:\)0.9 | 75.41\(\:\pm\:\)0.5 | 72.64\(\:\pm\:\)0.9 | 57.67\(\:\pm\:\)1.6 | 82.02\(\:\pm\:\)0.6 | 79.80\(\:\pm\:\)0.5 | 64.64\(\:\pm\:\)0.1 | 82.60\(\:\pm\:\)0.5 | 73.06\(\:\pm\:\)0.3 | |
| oracle | s. only | 45.43\(\:\pm\:\)0.9 | 68.91\(\:\pm\:\)1.4 | 79.53\(\:\pm\:\)0.3 | 55.59\(\:\pm\:\)0.7 | 57.42\(\:\pm\:\)1.2 | 65.23\(\:\pm\:\)0.8 | 59.32\(\:\pm\:\)0.7 | 40.80\(\:\pm\:\)0.9 | 75.80\(\:\pm\:\)1.2 | 69.88\(\:\pm\:\)0.9 | 47.20\(\:\pm\:\)0.9 | 77.31\(\:\pm\:\)0.1 | 61.87\(\:\pm\:\)0.3 |
| pada | 50.53\(\:\pm\:\)0.7 | 67.45\(\:\pm\:\)1.6 | 80.14\(\:\pm\:\)1.4 | 57.30\(\:\pm\:\)1.9 | 54.47\(\:\pm\:\)1.7 | 64.55\(\:\pm\:\)1.1 | 61.07\(\:\pm\:\)3.0 | 40.94\(\:\pm\:\)1.6 | 79.55\(\:\pm\:\)1.4 | 73.09\(\:\pm\:\)1.5 | 54.63\(\:\pm\:\)0.9 | 80.93\(\:\pm\:\)0.6 | 63.72\(\:\pm\:\)0.3 | |
| safn | 49.57\(\:\pm\:\)0.3 | 68.55\(\:\pm\:\)1.0 | 78.26\(\:\pm\:\)0.2 | 57.91\(\:\pm\:\)0.3 | 59.29\(\:\pm\:\)0.5 | 66.81\(\:\pm\:\)0.5 | 59.87\(\:\pm\:\)0.7 | 45.29\(\:\pm\:\)0.7 | 75.98\(\:\pm\:\)0.6 | 69.08\(\:\pm\:\)0.6 | 51.68\(\:\pm\:\)0.8 | 77.29\(\:\pm\:\)0.5 | 63.30\(\:\pm\:\)0.2 | |
| ba3us | 63.26\(\:\pm\:\)1.0 | 82.75\(\:\pm\:\)0.9 | 89.16\(\:\pm\:\)0.2 | 69.91\(\:\pm\:\)0.2 | 71.93\(\:\pm\:\)1.6 | 77.58\(\:\pm\:\)0.9 | 75.73\(\:\pm\:\)1.3 | 59.94\(\:\pm\:\)0.7 | 86.89\(\:\pm\:\)0.5 | 80.93\(\:\pm\:\)0.8 | 66.77\(\:\pm\:\)1.5 | 86.93\(\:\pm\:\)0.2 | 75.98\(\:\pm\:\)0.3 | |
| ar | 57.33\(\:\pm\:\)1.7 | 79.61\(\:\pm\:\)1.6 | 86.31\(\:\pm\:\)0.4 | 69.45\(\:\pm\:\)0.5 | 71.88\(\:\pm\:\)0.9 | 79.94\(\:\pm\:\)0.8 | 70.28\(\:\pm\:\)1.0 | 53.57\(\:\pm\:\)0.2 | 83.78\(\:\pm\:\)1.0 | 77.26\(\:\pm\:\)0.6 | 59.68\(\:\pm\:\)1.1 | 83.72\(\:\pm\:\)0.6 | 72.73\(\:\pm\:\)0.3 | |
| jumbot | 61.87\(\:\pm\:\)1.4 | 78.19\(\:\pm\:\)2.4 | 88.11\(\:\pm\:\)1.5 | 77.69\(\:\pm\:\)0.1 | 76.75\(\:\pm\:\)0.8 | 84.15\(\:\pm\:\)1.3 | 76.83\(\:\pm\:\)1.9 | 63.72\(\:\pm\:\)0.5 | 84.80\(\:\pm\:\)1.3 | 81.79\(\:\pm\:\)0.8 | 64.70\(\:\pm\:\)1.1 | 87.17\(\:\pm\:\)1.7 | 77.15\(\:\pm\:\)0.4 | |
| mpot | 64.48\(\:\pm\:\)1.2 | 80.88\(\:\pm\:\)3.3 | 86.78\(\:\pm\:\)0.5 | 76.22\(\:\pm\:\)0.1 | 77.95\(\:\pm\:\)1.3 | 82.59\(\:\pm\:\)0.7 | 75.18\(\:\pm\:\)1.3 | 64.60\(\:\pm\:\)0.0 | 84.87\(\:\pm\:\)1.4 | 80.59\(\:\pm\:\)0.6 | 67.04\(\:\pm\:\)0.6 | 86.52\(\:\pm\:\)1.2 | 77.31\(\:\pm\:\)0.5 |
In this section, we provide additional results that we could not add to the main paper due to the space constraints.
In Table 10, we show the accuracy per task on office-home averaged over three different seeds (2020, 2021, 2022) for all pairs of methods and model selection strategies. In Table 11, we compare previously reported results with ours on visda. While proposed methods reported results on office-home, only pada and ar results are reported in the original papers for visda. [14] ar) also report results for ba3us. Analysing the results, we see a 9 percentage point decrease in average task accuracy for pada, but our experiments show that there is a significant seed dependence which we discuss in detail below. This is particularly important since [4] (pada) report results from a single run. Comparing our best seeds for pada on the SR and RS tasks, we achieve 58.01% and 67.9% accuracy versus a reported 53.53% and 76.5%. Moreover, we point out that the official code repository for pada does not include the details to reproduce the visda experiments, so it is possible that minor tweaks (e.g learning rate) are necessary. As for ba3us, our results are within the standard deviation being better on the SR task and worse on the RS task. Finally as for ar we see a decrease in performance which, as the results on office-home show, can be explained by the differences in the neural network architecture.
Finally in Table 12, we show all the average task accuracies from all pairs of methods and model selection strategies on the office-home and visda datasets including the 50-rnd model selection strategy.
| Algorithm | SR | RS | Avg |
|---|---|---|---|
| s. only\(^\dagger\) | 45.26 | 64.28 | 54.77 |
| s. only (Ours) | 51.86 | 67.11 | 59.48 |
| pada\(^\dagger\) | 53.53 | 76.50 | 65.02 |
| pada (Ours) | 49.34 | 59.81 | 54.57 |
| safn\(^\dagger\) | 67.65 | - | - |
| safn (Ours) | 56.88 | 68.40 | 62.64 |
| ba3us\(^\dagger\) | 69.86 | 67.56 | 68.71 |
| ba3us (Ours) | 71.77 | 63.56 | 67.67 |
| ar\(^{\dagger *}\) | 85.30 | 74.82 | 80.06 |
| ar (Ours) | 76.33 | 71.36 | 73.85 |
| jumbot\(^\dagger\) | - | - | - |
| jumbot (Ours) | 90.55 | 77.46 | 84.01 |
| mpot\(^\dagger\) | - | - | - |
| mpot (Ours) | 87.23 | 86.67 | 86.95 |
| Dataset :=========================: office-home | Method :====================: s. only | s-acc :===================: 60.38\(\pm\)0.5 | ent :=================: 60.73\(\pm\)0.2 | dev :=================: 60.22\(\pm\)0.3 | snd :=================: 59.55\(\pm\)0.3 | 1-shot :====================: 58.92\(\pm\)0.4 | 50-rnd :====================: 60.28\(\pm\)0.4 | 100-rnd :=====================: 60.34\(\pm\)0.4 | oracle :====================: 61.87\(\pm\)0.3 |
| pada | 63.08\(\pm\)0.3 | 59.74\(\pm\)0.5 | 52.72\(\pm\)2.8 | 62.36\(\pm\)0.4 | 62.00\(\pm\)0.5 | 63.82\(\pm\)0.4 | 63.22\(\pm\)0.1 | 63.72\(\pm\)0.3 | |
| safn | 62.09\(\pm\)0.2 | 61.37\(\pm\)0.3 | 62.03\(\pm\)0.4 | 62.59\(\pm\)0.1 | 49.30\(\pm\)0.7 | 62.00\(\pm\)0.2 | 62.36\(\pm\)0.2 | 63.30\(\pm\)0.2 | |
| ba3us | 68.32\(\pm\)1.1 | 73.36\(\pm\)0.6 | 62.25\(\pm\)7.1 | 75.37\(\pm\)0.8 | 65.56\(\pm\)7.6 | 73.22\(\pm\)0.3 | 75.19\(\pm\)0.4 | 75.98\(\pm\)0.3 | |
| ar | 65.68\(\pm\)0.3 | 70.58\(\pm\)0.4 | 64.32\(\pm\)0.9 | 70.25\(\pm\)0.2 | 70.56\(\pm\)0.7 | 70.26\(\pm\)0.2 | 70.34\(\pm\)0.2 | 72.73\(\pm\)0.3 | |
| jumbot | 62.89\(\pm\)0.2 | 74.61\(\pm\)0.8 | 61.28\(\pm\)0.1 | 72.29\(\pm\)0.2 | 74.95\(\pm\)0.1 | 64.95\(\pm\)0.3 | 75.74\(\pm\)0.3 | 77.15\(\pm\)0.4 | |
| mpot | 66.24\(\pm\)0.1 | 64.46\(\pm\)0.1 | 61.37\(\pm\)0.2 | 46.92\(\pm\)0.4 | 68.28\(\pm\)0.2 | 69.90\(\pm\)0.5 | 73.06\(\pm\)0.3 | 77.31\(\pm\)0.5 | |
| visda | s. only | 55.15\(\pm\)2.4 | 55.24\(\pm\)3.2 | 55.07\(\pm\)1.2 | 55.02\(\pm\)2.9 | 55.72\(\pm\)2.2 | 57.90\(\pm\)1.1 | 58.16\(\pm\)0.6 | 59.48\(\pm\)0.4 |
| pada | 47.48\(\pm\)4.8 | 32.32\(\pm\)4.9 | 43.43\(\pm\)5.3 | 56.83\(\pm\)1.0 | 53.15\(\pm\)2.9 | 55.67\(\pm\)2.5 | 54.38\(\pm\)2.7 | 54.57\(\pm\)2.6 | |
| safn | 58.20\(\pm\)1.7 | 42.83\(\pm\)6.3 | 58.62\(\pm\)1.3 | 44.82\(\pm\)8.8 | 56.89\(\pm\)2.1 | 57.90\(\pm\)3.3 | 59.09\(\pm\)2.8 | 62.64\(\pm\)1.5 | |
| ba3us | 55.10\(\pm\)3.7 | 65.58\(\pm\)1.4 | 58.40\(\pm\)1.4 | 51.07\(\pm\)4.3 | 64.77\(\pm\)1.4 | 66.66\(\pm\)2.4 | 67.44\(\pm\)1.2 | 67.67\(\pm\)1.3 | |
| ar | 66.68\(\pm\)1.0 | 64.27\(\pm\)3.6 | 67.20\(\pm\)1.5 | 55.69\(\pm\)0.9 | 70.29\(\pm\)1.7 | 71.91\(\pm\)0.3 | 72.60\(\pm\)0.8 | 73.85\(\pm\)0.9 | |
| jumbot | 60.63\(\pm\)0.7 | 62.42\(\pm\)2.4 | 59.86\(\pm\)0.6 | 77.69\(\pm\)4.2 | 78.34\(\pm\)1.9 | 82.85\(\pm\)2.9 | 83.49\(\pm\)1.9 | 84.01\(\pm\)1.9 | |
| mpot | 70.02\(\pm\)2.0 | 74.64\(\pm\)4.4 | 61.62\(\pm\)1.3 | 78.40\(\pm\)3.9 | 70.96\(\pm\)3.7 | 86.65\(\pm\)5.1 | 86.69\(\pm\)5.1 | 86.95\(\pm\)5.0 |