Vision Transformers Learn Gestalt-Like Figure-Ground Cues from Natural Images

Matthias Tangemann\(^{1,2}\) Benjamin Lo\(^{3,4}\) Zygmunt Pizlo\(^{5}\) Kaleem Siddiqi\(^{3,4}\)
Dirk B. Walther\(^{1}\) Sven Dickinson\(^{1,2}\)

\(^1\)University of Toronto \(^2\)Vector Institute \(^3\)McGill University \(^4\)MILA \(^5\)UC Irvine

mtangemann@cs.toronto.edu


Abstract

Figure-ground organization in the human visual system relies on several shape-based cues, including surroundedness, convexity, and symmetry. While these cues have been extensively studied using abstract stimuli, little is known about how they operate under natural conditions or how they arise from the statistics of natural scenes. Deep neural networks offer a promising path forward: a model that relies on the same figure-ground cues as humans would provide tractable experimental access to the underlying mechanisms. In this study, we evaluate shape-based figure-ground organization in Vision Transformers (ViTs), for which prior work has demonstrated the emergence of object-based grouping. We test 25 ViTs spanning supervised and self-supervised training objectives, by fitting linear probes to predict figure-ground assignment from intermediate patch representations using both natural images and controlled artificial stimuli that isolate individual cues. Our results show that ViTs robustly encode surroundedness and convexity, and that probes trained on natural images generalize zero-shot to artificial stimuli across several models. For symmetry we observe mixed results: the cue is encoded for uniformly colored but not for textured regions. Taken together, our findings demonstrate that Gestalt-like figure-ground cues can be learned from natural scene statistics and position ViTs as a compelling model system for studying the computational mechanisms of perceptual organization.

Code and data is available at https://github.com/mtangemann/mlvbench.

1 Introduction↩︎

Figure-ground organization is one of the most fundamental processes in human visual perception. Research over the past century has identified several shape-based cues that drive this process, including surroundedness, convexity, and symmetry [1]. These cues have been extensively characterized using controlled, artificial stimuli (e.g., [2]), yet we lack a precise understanding of how they contribute to the perceptual organization of natural scenes. It has been hypothesized that many visual cues are rooted in the statistics of natural scenes [3], [4], but the mechanisms that enable learning general cues from experience remain unknown.

Figure 1: We extract patch representations from frozen, pre-trained ViTs and fit linear probes to predict figure-ground assignment. We evaluate probes on both natural images and controlled stimuli that isolate individual shape cues while semantic information, texture, and region size are uninformative. We exclude patches that span both foreground and background (shown in black).

Computational models that rely on the same figure-ground cues as humans could offer tractable experimental access to these open questions. Such models could serve a role analogous to model organisms in neuroscience: they allow for controlled experimentation that is difficult or impossible in human observers, while generating hypotheses that can subsequently be tested in humans. Vision Transformers (ViTs) are strong candidates for such model systems. Their self-attention mechanism enables global processing across the image, and several studies have demonstrated that structured scene representations emerge in their intermediate layers [5][7]. These studies have established that ViTs develop rich segmentation capabilities, but the question of what cues underlie these capabilities remains open. In particular, it is unclear whether ViTs rely on generalizable, shape-based cues as humans do, or whether they primarily rely on semantic and textural regularities.

In this work, we address this gap by studying whether ViTs encode specific, well-characterized shape cues for figure-ground organization. We fit linear probes to classify patches as foreground or background (Figure 1), evaluating not only whether shape cues are represented but also where in the processing hierarchy they emerge. We test 25 ViTs spanning diverse training objectives across both natural images and synthetic stimuli that isolate surroundedness, convexity, and symmetry individually. Crucially, our synthetic stimuli are constructed so that semantic content, texture, and region size are uninformative, ensuring that probe performance reflects genuine sensitivity to shape.

We find strong evidence for shape-based figure-ground processing across a broad range of ViTs. All tested models encode information about surroundedness and convexity, and for several models probes trained on natural images generalize zero-shot to synthetic stimuli. The results for symmetry are more nuanced and show an interaction between texture content and symmetry cues. Overall, these results demonstrate that generic, Gestalt-like figure-ground cues can be learned from the statistics of natural scenes. A layer-wise analysis further reveals that the strongest figure-ground representations emerge in intermediate layers rather than in the final representation, with notable differences between training objectives: self-supervised models retain more figure-ground information in later layers than supervised and CLIP-like models.

In summary, our paper makes the following contributions:

  • We present the first systematic study of generic, shape-based figure-ground cues in ViTs.

  • We demonstrate that probes trained on natural images generalize to synthetic stimuli isolating individual shape cues, showing that human-like figure-ground cues can be learned from natural scene statistics.

  • We provide a layer-wise analysis across 25 models, revealing that figure-ground information peaks in intermediate layers and that the training objective systematically affects where this information is retained.

Our findings establish ViTs as a viable model system for studying perceptual organization in silico. To support future research in this direction, we will publish data, code, and our pretrained probes for all experiments in this paper.

2 Related Work↩︎

2.0.0.1 Perceptual Organization.

Perceptual organization refers to the visual system’s ability to structure sensory input into coherent objects and surfaces, a process guided by several well-established cues [1], [2], [8][10]. Among these, surroundedness, convexity, and symmetry play central roles in figure–ground assignment and shape interpretation. Regions that are spatially surrounded are more likely to be perceived as figures rather than background, while convex regions tend to dominate over concave counterparts in determining object boundaries. Symmetry further biases perception by promoting the grouping of elements into unified forms, reflecting the visual system’s sensitivity to regularity and structural simplicity. Together, these cues illustrate how mid-level vision resolves ambiguity by leveraging probabilistic regularities in natural scenes.

2.0.0.2 Human vs Machine Vision.

Multiple lines of work investigate whether the strong capabilities of DNNs for computer vision tasks are enabled by internal processing similar to humans, mostly with a focus on core object recognition (see [11] for a recent review). The application of DNNs in neuroscience has been especially successful: DNNs outperform all other models for predicting the neuronal activity in visual areas of humans and other primates [12], [13]. The results are more nuanced in behavioral comparisons to human perception. DNNs have been found to be less robust than humans [14][16], to rely more on texture than shape cues ([17], but see [18]), and to make different errors than humans [19]. However, the gap is narrowing: by scaling model size and training data, and by using richer pretraining tasks than object recognition, models can more closely approximate human visual perception [20], [21].

Several authors have compared humans against machines in tasks which go beyond core object recognition, such as the perception of motion [22][24] and depth [25], [26]. Another line of work has investigated whether DNNs follow Gestalt principles, and in particular closure, reporting both successes and failure cases [6], [27][34]. However, drawing reliable conclusions can be difficult. Many studies rely on training or finetuning models which may alter their internal feature spaces. Moreover, the rapid pace of research in machine learning quickly makes studies obsolete, leaving us with limited insight into the current generation of vision encoders. Exceptions include recent studies which demonstrate the emergence of human-like, object-level grouping in Vision Transformers [7], [35].

2.0.0.3 Probing.

Probing is a standard paradigm in interpretability research. Hidden features are extracted from a pretrained network, and a small readout model is trained or applied to test whether a specific property is encoded in the internal representation. Probing has been traditionally used to evaluate the representation of self-supervised models, e.g., by predicting ImageNet classes using a linear or k-NN readout [5], [36], [37]. Moreover, several studies have probed the dense representation of the patches in Vision Transformers, showing that multiple foundation models encode mid-level properties such as correspondence, depth, or scene geometry [38][40]. Several works have further demonstrated that object segmentation can be decoded from internal features [5], [7], [35], [41][43]. To our knowledge, it has not been investigated yet whether the internal cues used for segmentation align with human visual perception.

3 Methods↩︎

We evaluate whether the intermediate features of pretrained ViTs encode surroundedness, convexity and symmetry as generic shape cues for figure-ground assignment. To this end, we train linear probes to classify each patch as figure or ground, based on frozen features from pre-trained ViTs, and assess their performance relative to a spatial prior baseline. We evaluate on both natural stimuli and synthetic datasets designed such that only a specific shape cue is informative for distinguishing figure and ground. Below, we describe the probe fitting and evaluation (Section 3.1), the stimuli (Section 3.2), and the models considered (Section 3.3).

3.1 Probe Fitting and Evaluation↩︎

We fit linear probes on a dataset \(D\) of images \(x \in \mathbb{R}^{H\times W\times 3}\) and ground truth segmentation masks \(s \in \{0,1\}^{H\times W}\), using the native input resolution of each model. For each layer \(l\), we extract intermediate features \(f^{(l)} \in \mathbb{R}^{h\times w\times C}\), where \((h,w) = (H/p, W/p)\) and \(p\) is the patch size of the model. We downscale the ground truth segmentation mask to the internal grid, yielding \(\tilde{s} \in \{0,1\}^{h\times w}\). Patches that overlap both foreground and background in the original mask receive ambiguous labels and are excluded from training and evaluation.

3.1.0.1 Spatial prior.

In all datasets, patches near the image center are more likely to belong to the figure than patches near the boundary. We quantify this center bias with a spatial prior that estimates the expected foreground probability at each grid location \((i, j)\):

\[\text{prior}_{i,j} = \frac{1}{|D_{train}|} \sum_{\tilde{s} \in D_{train}} \tilde{s}_{i,j}\]

This prior serves two purposes: it acts as a baseline against which we measure probe performance, and it informs the design of the probes themselves, as described next.

3.1.0.2 Probe definition.

For each model layer, we fit a logistic regression model that predicts the probability of a patch belonging to the foreground:

\[\text{probe}_{i,j}(f_{i,j}) = \sigma(w^Tf_{i,j} + b_{i,j})\]

The probe shares a single weight vector \(w\) across all spatial positions but uses a separate bias term \(b_{i,j}\) for each location, initialized to the log-odds of the spatial prior at that position. This design choice is important for the interpretability of our results. With a standard scalar bias, the probe would need to learn the spatial layout of foreground likelihood from the features themselves; a failure to do so could then be conflated with a lack of shape information. By providing explicit access to the spatial prior through position-dependent biases and allowing the model to override them during training, we ensure that any improvement over the prior can be cleanly attributed to shape information in the features.

3.1.0.3 Training.

Probes are trained with binary cross-entropy (BCE) loss using the Adam optimizer [44] with a batch size of 256, a learning rate of 0.0001, and no weight decay. We train each probe for 2000 steps. Inspection of validation loss curves confirmed that all probes converge well before this limit. We fit independent probes for each model layer and select the best layer per model on the validation set. All results are reported on held-out test sets. Probes for all models were fit using a single NVIDIA L40S GPU within 100 GPU hours.

3.1.0.4 Evaluation metric.

We evaluate probes using information gain explained (IGE), which measures how much the probe improves over the spatial prior baseline. We first define the information gain as the reduction in BCE loss relative to the prior:

\[\text{IG} = \text{BCE}_{prior} - \text{BCE}_{probe}\]

where both terms are computed over all patches in the test set. To allow direct comparison across datasets with different baseline difficulties, we normalize by the prior loss:

\[\text{IGE} = \frac{\text{IG}}{\text{BCE}_{prior}} = 1 - \frac{\text{BCE}_{probe}}{\text{BCE}_{prior}}\]

An IGE of 0% indicates that the probe performs no better than the spatial prior, while an IGE of 100% indicates perfect prediction. We additionally report accuracies in Appendix 7. Unlike accuracy, IGE accounts for the confidence of predictions and is therefore a more sensitive measure of the information encoded in the features.

3.2 Stimuli↩︎

3.2.0.1 Natural stimuli.

We use images and ground truth segmentation masks from the MSRA-10K dataset [45], which contains single foreground objects in background scenery. The dataset is split into 80% training, 10% validation, and 10% test images. Each image is resized so that the shorter side matches the model’s native input resolution, and is then center-cropped to a square.

3.2.0.2 Synthetic stimuli.

For each synthetic image, we generate a segmentation mask isolating a single shape cue and fill the foreground and background regions with randomly assigned textures from the DTD dataset [46]. Textures are split into non-overlapping train/validation/test sets (80/10/10%). The foreground region is scaled to cover exactly half the image area so that region size is uninformative. Since textures are also randomly assigned, the only valid cue for figure-ground assignment is the region shape. We construct three conditions (see Figures 1 and 2 for exampels):

  • Surroundedness: The foreground is a random shape from the Infinite DSprites dataset [47], positioned with at least a one-pixel margin to all image boundaries so that it is fully surrounded by the background.

  • Convexity: Foreground and background are separated by a random parabolic arc, \(v = ku^{2} + c\), in a randomly oriented coordinate frame, with \(k \sim \text{Uniform}(1.0, 3.0)\) and \(c\) chosen to ensure equal region size. The convex side is labeled as foreground.

  • Symmetry: The image is divided into four equally sized vertical columns by three Bézier-curve edges. Two consecutive edges are made mirror-symmetric by copying and inverting their control points, producing two symmetric columns (foreground) flanked by two asymmetric columns (background). This design follows the classic Bahnsen column paradigm [48].

Figure 2: Example probe predictions for each stimulus condition for representations extracted after layer 11 of BEiT-3 VIT-L [49]. Red and blue patches denote foreground and background, respectively. Black patches cover both foreground and background pixels and are thus ignored during training and evaluation.

3.2.0.3 Additional controls.

To empirically rule out texture-based strategies, we construct several control test sets. First, for each condition we create a texture-reversed variant in which foreground and background textures are swapped while keeping the same shape-defined ground truth. Second, for surroundedness and convexity, we generate variants where the foreground is split into two sub-regions by a straight line, disrupting local texture coherence (see Figure 2 for examples). For symmetry, pilot experiments revealed that models struggle to differentiate symmetric from asymmetric regions when filled with textures, suggesting interference between texture content and boundary processing. As a control, we therefore add a training condition in which textures are replaced with uniform colors.

3.3 Models↩︎

We evaluate a diverse set of 25 Vision Transformer models. For all models, we use the implementations and checkpoints as provided by the timm library [50] and test the ViT-B and ViT-L variants. We consider ViTs trained for object recognition on ImageNet [51], including an improved version of the original ViT (AugReg, [52]), DeIT III [53] and FlexiViT [54]. Vision-Language Alignment models have been pioneered by CLIP [55]. We additionally consider the more recent SigLIP 2 [56] and Perception Encoder models [57]. Masked Image Modelling is a self-supervised pretraining task where the model is trained to reconstruct missing patches in the input image. We include the standard Masked Auteoncoder [58] and BEiT [59]. Further, we consider EVA-02 [60], which is trained to reconstruct CLIP features for masked patches, and BEiT-3 [49], which is jointly trained on masked images and text. Self-Distillation is used by the DINO models. We include the original DINO model, DINOv2 with registers, and DINOv3 [5], [21], [61], [62]. Links to the precise model checkpoints are provided in Table 3 in the appendix.

4 Results↩︎

We present results for all 25 ViTs (listed in Section 3.3) across natural and synthetic conditions. We begin with probe performance on natural images and the cue-isolation conditions, then demonstrate that probes trained on natural images generalize zero-shot to synthetic stimuli, before analyzing the effects of pre-training objective, model size, and layer depth.

Figure 3: Overall probe performances for the different stimulus conditions. For each model, the best layer has been selected on the respective validation set. Performance is measured using information gain explained (IGE), where 0 corresponds to chance-level performance and 1 to a perfect prediction. The same data is provided as tables in Appendix 7.
Figure 4: Zero-shot generalization of probes trained on natural images to the surroundedness and convexity conditions. The x-axis shows performance of probes trained specifically for each condition; the y-axis shows performance of probes trained on natural images and evaluated on the same synthetic test sets. Symmetry conditions are excluded due to low performance in the textured condition.
Figure 5: Probe performance by model size (top row) and pretraining task (bottom row: OR = Object Recognition, VLA = Vision-Language Alignment, MIM = Masked Image Modeling, SD = Self-Distillation). Brackets indicate statistical significance (see Appendix 10) (*/**/***: p<.05/.01/.001, non-significant pairs are omitted for clarity).
Figure 6: Probe performance on standard test sets compared to reversed-texture and split-foreground variants (see Figure 2 for examples). Probes generalize almost perfectly to reversed textures and maintain strong performance on split foreground shapes not seen during training.
Figure 7: Left: Per-layer performances for each condition. The best layer selected on the validation set is marked with a colored dot. For most models, figure-ground information is best decoded from intermediate layers. More detailed information is provided in Figure 15 in the appendix. Right: Relative performance drop of the final layer compared to the best layer. Self-distillation models consistently retain the most shape-cue information across layers.

4.0.0.1 ViTs represent figure-ground assignment for natural images.

The results in Figure 3 show that patch tokens from all ViTs encode sufficient information to segment the foreground object in natural images. The supervised DeiT III leads with 87% IGE, followed closely by several masked-image models. The Perception Encoder performs worst in this setting at 75% IGE—still well above the spatial prior. Example predictions from the best and worst models (Section 8 in the appendix) illustrate strong performance across ViTs and reveal that many errors stem from genuinely ambiguous cases.

4.0.0.2 ViTs represent surroundedness and convexity.

The results in Figure 3 and example predictions in Figure 2 show that ViTs excel in the surroundedness condition, where probes for several models achieve near-perfect predictions. The results for convexity are more varied: the best model (BEiT-3) achieves near-perfect predictions, while the worst (FlexiViT) reaches 66% IGE.

The control experiments in Figure 6 confirm that these results reflect genuine shape sensitivity. Probes generalize to the reversed-texture test set without any loss in performance, ruling out spurious texture influence. Probes also generalize well when the foreground is randomly split into two textured sub-regions, although this configuration was not seen during training. Together, these results demonstrate that ViT patch representations encode shape cues related to surroundedness and convexity, rather than relying solely on semantic or texture information.

4.0.0.3 Some ViT representations generalize from natural images to synthetic surroundedness and convexity conditions.

To test whether surroundedness and convexity are linked to figure-ground organization more broadly, we evaluate whether probes trained on natural images generalize zero-shot to the synthetic conditions. The results in Figure 4 show that probes generalize well for several models. For surroundedness, the natural-image probes of several models perform only slightly below probes specifically trained on the synthetic condition, demonstrating a strong link between figure-ground organization and a generic, shape-based surroundedness cue. The generalization gap is more pronounced for convexity, but for many models, natural-image probes still perform substantially above the spatial prior when convexity is the only valid cue. This suggests a somewhat weaker but nonetheless substantial link between figure-ground organization and a generic notion of convexity.

4.0.0.4 Symmetry is represented only in the absence of texture.

Probe performance drops substantially in the symmetry condition (Figure 3). No model’s representation supports differentiation of symmetric from asymmetric regions above the spatial prior baseline. When textures are removed, however, several models allow identification of symmetric regions with near-perfect accuracy (e.g., DINOv3 and BEiT-3). We therefore hypothesize that the failure in the textured condition points to interference between texture content and boundary processing, rather than a fundamental inability to encode symmetry. This pattern is broadly consistent with symmetry being a weaker figure-ground cue in human vision as well [1], [10].

4.0.0.5 Both model size and pre-training influence figure-ground representation.

In Figure 5, we analyze probe performance by model size and pre-training task. In all conditions, ViT-L models performed significantly better than the smaller ViT-B models. Furthermore, we observe pre-training task and dataset influencing performance. Due to the small sample size, we only see significant differences in a few cases. Overall, masked image models outperform models trained for object recognition and vision-language alignment.

4.0.0.6 Intermediate layers encode figure-ground cues more strongly than the final layer.

All preceding results used the best layer per model, selected on the validation set. Figure 7 analyzes where in the network figure-ground information is most accessible. Across models and conditions, the strongest representations are found in intermediate layers, at a median depth of approximately 60% of the network. We further examine the performance drop between the best intermediate layer and the final layer. Self-distillation models (DINO) retain nearly all figure-ground information through to the final layer, while other training objectives produce more pronounced drops. These results highlight the importance of layer-wise analysis: evaluating only the final layer would systematically favor self-supervised models and underestimate the prominence of shape-cue representations of other models.

5 Limitations↩︎

We rely on patch-wise, linear probes to decode figure-ground information. This ensures that any positive result reflects readily accessible information, that can be used by the key and query projections in subsequent layers. For the symmetry condition, where regions could not be linearly separated, it is nevertheless possible that a more powerful probe could successfully decode symmetry information.

Moreover, whereas we focus on three well-studied shape cues, human figure-ground organization involves additional factors such as small area, lower region, and familiarity. Our framework extends naturally to these other cues, and we see their investigation as a promising direction for future work.

Finally, the present study does not include a direct comparison with human behavioral data. For the well-established cues studied here, the qualitative alignment with human findings provides a solid foundation. Direct human comparisons will become particularly valuable as future work moves beyond the examination of these canonical cues, towards more fine-grained questions to examine cue interactions and relative cue strengths.

6 Discussion↩︎

We performed a systematic probing study evaluating whether ViTs encode shape-based figure-ground cues known from human vision. Our results demonstrate that surroundedness and convexity are robustly represented across a diverse set of 25 ViTs, and that probes trained on natural images generalize zero-shot to synthetic stimuli isolating these cues in several models. This provides direct evidence that generic, Gestalt-like figure-ground cues can be learned from the statistics of natural scenes and provides computational support to the longstanding hypothesis that perceptual organization is rooted in ecological statistics [3], [4].

These findings also challenge the common characterization of deep neural networks as primarily texture-driven. Our results show that shape-based processing sufficient for figure-ground organization coexists with texture information in the same representations. The emergence of these cues across supervised, self-supervised, and vision-language models suggests that learning shape-based figure-ground cues is a robust phenomenon rather than an artefact of any particular training paradigm.

The results for symmetry point to an interesting direction for future work. The failure in the textured condition, combined with near-perfect performance when textures are removed, suggests a specific interference between texture content and boundary-based symmetry processing. Understanding this interaction, and whether it parallels known limitations of symmetry as a figure-ground cue in humans, could yield insights into how different cues interact and compete during perceptual organization.

More broadly, our work establishes ViTs as a viable model system for studying perceptual organization in silico. Our framework of combining natural-image probing with controlled synthetic conditions that isolate individual cues can be extended to additional cues such as small area, lower region, and parallelism, and to studying cue combination and competition. Our layer-wise analyses open the door to investigating how figure-ground cues are computed across processing stages, with potential parallels to hierarchical processing in human vision. We envision that this line of research will foster closer collaboration between research in interpretability and perceptual science.

This work was supported by the Natural Sciences and Engineering Research Council of Canada (NSERC) and Samsung. The authors thank the Digital Research Alliance of Canada (alliancecan.ca) for providing computing resources.

7 Detailed Results↩︎

Model ACC IGE \(\downarrow\)
DeiT-3-L 98.1 87.4
image
BEiT-3-L 98.0 87.0
image
BEiT-L 98.0 86.7
image
MAE-L 98.1 86.6
image
DeiT-3-B 97.9 85.9
image
BEiT-B 97.9 85.4
image
BEiT-3-B 97.6 85.3
image
EVA-02-L 97.9 85.2
image
MAE-B 97.8 85.1
image
EVA-02-B 97.5 84.0
image
DINO-B 97.5 83.8
image
SigLIP-2-L 97.6 83.8
image
DINOv3-L 97.5 83.5
image
FlexiViT-L 97.4 83.1
image
DINOv3-B 97.3 82.4
image
CLIP-L 97.2 82.3
image
AugReg-B 97.2 81.9
image
AugReg-L 97.2 81.3
image
DINOv2-L 96.8 81.2
image
CLIP-B 96.8 79.9
image
PE-B 96.6 79.0
image
SigLIP-2-B 96.8 78.8
image
DINOv2-B 96.3 78.5
image
FlexiViT-B 96.2 76.3
image
PE-L 95.9 74.9
image
Trained for surroundedness Trained for natural images
Model ACC IGE \(\downarrow\) ACC IGE
BEiT-3-L 100.0 99.6
image
99.0 93.4
image
DeiT-3-L 99.9 99.1
image
97.8 88.6
image
BEiT-3-B 99.9 98.9
image
96.5 82.1
image
MAE-L 99.9 98.8
image
97.4 84.5
image
DINOv3-L 99.9 98.7
image
95.6 78.8
image
EVA-02-L 99.9 98.6
image
98.5 91.3
image
BEiT-L 99.9 98.6
image
97.4 86.1
image
DeiT-3-B 99.8 98.3
image
95.9 80.2
image
SigLIP-2-L 99.9 97.9
image
98.5 88.0
image
MAE-B 99.8 97.8
image
96.9 80.9
image
CLIP-L 99.8 97.5
image
96.7 83.7
image
DINOv3-B 99.7 97.3
image
94.4 71.0
image
DINOv2-L 99.6 97.0
image
96.4 81.0
image
BEiT-B 99.7 96.7
image
95.1 74.0
image
EVA-02-B 99.7 96.5
image
96.6 82.7
image
SigLIP-2-B 99.7 96.1
image
97.1 79.6
image
DINOv2-B 99.5 96.0
image
95.7 77.3
image
DINO-B 99.6 95.3
image
94.6 73.0
image
CLIP-B 99.7 95.1
image
97.4 83.9
image
FlexiViT-L 99.6 94.6
image
90.8 62.8
image
AugReg-L 99.5 94.1
image
91.9 60.9
image
PE-B 99.4 93.4
image
93.5 66.9
image
AugReg-B 99.4 93.3
image
93.3 66.2
image
PE-L 98.7 89.0
image
88.3 39.7
image
FlexiViT-B 98.6 83.4
image
84.2 35.5
image
Trained for convexity Trained for natural images
Model ACC IGE \(\downarrow\) ACC IGE
BEiT-3-L 99.3 96.1
image
87.8 65.4
image
DeiT-3-L 98.7 94.5
image
76.8 25.2
image
BEiT-3-B 98.7 94.1
image
80.1 39.1
image
DeiT-3-B 98.1 92.2
image
71.8 8.1
image
DINOv3-L 98.3 92.1
image
76.8 29.5
image
BEiT-L 97.6 89.2
image
80.9 41.7
image
EVA-02-L 97.6 89.2
image
84.7 56.5
image
DINO-B 97.2 87.2
image
77.5 30.2
image
DINOv3-B 97.1 86.6
image
73.7 12.2
image
MAE-L 96.7 85.4
image
77.9 24.3
image
DINOv2-L 96.2 83.9
image
80.9 41.9
image
CLIP-L 96.0 83.7
image
76.0 27.8
image
SigLIP-2-L 95.8 81.7
image
81.5 48.7
image
FlexiViT-L 95.7 81.2
image
74.3 27.5
image
BEiT-B 95.3 79.8
image
75.2 15.3
image
MAE-B 95.2 79.6
image
75.5 17.1
image
EVA-02-B 94.9 79.6
image
79.9 39.9
image
CLIP-B 95.3 78.4
image
75.4 35.7
image
DINOv2-B 94.6 77.3
image
77.2 32.9
image
AugReg-B 94.5 76.8
image
75.3 25.4
image
AugReg-L 94.1 76.6
image
75.8 23.1
image
SigLIP-2-B 94.1 75.2
image
76.0 34.5
image
PE-B 92.9 72.2
image
71.3 17.6
image
PE-L 92.3 69.4
image
69.5 -4.5
image
FlexiViT-B 91.4 66.1
image
67.1 9.0
image
Texture No Texture
Model ACC IGE ACC IGE \(\downarrow\)
DINOv3-L 61.1 5.9
image
99.8 99.2
image
BEiT-3-L 63.7 8.4
image
99.7 98.5
image
BEiT-L 60.8 5.6
image
99.6 98.1
image
DINOv2-L 57.7 2.6
image
99.7 98.0
image
DINOv3-B 56.1 1.6
image
99.5 97.6
image
MAE-L 62.8 7.6
image
98.7 94.4
image
BEiT-3-B 57.4 2.4
image
98.1 92.8
image
DeiT-3-L 56.4 1.1
image
96.9 89.1
image
BEiT-B 56.3 1.7
image
97.3 88.3
image
MAE-B 61.1 6.0
image
96.3 85.1
image
DINOv2-B 51.8 -0.1
image
94.8 79.8
image
CLIP-L 56.0 1.8
image
92.1 71.0
image
PE-L 55.0 1.2
image
91.2 70.6
image
EVA-02-L 61.4 6.0
image
91.8 69.2
image
SigLIP-2-L 59.1 3.6
image
92.3 66.9
image
PE-B 54.8 1.1
image
89.3 65.5
image
EVA-02-B 56.9 2.6
image
89.5 65.1
image
DeiT-3-B 48.4 -0.4
image
86.3 57.6
image
DINO-B 53.4 0.3
image
87.3 54.8
image
FlexiViT-L 54.8 1.3
image
86.0 51.7
image
AugReg-L 56.5 2.3
image
81.5 49.4
image
CLIP-B 54.1 0.8
image
83.5 46.6
image
SigLIP-2-B 54.4 1.0
image
78.2 36.1
image
AugReg-B 52.4 0.2
image
74.5 30.5
image
FlexiViT-B 50.9 -0.2
image
73.0 22.7
image

8 Visualization of Predictions↩︎

Figure 8: Probe predictions for the best and worst models in the natural condition
Figure 9: Probe predictions for the best and worst models in the surroundedness condition
Figure 10: Zero-shot probe predictions for the best and worst models in terms of generalization from natural images to the surroundedness condition.
Figure 11: Probe predictions for the best and worst models in the convexity condition.
Figure 12: Zero-shot probe predictions for the best and worst models in terms of generalization from natural images to the convexity condition.
Figure 13: Probe predictions for the best and worst models in the symmetry condition.
Figure 14: Probe predictions for the best and worst models in the symmetry condition without textures.

9 Performance by Layer↩︎

Figure 15: Individual probe performance for each model across layers. Performance is measured in Information Gain Explained (IGE).

10 Detailed significance results for Figure 5↩︎

Table 1: Summary statistics per condition. \(W\): Wilcoxon signed-rank statistic (base vs. large, matched by model family); \(r\): effect size \(|Z|/\sqrt{n}\), with \(Z\) the normal approximation of \(W\) and \(n\) the number of non-zero pairs; \(H\): Kruskal-Wallis statistic (pre-training objective); \(\epsilon^2 = H/(N-1)\). Significance: \(^{*}p<0.05\), \(^{**}p<0.01\), \(^{***}p<0.001\).
Condition N N pairs W p(W) r H df p(H) \(\epsilon^2\)
Natural 25 12 11.0 0.027* 0.634 12.372 3 0.006** 0.516
Surroundedness 25 12 11.0 0.027* 0.634 7.452 3 0.059 0.311
Convexity 25 12 5.0 0.005** 0.770 6.319 3 0.097 0.263
Symmetry (Texture) 25 12 0.0 <0.001*** 0.883 12.269 3 0.007** 0.511
Symmetry (No Texture) 25 12 0.0 <0.001*** 0.883 11.369 3 0.010** 0.474
Table 2: Dunn post-hoc p-values (Holm-corrected) for all tasks. Significance: \(^{*}p<0.05\), \(^{**}p<0.01\), \(^{***}p<0.001\).
Pair Natural Surroundedness Convexity Symmetry (Texture) Symmetry (No Texture)
MIM vs OR 0.306 0.118 0.849 0.006** 0.039*
OR vs SD 0.964 1.000 0.923 0.906 0.055
SD vs VLA 0.964 1.000 0.318 0.906 0.144
MIM vs SD 0.082 0.760 0.923 0.169 1.000
OR vs VLA 0.440 1.000 0.923 0.906 1.000
MIM vs VLA 0.005** 0.113 0.116 0.087 0.144

11 Licenses↩︎

We use the DTD dataset [46], which does not provide a standard license but “is made available to the computer vision community for research purposes.” (https://www.robots.ox.ac.uk/ vgg/data/dtd/, 2026-05-06).

Similarly, the MSRA-10K dataset [45] does not provide a standard license but requires researchers to cite their paper (https://mmcheng.net/msra10k/, 2026-05-05).

We further use shapes from the Infinite DSprites dataset [47] which is licensed under the MIT license (https://github.com/sbdzdz/idsprites/, 2026-05-05).

The licenses and model cards for all models evaluated in our work are listed in Table 3.

Table 3: Vision Transformer checkpoints used in this work and their corresponding Hugging Face model cards and licenses.
Model Checkpoint License
AugReg-B timm/vit_base_patch16_224.augreg_in21k_ft_in1k Apache-2.0
AugReg-L timm/vit_large_patch16_224.augreg_in21k_ft_in1k Apache-2.0
BEiT-3-B timm/beit3_base_patch16_224.in22k_ft_in1k MIT
BEiT-3-L timm/beit3_large_patch16_224.in22k_ft_in1k MIT
BEiT-B timm/beit_base_patch16_224.in22k_ft_in22k MIT
BEiT-L timm/beit_large_patch16_224.in22k_ft_in22k MIT
CLIP-B timm/vit_base_patch16_clip_224.openai MIT
CLIP-L timm/vit_large_patch14_clip_224.openai MIT
DINO-B timm/vit_base_patch16_224.dino Apache-2.0
DINOv2-B timm/vit_base_patch14_reg4_dinov2.lvd142m Apache-2.0
DINOv2-L timm/vit_large_patch14_reg4_dinov2.lvd142m Apache-2.0
DINOv3-B timm/vit_base_patch16_dinov3.lvd1689m DINOv3 License
DINOv3-L timm/vit_large_patch16_dinov3.lvd1689m DINOv3 License
DeiT-3-B timm/deit3_base_patch16_224.fb_in22k_ft_in1k Apache-2.0
DeiT-3-L timm/deit3_large_patch16_224.fb_in22k_ft_in1k Apache-2.0
EVA-02-B timm/eva02_base_patch14_224.mim_in22k MIT
EVA-02-L timm/eva02_large_patch14_224.mim_in22k MIT
FlexiViT-B timm/flexivit_base.1200ep_in1k Apache-2.0
FlexiViT-L timm/flexivit_large.1200ep_in1k Apache-2.0
MAE-B timm/vit_base_patch16_224.mae CC-BY-NC-4.0
MAE-L timm/vit_large_patch16_224.mae CC-BY-NC-4.0
PE-B timm/vit_pe_core_base_patch16_224.fb Apache-2.0
PE-L timm/vit_pe_core_large_patch14_336.fb Apache-2.0
SigLIP-2-B timm/vit_base_patch16_siglip_224.v2_webli Apache-2.0
SigLIP-2-L timm/vit_large_patch16_siglip_256.v2_webli Apache-2.0

References↩︎

[1]
J. Wagemans, J. H. Elder, M. Kubovy, S. E. Palmer, M. A. Peterson, M. Singh, and R. von der Heydt. A century of Gestalt psychology in visual perception: I. Perceptual grouping and figure–ground organization. Psychological Bulletin, 138 (6): 1172–1217, Nov. 2012. ISSN 1939-1455. .
[2]
M. A. Peterson and E. Salvagio. Inhibitory competition in figure-ground perception: Context and convexity. Journal of Vision, 8 (16): 4, Dec. 2008. ISSN 1534-7362. .
[3]
W. S. Geisler, J. S. Perry, B. J. Super, and D. P. Gallogly. Edge co-occurrence in natural images predicts contour grouping performance. Vision Research, 41 (6): 711–724, Mar. 2001. ISSN 0042-6989. .
[4]
J. H. Elder and R. M. Goldberg. Ecological statistics of Gestalt laws for the perceptual organization of contours. Journal of Vision, 2 (4): 5, Aug. 2002. ISSN 1534-7362. .
[5]
M. Caron, H. Touvron, I. Misra, H. Jégou, J. Mairal, P. Bojanowski, and A. Joulin. Emerging Properties in Self-Supervised Vision Transformers. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pages 9650–9660, Oct. 2021.
[6]
T. Li, Z. Wen, L. Song, J. Liu, Z. Jing, and T. S. Lee. From Local Cues to Global Percepts: Emergent Gestalt Organization in Self-Supervised Vision Models, May 2025.
[7]
H. Adeli, S. Ahn, A. Luo, M. Zhang, N. Kriegeskorte, and G. Zelinsky. Human-like Object Grouping in Self-supervised Vision Transformers, Mar. 2026.
[8]
S. E. Palmer. Vision Science by Stephen E. Palmer . MIT Press, Cambridge, MA, Apr. 1999. ISBN 978-0-262-16183-1.
[9]
J. M. Wolfe, K. R. Kluender, D. M. Levi, L. M. Bartoshuk, R. S. Herz, R. L. Klatzky, and D. M. Merfeld. Sensation & Perception. Oxford University Press, New York, NY, USA, sixth edition edition, 2021. ISBN 978-1-60535-972-4.
[10]
G. Kanisza, W. Gerbino, and M. Henle. Convexity and symmetry in figure-ground organization. In Vision and Artifact, pages 25–32. Springer, New York, 1976.
[11]
F. A. Wichmann and R. Geirhos. Are Deep Neural Networks Adequate Behavioral Models of Human Visual Perception? Annual Review of Vision Science, 9: 501–524, Sept. 2023. ISSN 2374-4642, 2374-4650. .
[12]
D. L. K. Yamins, H. Hong, C. F. Cadieu, E. A. Solomon, D. Seibert, and J. J. DiCarlo. Performance-optimized hierarchical models predict neural responses in higher visual cortex. Proceedings of the National Academy of Sciences, 111 (23): 8619–8624, May 2014. .
[13]
C. Conwell, J. S. Prince, K. N. Kay, G. A. Alvarez, and T. Konkle. A large-scale examination of inductive biases shaping high-level visual representation in brains and machines. Nature Communications, 15 (1): 9383, Oct. 2024. ISSN 2041-1723. .
[14]
R. Geirhos, C. R. M. Temme, J. Rauber, H. H. Schütt, M. Bethge, and F. A. Wichmann. Generalisation in humans and deep neural networks. In S. Bengio, H. Wallach, H. Larochelle, K. Grauman, N. Cesa-Bianchi, and R. Garnett, editors, Advances in Neural Information Processing Systems, volume 31. Curran Associates, Inc., Dec. 2018.
[15]
D. Hendrycks and T. Dietterich. Benchmarking Neural Network Robustness to Common Corruptions and Perturbations. In International Conference on Learning Representations, Sept. 2018.
[16]
D. Hendrycks, S. Basart, N. Mu, S. Kadavath, F. Wang, E. Dorundo, R. Desai, T. Zhu, S. Parajuli, M. Guo, D. Song, J. Steinhardt, and J. Gilmer. The Many Faces of Robustness: A Critical Analysis of Out-of-Distribution Generalization. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 8340–8349, 2021.
[17]
R. Geirhos, P. Rubisch, C. Michaelis, M. Bethge, F. A. Wichmann, and W. Brendel. are biased towards texture; increasing shape bias improves accuracy and robustness. In 7th International Conference on Learning Representations (ICLR) 2019, May 2019.
[18]
T. Burgert, O. Stoll, P. Rota, and B. Demir. are not biased towards texture: Revisiting feature reliance through controlled suppression, Oct. 2025.
[19]
R. Geirhos, K. Meding, F. A. Wichmann, H. Larochelle, M. Ranzato, R. Hadsell, M. F. Balcan, and H. Lin. Beyond accuracy: Quantifying trial-by-trial behaviour of CNNs and humans by measuring error consistency. In Advances in Neural Information Processing Systems, volume 33, pages 13890–13902. Curran Associates, Inc., Dec. 2020.
[20]
M. Dehghani, J. Djolonga, B. Mustafa, P. Padlewski, J. Heek, J. Gilmer, A. P. Steiner, M. Caron, R. Geirhos, I. Alabdulmohsin, R. Jenatton, L. Beyer, M. Tschannen, A. Arnab, X. Wang, C. R. Ruiz, M. Minderer, J. Puigcerver, U. Evci, M. Kumar, S. V. Steenkiste, G. F. Elsayed, A. Mahendran, F. Yu, A. Oliver, F. Huot, J. Bastings, M. Collier, A. A. Gritsenko, V. Birodkar, C. N. Vasconcelos, Y. Tay, T. Mensink, A. Kolesnikov, F. Pavetic, D. Tran, T. Kipf, M. Lucic, X. Zhai, D. Keysers, J. J. Harmsen, and N. Houlsby. Scaling Vision Transformers to 22 Billion Parameters. In Proceedings of the 40th International Conference on Machine Learning, volume 202 of Proceedings of Machine Learning Research, pages 7480–7512. PMLR, July 2023.
[21]
O. Siméoni, H. V. Vo, M. Seitzer, F. Baldassarre, M. Oquab, C. Jose, V. Khalidov, M. Szafraniec, S. Yi, M. Ramamonjisoa, F. Massa, D. Haziza, L. Wehrstedt, J. Wang, T. Darcet, T. Moutakanni, L. Sentana, C. Roberts, A. Vedaldi, J. Tolan, J. Brandt, C. Couprie, J. Mairal, H. Jégou, P. Labatut, and P. Bojanowski. , Aug. 2025.
[22]
Y.-H. Yang, T. Fukiage, Z. Sun, and S. Nishida. Psychophysical measurement of perceived motion flow of naturalistic scenes. iScience, 26 (12), Dec. 2023. ISSN 2589-0042. .
[23]
M. Tangemann, M. Kümmerer, and M. Bethge. Object segmentation from common fate: Motion energy processing enables human-like zero-shot generalization to random dot stimuli, Nov. 2024.
[24]
Z. Sun, Y.-J. Chen, Y.-H. Yang, Y. Li, and S. Nishida. Machine Learning Modeling for Multi-order Human Visual Motion Processing, Jan. 2025.
[25]
Y. Kubota and T. Fukiage. Accuracy Does Not Guarantee Human-Likeness in Monocular Depth Estimators, Dec. 2025.
[26]
Y. Kubota and T. Fukiage. Human-like monocular depth biases in deep neural networks. PLOS Computational Biology, 21 (8): e1013020, Aug. 2025. ISSN 1553-7358. .
[27]
G. Ehrensperger, S. Stabinger, and A. R. Sánchez. Evaluating CNNs on the Gestalt Principle of Closure. In I. V. Tetko, V. Kůrková, P. Karpov, and F. Theis, editors, Artificial Neural Networks and Machine LearningICANN 2019: Theoretical Neural Computation, pages 296–301, Cham, Sept. 2019. Springer International Publishing. ISBN 978-3-030-30487-4. .
[28]
B. Kim, E. Reif, M. Wattenberg, S. Bengio, and M. C. Mozer. Neural Networks Trained on Natural Scenes Exhibit Gestalt Closure. Computational Brain & Behavior, 4 (3): 251–263, Sept. 2021. ISSN 2522-087X. .
[29]
V. Biscione and J. S. Bowers. Mixed Evidence for Gestalt Grouping in Deep Neural Networks. Computational Brain & Behavior, 6 (3): 438–456, Sept. 2023. ISSN 2522-087X. .
[30]
L. Tang and D. Ley. Degraded Polygons Raise Fundamental Questions of Neural Network Perception. In A. Oh, T. Naumann, A. Globerson, K. Saenko, M. Hardt, and S. Levine, editors, Advances in Neural Information Processing Systems, volume 36 of Datasets and Benchmarks Track, pages 9695–9706. Curran Associates, Inc., Dec. 2023.
[31]
Y. Zhang, D. Soydaner, F. Behrad, L. Koßmann, and J. Wagemans. Investigating the Gestalt Principle of Closure in Deep Convolutional Neural Networks. In 32nd European Symposium on Artificial Neural Networks, Computational Intelligence and Machine Learning Bruges, Belgium October 09 - 11, pages 679–684. i6docs.com, Oct. 2024. ISBN 978-2-87587-090-2. .
[32]
Y. Zhang, D. Soydaner, L. Koßmann, F. Behrad, and J. Wagemans. Finding Closure: A Closer Look at the Gestalt Law of Closure in Convolutional Neural Networks. Computational Brain & Behavior, June 2025. ISSN 2522-087X. .
[33]
J. Sha, H. Shindo, K. Kersting, and D. S. Dhami. Gestalt Vision: A Dataset for Evaluating Gestalt Principles in Visual Perception. In 19th International Conference on Neurosymbolic Learning and Reasoning, Sept. 2025.
[34]
B. Lonnqvist, E. Scialom, A. Gokce, Z. Merchant, M. Herzog, and M. Schrimpf. Contour Integration Underlies Human-Like Vision. In Proceedings of the 42nd International Conference on Machine Learning, pages 40290–40311. PMLR, July 2025.
[35]
Y. Li, S. Salehi, L. Ungar, and K. P. Kording. Does Object Binding Naturally Emerge in Large Pretrained Vision Transformers?, Dec. 2025.
[36]
T. Chen, S. Kornblith, M. Norouzi, and G. Hinton. A Simple Framework for Contrastive Learning of Visual Representations. In H. D. III and A. Singh, editors, Proceedings of the 37th International Conference on Machine Learning, volume 119 of Proceedings of Machine Learning Research, pages 1597–1607. PMLR, July 2020.
[37]
J.-B. Grill, F. Strub, F. Altché, C. Tallec, P. Richemond, E. Buchatskaya, C. Doersch, B. Avila Pires, Z. Guo, M. Gheshlaghi Azar, B. Piot, k. kavukcuoglu, R. Munos, and M. Valko. Bootstrap Your Own Latent - A New Approach to Self-Supervised Learning. In H. Larochelle, M. Ranzato, R. Hadsell, M. Balcan, and H. Lin, editors, Advances in Neural Information Processing Systems, volume 33, pages 21271–21284. Curran Associates, Inc., Dec. 2020.
[38]
M. El Banani, A. Raj, K.-K. Maninis, A. Kar, Y. Li, M. Rubinstein, D. Sun, L. Guibas, J. Johnson, and V. Jampani. Probing the 3D Awareness of Visual Foundation Models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 21795–21806, June 2024.
[39]
X. Chen, M. Marks, and Z. Cheng. Probing the Mid-level Vision Capabilities of Self-Supervised Learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 30095–30105, June 2025.
[40]
D. Danier, M. Aygün, C. Li, H. Bilen, and O. Mac Aodha. : Evaluating Monocular Depth Perception in Large Vision Models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 20049–20059, June 2025.
[41]
O. Siméoni, G. Puy, H. V. Vo, S. Roburin, S. Gidaris, A. Bursuc, P. Pérez, R. Marlet, and J. Ponce. Localizing Objects with Self-Supervised Transformers and no Labels. In BMVC 2021, Sept. 2021. .
[42]
J. Zhou, C. Wei, H. Wang, W. Shen, C. Xie, A. Yuille, and T. Kong. : Image BERT Pre-Training with Online Tokenizer, Jan. 2022.
[43]
Y. Wang, X. Shen, S. X. Hu, Y. Yuan, J. L. Crowley, and D. Vaufreydaz. Self-Supervised Transformers for Unsupervised Object Discovery Using Normalized Cut. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 14543–14553, June 2022.
[44]
D. P. Kingma and J. Ba. Adam: A Method for Stochastic Optimization. In 3rd International Conference on Learning Representations (ICLR) 2015. arXiv, May 2015. .
[45]
M.-M. Cheng, N. J. Mitra, X. Huang, P. H. S. Torr, and S.-M. Hu. Global Contrast Based Salient Region Detection. IEEE Transactions on Pattern Analysis and Machine Intelligence, 37 (3): 569–582, Mar. 2015. ISSN 1939-3539. .
[46]
M. Cimpoi, S. Maji, I. Kokkinos, S. Mohamed, and A. Vedaldi. Describing Textures in the Wild. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 3606–3613, 2014.
[47]
S. Dziadzio, Ç. Yıldız, G. M. van de Ven, T. Trzciński, T. Tuytelaars, and M. Bethge. Infinite dSprites for Disentangled Continual Learning: Separating Memory Edits from Generalization. In 3rd Conference on Lifelong Learning Agents (CoLLAs), July 2024.
[48]
P. Bahnsen. Eine Untersuchungüber Symmetrie und Asymmetrie bei visuellen Wahrnehmungen. Zeitschrift für Psychologie, 108: 129–154, 1928.
[49]
W. Wang, H. Bao, L. Dong, J. Bjorck, Z. Peng, Q. Liu, K. Aggarwal, O. K. Mohammed, S. Singhal, S. Som, and F. Wei. Image as a Foreign Language: BEiT Pretraining for Vision and Vision-Language Tasks. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 19175–19186, June 2023.
[50]
R. Wightman. , 2019.
[51]
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei. : A large-scale hierarchical image database. In 2009 IEEE Conference on Computer Vision and Pattern Recognition, pages 248–255, June 2009. .
[52]
A. P. Steiner, A. Kolesnikov, X. Zhai, R. Wightman, J. Uszkoreit, and L. Beyer. How to train your ViT? Data, Augmentation, and Regularization in Vision Transformers. Transactions on Machine Learning Research, 2022. ISSN 2835-8856.
[53]
H. Touvron, M. Cord, and H. Jégou. : Revenge of the ViT. In S. Avidan, G. Brostow, M. Cissé, G. M. Farinella, and T. Hassner, editors, Computer VisionECCV 2022, pages 516–533, Cham, Oct. 2022. Springer Nature Switzerland. ISBN 978-3-031-20053-3. .
[54]
L. Beyer, P. Izmailov, A. Kolesnikov, M. Caron, S. Kornblith, X. Zhai, M. Minderer, M. Tschannen, I. Alabdulmohsin, and F. Pavetic. : One Model for All Patch Sizes. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 14496–14506, 2023.
[55]
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, G. Krueger, and I. Sutskever. Learning Transferable Visual Models From Natural Language Supervision. In M. Meila and T. Zhang, editors, Proceedings of the 38th International Conference on Machine Learning, volume 139 of Proceedings of Machine Learning Research, pages 8748–8763. PMLR, July 2021.
[56]
M. Tschannen, A. Gritsenko, X. Wang, M. F. Naeem, I. Alabdulmohsin, N. Parthasarathy, T. Evans, L. Beyer, Y. Xia, B. Mustafa, O. Hénaff, J. Harmsen, A. Steiner, and X. Zhai. 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features, Feb. 2025.
[57]
D. Bolya, P.-Y. Huang, P. Sun, J. H. Cho, A. Madotto, C. Wei, T. Ma, J. Zhi, J. Rajasegaran, H. Rasheed, J. Wang, M. Monteiro, H. Xu, S. Dong, N. Ravi, D. Li, P. Dollár, and C. Feichtenhofer. Perception Encoder: The best visual embeddings are not at the output of the network, Apr. 2025.
[58]
K. He, X. Chen, S. Xie, Y. Li, P. Dollár, and R. Girshick. Masked Autoencoders Are Scalable Vision Learners. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 16000–16009, June 2022.
[59]
H. Bao, L. Dong, S. Piao, and F. Wei. : BERT Pre-Training of Image Transformers. In The Tenth International Conference on Learning Representations (ICLR) 2022, Apr. 2022.
[60]
Y. Fang, Q. Sun, X. Wang, T. Huang, X. Wang, and Y. Cao. : A visual representation for neon genesis. Image and Vision Computing, 149: 105171, Sept. 2024. ISSN 0262-8856. .
[61]
M. Oquab, T. Darcet, T. Moutakanni, H. Vo, M. Szafraniec, V. Khalidov, P. Fernandez, D. Haziza, F. Massa, A. El-Nouby, M. Assran, N. Ballas, W. Galuba, R. Howes, P.-Y. Huang, S.-W. Li, I. Misra, M. Rabbat, V. Sharma, G. Synnaeve, H. Xu, H. Jegou, J. Mairal, P. Labatut, A. Joulin, and P. Bojanowski. : Learning Robust Visual Features without Supervision, Feb. 2024.
[62]
T. Darcet, M. Oquab, J. Mairal, and P. Bojanowski. Vision Transformers Need Registers, Apr. 2024.