July 15, 2026
Multimodal large language models (MLLMs) have demonstrated remarkable performance across a wide range of 2D vision-language tasks [1]–[3]. However, a critical capability remains underdeveloped: 3D spatial reasoning. Humans naturally infer depth, relative positions, and occlusions from visual scenes, yet existing MLLMs still struggle with complex spatial reasoning tasks [4], [5]. This deficiency in such 3D understanding impedes their deployment in geometry-sensitive downstream applications, including robotics [6], [7], autonomous driving [8]–[10], and VR/AR systems [11]. Existing efforts to enhance spatial reasoning capabilities generally follow two paradigms. The first paradigm leverages large-scale datasets, enabling models to implicitly learn 3D relational patterns through statistical regularities in the data [4], [12]–[15]. The second paradigm augments 2D visual inputs with explicit 3D or 2.5D signals (e.g., point clouds, depth maps, or cognitive maps), thereby providing richer global geometric structure for the model [16]–[20]. Despite their contributions, both methods are constrained by reasoning within a discrete text space: geometric evidence is verbalized into discrete tokens, inevitably leading to the loss of fine-grained spatial cues and biasing reasoning toward linguistic priors. Furthermore, the discrete nature of text tokens fails to capture the continuous characteristics of the physical world, rendering autoregressive generation unreliable for precise numerical tasks like distance estimation and object localization [21].
Recent studies have explored latent reasoning to address the inherent limitations of text-based reasoning. By conducting intermediate reasoning within non-linguistic latent representations, rather than verbalizing every step in natural language, these methods preserve continuous semantic information, facilitating more coherent multimodal reasoning [22]–[25]. When applied to spatial reasoning, researchers typically design geometrically-aware latent tokens supervised by geometric priors (e.g., depth maps) [26], [27]. Nevertheless, directly adopting such latent tokens for 3D spatial reasoning introduces two key limitations. First, existing methods rely on a single type of latent representation, which struggles to accommodate the diverse demands of spatial reasoning. For example, while latents decoded into depth maps effectively capture near-far object relationships, they are limited in representing complex spatial information beyond depth. Second, these latent representations primarily focus on global geometric context, offering limited insight into local object interactions. Consequently, when tasked with precise tasks such as estimating the absolute distance between two points, global latents fail to provide sufficient local evidence.
In this work, we propose GeoAnchor, a text–latent interleaved framework that avoids compressing geometric information into a single entangled latent representation or verbalizing it into discrete text tokens. Instead, GeoAnchor decomposes 3D spatial information into three complementary basic latent components: a position token for precise object grounding, a direction token for modeling relational orientation, and a geometry token for capturing global scene structure. The position and direction tokens provide explicit and traceable local evidence for target objects, while the geometry latent encodes global scene context. This design enables dynamic recombination of these latents within a structured space, eliminating the need for a uniform latent token design across all types of questions. Consequently, GeoAnchor achieves an interpretable latent reasoning process that adaptively accommodates the diverse demands of spatial reasoning tasks.
To enable such structured reasoning, we introduce a four-stage collaborative training strategy. First, we align the model with local 3D perception using large-scale object-grounding data, thereby establishing robust object-level spatial cues. Second, we jointly optimize local and global latent components, facilitating the evolution of reasoning from local perception to holistic spatial understanding. Third, we refine the latent representations via text-only supervision, encouraging the model to internalize spatial reasoning logic without relying on intermediate guidance. Finally, we incorporate reinforcement learning [28] with pattern-specific rewards to enable adaptive selection of latent tokens, allowing the model to dynamically switch between local-only and local-plus-global reasoning patterns. Built on Qwen3-VL-2B [1], GeoAnchor achieves state-of-the-art performance: 68.4% and 69.7% accuracy on the SPAR-Bench [29] and SPBench [13], respectively. It surpasses the base model by a significant 21.2% margin, and outperforms GPT-4o [30] and Gemini-2.5-Flash [31] by 18.0% and 13.6%, respectively. Moreover, a 10.7% performance gain on the out-of-domain ViewSpatial Bench [32] confirms GeoAnchor’s robust generalization capability.
The main contributions of this paper are summarized as follows:
We propose GeoAnchor, a new 3D spatial reasoning framework that decomposes geometric reasoning into three complementary and physically meaningful latent components: position, direction, and geometry. Such a design enables dynamic latent recombination toward interpretable and query-adaptive spatial understanding.
We introduce a collaborative four-stage training strategy that enables the model to perform structured local-plus-global spatial reasoning and adaptively select latent tokens.
Extensive experiments on challenging 3D reasoning benchmarks, including SPAR-Bench, SPBench and ViewSpatial, demonstrate that GeoAnchor surpasses the base model by 21.2% and outperforms GPT-4o and Gemini-2.5-Flash, showing competitive performance in 3D reasoning tasks.
With rapid progress in embodied intelligence, spatial understanding in MLLMs has drawn increasing attention. Recent efforts to improve spatial reasoning in MLLMs can be broadly grouped into two categories. The first leverages curated datasets and multi-stage training to learn spatial relations from large-scale examples [12]–[15], [33]–[37]. For example, SpatialLadder [13] constructs SpatialLadder-26k and adopts a progressive curriculum from perception to reasoning. SpaceR [37] further improves spatial reasoning by introducing spatially tailored reward signals in reinforcement learning. However, direct training often promotes memorization over reasoning, limiting generalization.
Recognizing the limitations of purely 2D priors, the second category augments the input with 2.5D or 3D information to strengthen global spatial understanding [16]–[19], [38], [39]. For example, SpatialMind [38] encodes object locations on a 2D grid to form a cognitive map that is provided as additional input. Spatial-MLLM [19] incorporates a frozen geometry encoder (e.g. VGGT) to provide complementary 3D representations. Nevertheless, methods that primarily generate text still struggle to faithfully express continuous spatial relations in discrete language. This mismatch causes substantial information loss when continuous geometry is mapped to textual descriptions, ultimately limiting spatial understanding.
To overcome the limitations of a text-only output modality, one straightforward solution is to “think with images” [40]–[43]. For example, ViLaSR [42] facilitates spatial localization by rendering auxiliary guide lines, while MindJourney [40] leverages a diffusion model to synthesize alternative viewpoints and feeds the generated images back into the model. However, these methods rely heavily on external tools, reducing flexibility and limiting generalization.
Inspired by latent reasoning in LLMs [44]–[46], which replaces discrete text tokens with self-generated continuous embeddings, several studies have extended latent reasoning to MLLMs [22]–[25], [47] and further to spatial reasoning [26], [27], [48], [49] to improve reasoning flexibility. LVR [24] and Mirage [23] use image supervision to guide latent tokens, encouraging the model to develop an interleaved visual–textual reasoning process in latent space. For spatial reasoning, Aurora [26] and SSR [27] incorporate depth maps into latent tokens to strengthen geometric reconstruction. Nevertheless, the complexity of spatial reasoning is unlikely to be fully captured by a single global latent token. Such a bottleneck often lacks sufficient capacity and structure for diverse spatial problems, while also reducing interpretability when the model must resolve heterogeneous geometric relations.
In this section, we present GeoAnchor, a framework designed to tackle complex spatial reasoning tasks via an interleaved text-latent paradigm. Section 3.1 presents the overall architecture, while Section 3.2 details the design of decomposed spatial tokens. Section 3.3 then describes the collaborative training strategy, which facilitates a structured evolution from local perception to holistic spatial understanding. Finally, Section 3.4 summarizes the construction of the grounding and reasoning datasets employed across training stages.
GeoAnchor employs a text-latent interleaved reasoning framework to address the loss of continuous 3D information inherent in relying solely on discrete textual representations. Specifically, the model utilizes discrete text tokens for semantic planning and logical reasoning, while leveraging latent representations to encode and capture over continuous, multi-grained 3D spatial cues. The overall structure is illustrated in Figure 2. Given an image \(\mathcal{I}\) and query \(Q\), GeoAnchor outputs the reasoning trajectory \(\mathcal{O}\), which consists of an interleaved sequence of text and latent tokens:
\[\mathcal{O} = t_1 \oplus z_1 \oplus \cdots \oplus z_{k-1} \oplus t_k,\] where \(\mathbf{t} = \{t_1, \ldots, t_k\}\) denotes the sequence of text tokens, and \(\mathbf{z} = \{z_1, \ldots, z_{k-1}\}\) represents the sequence of latent tokens.
Unlike text tokens retrieved via embedding lookups, each latent token \(z_i\) is defined as a fixed-length sequence of \(M\) continuous hidden states, i.e., \(z_i = \{h_{i,1}, h_{i,2}, \dots, h_{i,M}\}\), where each \(h_{i,j} \in \mathbb{R}^d\) resides in the model’s hidden space. Specifically, at the \(j\)-th internal step of generating \(z_i\), the MLLM \(f_\theta(\cdot)\) produces the hidden state \(h_{i,j}\) conditioned on the preceding context: \[h_{i,j} = f_{\theta}^{\mathrm{hidden}}(Q, \mathcal{I}, \mathcal{O}_{<z_i}, h_{i,1:j-1}),\] where \(\mathcal{O}_{<z_i}\) denotes the entire reasoning trajectory generated before \(z_i\), and \(h_{i,1:j-1}\) denotes the internal hidden states generated up to step \(j-1\). However, the output hidden-state space and the input embedding space typically lie on different manifolds. Directly feeding \(h_{i,j}\) back into the input trajectory without alignment can lead to severe latent drift and autoregressive instability [50]. To bridge this gap, we introduce a latent projector \(\mathcal{P}(\cdot)\) that maps the output hidden state into the input embedding space. Specifically, the projected continuous embedding \(e_{i,j} \in \mathbb{R}^d\) is computed as: \[e_{i,j} = \mathcal{P}(h_{i,j}) = \text{LayerNorm}\Big(h_{i,j} + \text{MLP}\big(\text{LayerNorm}(h_{i,j})\big)\Big).\] The projected embedding \(e_{i,j}\) is then fed back as the input representation for the subsequent generation step, ensuring a consistent and stable reasoning trajectory.
To address the limitation that a single latent type struggles to handle diverse spatial reasoning tasks, we decompose the spatial reasoning process into two distinct stages: obtaining local atomic 3D cues and building global geometric context. Accordingly, we introduce three complementary latent tokens: the position token \(z^{\mathrm{pos}}\), the direction token \(z^{\mathrm{dir}}\), and the geometry token \(z^{\mathrm{geo}}\). The first two serve as local tokens by encoding atomic 3D cues for target objects, including object grounding and relational orientation, whereas the geometry token is a global token to capture the overall scene structure. By disentangling the uniform latent representation into structured 3D evidence with distinct functions, the model can flexibly compose diverse spatial evidence within the latent space, thereby enabling more interpretable spatial reasoning.
Explicit Supervision of Local 3D Cues. In conventional text-based reasoning, local spatial information, such as positional coordinates and directional vectors, is typically represented as discrete numerical text tokens. To preserve the continuity of 3D spatial information, we encode these cues into latent tokens and employ decoders to map them into numerical coordinates for alignment. We map latent tokens to the target 3D space using lightweight linear heads, similar to the text prediction module. Given that each latent token comprises a sequence of hidden states (Section 3.1), \(z^{\mathrm{pos}}\) and \(z^{\mathrm{dir}}\) are defined as: \[z^{\mathrm{pos}}=\{\mathbf{h}^{\mathrm{pos}}_1,\ldots,\mathbf{h}^{\mathrm{pos}}_{l_{\mathrm{pos}}}\},\qquad z^{\mathrm{dir}}=\{\mathbf{h}^{\mathrm{dir}}_1,\ldots,\mathbf{h}^{\mathrm{dir}}_{l_{\mathrm{dir}}}\},\] where \(\mathbf{h}^{\mathrm{pos}}_j\) and \(\mathbf{h}^{\mathrm{dir}}_j\in\mathbb{R}^D\) are the hidden states at the \(j\)-th internal step, and \(l_{\mathrm{pos}}\) and \(l_{\mathrm{dir}}\) denote the numbers of hidden states for each latent token. The hidden states within each latent token are first averaged along the sequence dimension and then mapped to 3D outputs through two separate linear heads: \[\hat{\mathbf{p}}=W_{\mathrm{pos}}\frac{1}{l_{\mathrm{pos}}}\sum_{j=1}^{l_{\mathrm{pos}}}\mathbf{h}^{\mathrm{pos}}_j, \quad \hat{\mathbf{d}}=W_{\mathrm{dir}}\frac{1}{l_{\mathrm{dir}}}\sum_{j=1}^{l_{\mathrm{dir}}}\mathbf{h}^{\mathrm{dir}}_j,\] where \(W_{\mathrm{pos}}, W_{\mathrm{dir}} \in \mathbb{R}^{3 \times D}\) are projection heads, while \(\hat{\mathbf{p}} \in \mathbb{R}^3\) and \(\hat{\mathbf{d}} \in \mathbb{R}^3\) denote the predicted position and direction, respectively. To explicitly supervise these two tokens, we apply the Smooth L1 loss [51] for position prediction and the cosine similarity loss for direction prediction: \[\mathcal{L}_{\mathrm{pos}} = \begin{cases} \frac{1}{2}(\hat{\mathbf{p}} - \mathbf{p})^2, & \text{if } |\hat{\mathbf{p}} - \mathbf{p}| < 1.0 \\ |\hat{\mathbf{p}} - \mathbf{p}| - \frac{1}{2}, & \text{otherwise} \end{cases}, \mathcal{L}_{\mathrm{dir}} = 1 - \frac{\hat{\mathbf{d}} \cdot \mathbf{d}}{\|\hat{\mathbf{d}}\|_2 \|\mathbf{d}\|_2},\] where \(\mathbf{p}\) and \(\mathbf{d}\) denote the ground-truth 3D position coordinates and direction vectors, respectively.
Coverage-Based Supervision of Global Geometry. We supervise the global token \(z^{\mathrm{geo}}\) using geometry features from the last layer of the VGGT model [52], since these features can be decoded into 3D point clouds and inherently capture global geometric structure. However, VGGT generates a high-resolution feature grid, whereas the global token is a compact representation composed of a few hidden states. Consequently, enforcing strict token-wise alignment would impose overly fine-grained constraints on the global token, potentially undermining its role as a compact representation of scene-level geometry. Therefore, a coarse soft-coverage alignment strategy is adopted instead of strict dense alignment.
We first transform the dense VGGT feature map into a supervision feature by multi-scale average pooling. Specifically, for the final VGGT feature map of size \(H\times W\times D\), average pooling is applied at multiple spatial resolutions \(\{r_1,\dots,r_L\}\). Each pooled feature grid is flattened into \(r_l^2\) feature vectors, and the vectors from all scales are concatenated to form the coarse-grained VGGT feature map \(V \in \mathbb{R}^{l_{\mathrm{vggt}} \times D}\), where \(l_{\mathrm{vggt}} = \sum_{l=1}^{L} r_l^2\). In parallel, the global token \(z^{\mathrm{geo}}\) is projected into geometry tokens \(G \in \mathbb{R}^{l_{\mathrm{geo}} \times D}\) via a linear layer, where \(l_{\mathrm{geo}}\) denotes the number of geometry tokens. We then compute token similarities as: \[A_{i,j}=\frac{\langle \mathbf{v}_{i}, \mathbf{g}_{j}\rangle}{\tau}, \qquad u_j=\frac{1}{l_{\mathrm{vggt}}}\sum_{i=1}^{l_{\mathrm{vggt}}}\frac{\exp(A_{i,j})}{\sum_{j'=1}^{l_{\mathrm{geo}}}\exp(A_{i,j'})},\] where \(\mathbf{v}_i \in \mathbb{R}^D\) is the \(i\)-th coarse-grained VGGT feature and \(\mathbf{g}_j \in \mathbb{R}^D\) is the \(j\)-th geometry token, \(\tau\) represents the temperature parameter, and \(u_j\) denotes the average soft assignment received by the \(j\)-th geometry token. The alignment loss is defined as: \[\mathcal{L}_{\mathrm{geo}} = -\frac{1}{l_{\mathrm{vggt}}}\sum_{i=1}^{l_{\mathrm{vggt}}}\log\sum_{j=1}^{l_{\mathrm{geo}}}\exp(A_{i,j}) + \lambda_{\mathrm{bal}} \sum_{j=1}^{l_{\mathrm{geo}}}u_j\left(\log u_j-\log\frac{1}{l_{\mathrm{geo}}}\right).\] The first term encourages that every VGGT feature aligns with at least one geometry token, thereby achieving coarse coverage without dense alignment. The second term enforces balanced token utilization to prevent representation collapse, controlled by the coefficient \(\lambda_{\mathrm{bal}}\). Collectively, these objectives guide the geometry tokens to capture complementary global evidence that integrates seamlessly with local evidence for downstream reasoning.
To effectively train the model for spatial reasoning with latent tokens, we design a collaborative training framework that systematically builds up spatial intelligence, as illustrated in Figure 3.
Stage 1: Local Perception Warm-up. We first warm up the base model by training its 3D grounding capability. Specifically, we construct a large-scale 3D grounding dataset and train the model with two complementary abilities: predicting the 3D position of a target object and predicting the relative direction between two objects. By learning these two signals jointly, the model is encouraged to collaboratively organize the local tokens into coherent local spatial evidence. The local tokens are optimized with \(\mathcal{L}_{\text{pos}}\) and \(\mathcal{L}_{\text{dir}}\), while the text tokens are trained with the next-token prediction loss \(\mathcal{L}_{\text{NTP}}\). Formally, the training objective for Stage 1 is: \[\mathcal{L}_{\mathrm{stage1}} = \lambda_t \mathcal{L}_{\text{NTP}} + \lambda_l \left( \mathcal{L}_{\mathrm{pos}} + \mathcal{L}_{\mathrm{dir}} \right),\] where \(\lambda_t\) and \(\lambda_l\) are the corresponding balance coefficients.
Stage 2: Spatial Latent Reasoning. In the second stage, the model is further trained on a spatial reasoning dataset covering diverse spatial tasks. During this stage, it learns to conduct an interpretable spatial reasoning process in the latent space by collaboratively integrating local 3D cues with global geometric context. It also learns to dynamically utilize the position and the direction tokens according to the input question, thereby developing a basic ability for adaptive latent token selection. The position, direction, and geometry tokens are supervised by their corresponding losses. Denoting \(\lambda_g\) as the coefficient for \(\mathcal{L}_{\mathrm{geo}}\), the total loss for Stage 2 is defined as follows: \[\mathcal{L}_{\mathrm{stage2}}=\lambda_t\mathcal{L}_{\text{NTP}} + \lambda_l(\mathcal{L}_{\mathrm{pos}} + \mathcal{L}_{\mathrm{dir}}) + \lambda_g\mathcal{L}_{\mathrm{geo}}.\]
Stage 3: Latent Relaxation. Although the explicit supervision in Stage 2 effectively aligns the latent tokens with 3D information, overly strong alignment may shift them away from the original language manifold and hinder subsequent reasoning. Inspired by [23], [47], we therefore introduce a latent relaxation stage that removes explicit latent supervision and retains only text supervision. \(\lambda_{l}\) and \(\lambda_{g}\) are set to 0, and the latent tokens are thus updated only through the text supervision, which helps the learned spatial representations better fit the language distribution while preserving the spatial information acquired in earlier stages.
Stage 4: Adaptive Latent Reinforcement Learning. Although the preceding stages enable reasoning with latent tokens, they still rely on a fixed usage pattern where both local and global tokens are invoked for every query. However, different spatial reasoning tasks require different levels of spatial information, making such fixed invocation suboptimal. To address this issue, we introduce an adaptive latent reinforcement learning stage based on Group Relative Policy Optimization (GRPO) [28]. We consider two reasoning patterns: local-only and local-plus-global. Then in addition to format and accuracy rewards, we introduce a pattern-specific reward to evaluate the effectiveness of each reasoning pattern. Within each group, the model samples \(N\) responses, each generated by randomly selecting one of the two patterns. For pattern \(t \in \{\text{local}, \text{local+global}\}\), its effectiveness is estimated as: \[Acc_t = \frac{n_t^{\mathrm{correct}} + \kappa Acc_t^{\mathrm{hist}}}{n_t + \kappa}, Acc_t^{\mathrm{hist}} \leftarrow (1-\mu)\, Acc_t^{\mathrm{hist}} + \mu \frac{n_t^{\mathrm{correct}}}{n_t},\] where \(n_t\) and \(n_t^{\mathrm{correct}}\) denote the numbers of sampled and correct responses under pattern \(t\), respectively, \(Acc_t^{\mathrm{hist}}\) is the EMA-based historical accuracy [53], \(\kappa\) is the smoothing coefficient, and \(\mu\) is the update rate. The pattern with the higher smoothed accuracy receives an additional pattern reward \(r_{\text{pattern}}\). This encourages the model to use global token only when local reasoning is insufficient, leading to more efficient and more interpretable reasoning.
| SPAR-Bench | SPBench | ViewSpatial | |||||||||||||||||||||||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 3-8(lr)9-11(lr)12-14 Model | Param | Avg. | Dep. | Dis. | Prox. | Rel. | View | Avg. | Rel. | Abs. | Avg. | Cam. | Per. | ||||||||||||||||||||||||
| Proprietary Models | |||||||||||||||||||||||||||||||||||||
| GPT-4o [30] | - | 40.1 | 34.3 | 43.4 | 57.1 | 45.9 | 31.0 | 53.4 | 49.4 | 56.0 | 37.5 | 33.5 | 43.6 | ||||||||||||||||||||||||
| Gemini-2.5-Flash [31] | - | 37.4 | 45.2 | 63.2 | 51.5 | 44.7 | 56.0 | 44.0 | 43.0 | ||||||||||||||||||||||||||||
| General Models | |||||||||||||||||||||||||||||||||||||
| MiniCPM-V-4.5 [54] | 8B | 37.7 | 32.4 | 32.6 | 62.1 | 50.6 | 29.8 | 40.9 | 47.4 | 36.7 | 39.0 | 43.0 | 32.9 | ||||||||||||||||||||||||
| LLaVA-OneVision-1.5 [55] | 8B | 34.5 | 29.3 | 32.6 | 58.5 | 40.7 | 26.7 | 40.8 | 45.1 | 38.0 | 38.7 | 40.6 | 35.3 | ||||||||||||||||||||||||
| Qwen3-VL [1] | 8B | 41.1 | 32.5 | 34.6 | 71.2 | 62.1 | 31.5 | 55.1 | 47.6 | 44.0 | 44.6 | 43.1 | |||||||||||||||||||||||||
| Molmo2 [56] | 8B | 24.6 | 16.7 | 3.2 | 62.4 | 48.4 | 25.1 | 29.7 | 47.6 | 18.1 | 45.5 | 46.0 | 45.1 | ||||||||||||||||||||||||
| GLM-4.1V [57] | 9B | 48.0 | 53.2 | 65.6 | 57.7 | 34.3 | 49.1 | 43.6 | 52.6 | 40.4 | 37.6 | 44.7 | |||||||||||||||||||||||||
| Qwen3.5 [58] | 9B | 48.6 | 39.0 | 53.0 | 73.8 | 33.4 | 48.8 | 48.6 | 48.8 | 43.2 | |||||||||||||||||||||||||||
| Kimi-VL [59] | 16B-A3B | 33.9 | 28.6 | 26.7 | 60.6 | 44.5 | 28.3 | 42.1 | 40.8 | 43.0 | 38.1 | 36.5 | 40.4 | ||||||||||||||||||||||||
| Internvl3.5 [60] | 38B | 38.0 | 31.5 | 35.5 | 60.9 | 58.0 | 25.5 | 53.1 | 49.4 | 55.5 | 41.4 | 43.3 | 38.6 | ||||||||||||||||||||||||
| Specialized Models | |||||||||||||||||||||||||||||||||||||
| 2B | 41.7 | 41.8 | 37.7 | 26.4 | 24.2 | 14.3 | 25.2 | 7.2 | 18.3 | 16.0 | 21.7 | ||||||||||||||||||||||||||
| 3B | 23.8 | 15.8 | 18.6 | 49.7 | 24.2 | 25.1 | 13.8 | 32.2 | 1.7 | 29.2 | 27.7 | 31.5 | |||||||||||||||||||||||||
| 3B | 32.4 | 22.1 | 29.8 | 53.8 | 43.7 | 29.5 | 43.3 | 41.8 | 45.4 | ||||||||||||||||||||||||||||
| 4B | 33.7 | 25.3 | 34.2 | 53.8 | 36.5 | 31.2 | 54.2 | 55.7 | 53.3 | 39.7 | 39.8 | 39.6 | |||||||||||||||||||||||||
| 7B | 38.6 | 28.9 | 43.1 | 61.8 | 46.4 | 28.2 | 52.2 | 54.4 | 50.8 | 34.3 | 33.8 | 34.9 | |||||||||||||||||||||||||
| 13B | 29.6 | 26.9 | 29.0 | 52.7 | 21.7 | 25.7 | 32.1 | 36.8 | 29.0 | 25.6 | 26.2 | 24.7 | |||||||||||||||||||||||||
| Qwen3-VL [1] (Baseline Model) | 2B | 32.4 | 21.9 | 24.0 | 60.6 | 49.2 | 28.3 | 52.9 | 51.4 | 54.8 | 36.3 | 39.5 | 37.2 | ||||||||||||||||||||||||
| GeoAnchor | 2B | 58.7 | |||||||||||||||||||||||||||||||||||
| Improvement | - | +36.0 | +33.6 | +43.4 | +19.7 | +35.4 | +40.5 | +16.8 | +35.3 | +3.9 | +10.7 | +8.0 | +8.8 | ||||||||||||||||||||||||
We construct two training sets for GeoAnchor: a large-scale 3D grounding dataset for Stage 1 and a spatial reasoning dataset for the subsequent supervised and reinforcement learning stages. The former provides explicit supervision for local spatial perception, while the latter supports joint local-global spatial reasoning.
For the grounding dataset, we build on top of ScanNet [61] by sampling 10k scenes from its 2.5 million views over 1,500+ scans. For each image, we use Qwen3-VL-32B [1] to identify up to five objects, along with their descriptions and 2D bounding boxes. We then employ Depth Anything v3 [62] to estimate depth maps and camera poses, and recover the 3D coordinates of the extracted objects through back-projection. Ground-truth directions are derived from coordinate differences. We further diversify the queries using point-coordinate, bounding-box, and natural-language formulations, with each sample involving one to three objects. This process yields 550k 3D grounding samples in total.
For the spatial reasoning dataset, we use SPAR [29] as the primary data source, since it is built on ScanNet [61], ScanNet++ [63], and Structured3D [64], and provides 7M samples covering diverse spatial tasks. We sample 100k questions from SPAR and use the bounding boxes provided in the dataset to obtain the ground-truth positions and directions of the target objects, following the same back-projection procedure as in the grounding dataset. In addition, we extract global geometric features for each image using VGGT [52]. To better cover centimeter-based answer formats, we further incorporate 5k samples from SpatialLadder-26k [13]. In total, the resulting spatial reasoning dataset contains 105k samples.
Implementation Details. GeoAnchor is built on Qwen3-VL-2B [1], with local token length \(l_{\mathrm{pos}}=l_{\mathrm{dir}}=2\) and global token length \(l_{\mathrm{geo}}=8\). Stage 1 is trained on the grounding dataset for one epoch using a batch size of 64 and a learning rate of \(1\times10^{-4}\). Stages 2 and 3 are each trained for one epoch on the spatial reasoning dataset with a batch size of 32 and a learning rate of \(2\times10^{-5}\). The loss weights are set to \(\lambda_t=\lambda_l=1\) and \(\lambda_g=0.1\). For global geometry supervision, the VGGT feature map is pooled at \(L=3\), with \(\{r_1,r_2,r_3\}=\{1,2,4\}\). In the RL stage, 8 rollouts are performed per question with a sampling temperature of 1.0, a pattern reward \(r_{\text{pattern}}\) of 0.5, a KL-divergence coefficient \(\beta\) of 0.01, and a learning rate of \(5\times10^{-7}\). Additional details are provided in the supplementary material.
Evaluation Benchmarks and Metrics. We evaluate GeoAnchor on SPAR-Bench [29], SPBench [13], and ViewSpatial [32], where the first two serve as in-domain benchmarks and the last assesses out-of-domain generalization. For multiple-choice questions, we directly judge correctness and report the mean accuracy. For numerical questions, the mean accuracy is computed as the average accuracy across confidence thresholds from 0.5 to 0.9 with a step size of 0.05.
Table 1 compares GeoAnchor with state-of-the-art models across three benchmarks. Despite its compact 2B parameter size, GeoAnchor outperforms proprietary, open-source, and specialized MLLMs, including those with up to 38B parameters. On in-domain benchmarks, it achieves state-of-the-art mean accuracy on both SPAR-Bench (68.4%) and SPBench (69.7%), surpassing the base Qwen3-VL-2B by 36.0% and 16.8%, respectively. While SpatialLadder slightly outperforms GeoAnchor on the SPBench Abs. metric, its inconsistent performance across other benchmarks suggests a limitation in generalization. Notably, on the out-of-domain ViewSpatial benchmark, GeoAnchor attains 47.0% mean accuracy, ranking first in both sub-tasks. The substantial 10.7% gain over the base model demonstrates our method’s robust generalization capability beyond the training distribution.
Figure 4:
.
Effects of Different Reasoning Paradigms. Table 4 compares the performance of different reasoning paradigms on spatial tasks. Compared with vanilla SFT, text CoT SFT does not yield consistent gains across benchmarks, suggesting that spatial information is difficult to incorporate effectively through textual reasoning. Therefore, the reasoning process is more easily influenced by language patterns than by geometric evidence. In contrast, by encoding continuous 3D information into latent tokens, GeoAnchor outperforms both vanilla SFT and text CoT SFT, indicating that latent reasoning provides a more effective way to leverage spatial information. Furthermore, GeoAnchor also surpasses the single-latent reasoning baseline. Performance drops when the model is restricted to one type of analysis, showing that the decomposed token design plays a more important role than latent reasoning. Our design reduces the semantic burden of a single latent and enables more task-adaptive use of spatial cues, leading to more interpretable and robust reasoning than a single-latent design.
Effects of Collaborative Training Stages. Table [tab:unified95ablation] illustrates the effectiveness of each training stage. Stages 1, 2, and 3 improve the performance by 5.0%, 14.6%, and 23.0%, respectively, validating the effectiveness of our multi-stage design. Figure 5 shows that Stage 3 yields significantly higher gains than merely extending Stage 2, confirming its value lies in integrating latent tokens into reasoning rather than additional training iterations. Moreover, while applying GRPO directly as stage 4 improves in-domain performance, it degrades out-of-domain results on ViewSpatial. After adding the pattern reward, the model achieves consistent gains across all three benchmarks. This suggests that applying a fixed reasoning pattern introduces redundant information and thereby interferes with answer prediction. In contrast, the pattern reward encourages a more concise and efficient reasoning process, enabling the model to make more accurate use of spatial evidence.
Effects of Local Position and Direction Tokens. Table ¿tbl:tab:local95token95interpretability? groups the three benchmarks into three categories: Position, which focuses on relative object locations; Direction, which focuses on object orientation; and Mixed, which requires both types of local evidence. The results show that using both position and direction tokens yields the best overall performance. When only the position token is used during reasoning, the model performs well on position-related questions but shows clear degradation on direction-related and mixed categories. Using only the direction token leads to the opposite trend. These observations indicate that the local tokens indeed encode their corresponding spatial information and provide strong evidence for final answer generation.
Figure 6: Comparison of the top text attention weights assigned by the final answer under different reasoning paradigms.. a — Text CoT, b — GeoAnchor w/o Stage 1, c — GeoAnchor
Figure 7:
.
Effects of Alignment Strategies in Global Tokens. Table [tbl:tab:vggt95ablation] compares different VGGT alignment strategies. The proposed soft coverage alignment is evaluated against three dense alignment baselines. Specifically, Mean Pooling averages the VGGT features and the global token along the sequence dimension. Adaptive Pooling and Linear Interpolation resample the VGGT features to match the length of the global token using average pooling and interpolation, respectively. The results show that soft coverage alignment achieves the best performance across all three benchmarks, suggesting that coverage-based supervision provides a more suitable strategy for learning a compact representation of scene-level geometry while preserving sufficient information for downstream reasoning.
Latent reasoning improves aggregation of answer-relevant information, with Stage 1 playing a key role. As shown in Figure 6, GeoAnchor assigns stronger attention to latent tokens when generating the final answer. In contrast, although Text CoT also produces numerical coordinates, its answer token attends less to spatially meaningful cues and more to semantically weak symbols. This suggests that discretized textual reasoning weakens spatial semantics, whereas GeoAnchor allows the answer to directly leverage compact latent representations. The attention analysis further highlights that Stage 1 is important for developing a strong dependence on latent tokens, indicating that this stage establishes robust grounded latent evidence that supports subsequent reasoning.
Final-answer attention is more accurately grounded in task-relevant visual regions. As shown in the upper part of Figure 7, GeoAnchor focuses more precisely on the visual regions relevant to the queried spatial target when generating the final answer. By comparison, the baseline exhibits more diffuse attention, while Text CoT spreads attention more broadly without clearly isolating the most relevant object or location. This suggests that GeoAnchor not only aggregates reasoning information more effectively, but also grounds its predictions more faithfully in visual evidence.
Local tokens show clear spatial specialization. The lower part of Figure 7 shows that GeoAnchor produces distinct visual attention patterns when generating latent tokens for different positions. Each local token consistently attends to the region associated with its corresponding spatial target, indicating that these tokens encode disentangled spatial semantics rather than entangled global patterns. This behavior suggests that the latent space is organized in a semantically meaningful way, making the intermediate reasoning variables both interpretable and useful for final answer prediction.
The global token indeed encodes VGGT information. Following [23], we use t-SNE [65] to visualize the relationships among text, image, VGGT features, and the global token. As shown in Figure [fig:tsne95visual], on both SPAR-Bench and SPBench, the global token lies between the VGGT and text tokens. This suggests that it remains close to the geometry manifold while preserving textual semantics, consistent with the role of Stage 3 discussed in Section 3.3.
This paper presents GeoAnchor, a text-latent interleaved framework for spatial reasoning. By decomposing spatial reasoning into position, direction, and geometry latents, GeoAnchor enables decomposed local-plus-global reasoning beyond text-only representations. A collaborative multi-stage training strategy further improves the learning of grounded spatial evidence. Experiments on multiple benchmarks show that GeoAnchor consistently outperforms the base model and generalizes well across diverse 3D reasoning tasks.
This work is supported by the National Natural Science Foundation of China (Grant Nos. 92270201 and 62125204), and National Natural Science Foundation of China (62522102, 62432001), Beijing Natural Science Foundation (L247006).
Our training data mainly come from SPAR-7M and SpatialLadder-26K. In the experiments, we use SPAR-Bench, SPBench, and ViewSpatial-Bench as the main benchmarks to evaluate our method. In this section, we first introduce the datasets and benchmarks used in our study, and then summarize the task definitions and sample counts of the evaluated benchmark subsets in the corresponding tables.
SPAR-7M is a large-scale dataset designed for spatial understanding, constructed from more than 4,000 indoor 3D scenes collected from public scene datasets with 3D ground truth, including ScanNet, ScanNet++, and Structured3D. Through a 3D-driven data generation pipeline, the authors derive image sequences, camera parameters, and depth information from these scenes, and further generate diverse spatial question-answer pairs from the associated 3D scene metadata. As a result, SPAR-7M covers 33 task types and contains over 7 million QA pairs, spanning a broad range of abilities from low-level spatial perception to high-level spatial reasoning, and supporting single-view, multi-view, and video settings.
Importantly, the dataset is accompanied by its own benchmark, namely SPAR-Bench, which is explicitly introduced for evaluation rather than training alone. SPAR-Bench is constructed from the SPAR-7M split by selecting representative spatial tasks and manually verifying the resulting samples for quality control. In our experiments, we use the single-image subset of SPAR-Bench, which contains 2,866 samples in total. Detailed task definitions and per-task sample counts are summarized in Table 2.
| Task Name | Description | #Samples |
|---|---|---|
| depth_prediction_oc | Given the depth of a reference point, estimate the depth of another queried point. | 360 |
| depth_prediction_oo | Given the depth of a reference point, estimate the depth difference between two other queried points. | 372 |
| distance_prediction_oc | Estimate the distance between a queried point and the camera. | 393 |
| distance_prediction_oo | Estimate the distance between two queried points. | 363 |
| distance_infer_center_oo | Determine which of two queried points is farther away from a given reference point. | 340 |
| obj_spatial_relation | Under the camera view, determine the relative position of one queried point with respect to another. | 364 |
| spatial_imagination_oc | First determine the position of a queried point under the camera view, then re-judge its position after a viewpoint transformation. | 372 |
| spatial_imagination_oo | First determine the relative position between two queried points under the camera view, then re-judge their relation after a viewpoint transformation. | 302 |
| Total | 2866 |
3pt
SpatialLadder-26K is a multimodal spatial reasoning dataset designed to support progressive learning from perceptual grounding to more advanced reasoning. It contains 26,610 samples across four complementary task categories, including object localization, single-image spatial reasoning, multi-view spatial reasoning, and video spatial reasoning. In terms of task coverage, the dataset spans seven spatial dimensions, namely relative direction, relative distance, absolute distance, object size, counting, room size, and appearance order, thereby covering a broad range of spatial understanding skills across image, multi-view, and video modalities. In terms of data sources, the object localization, single-image, and multi-view portions are constructed from ScanNet 3D scene reconstructions, while the video subset is sampled from SR-91k.
The paper further introduces two dedicated benchmarks built using the same pipeline on the ScanNet validation set, namely SPBench-SI and SPBench-MV. In our experiments, we use only the single-image benchmark, namely SPBench-SI, which contains 1,009 samples in total. Detailed task definitions and per-task sample counts are summarized in Table 3.
| Task Name | Description | #Samples |
|---|---|---|
| object_rel_direction | Under the camera view, determine the relative position between two objects. | 306 |
| object_rel_distance | Determine which object is farther away from a given reference object. | 91 |
| object_abs_distance | Under the camera view, estimate the distance between two objects. | 149 |
| object_size_estimation | Estimate the size of an object along a given dimension in centimeters. | 463 |
| Total | 1009 |
3pt
ViewSpatial-Bench is a dedicated benchmark for evaluating multi-perspective spatial localization in vision-language models. It contains over 5,700 multiple-choice question-answer pairs spanning more than 1,000 unique 3D scenes, with source images drawn from the validation sets of ScanNet and MS-COCO. In terms of task coverage, the benchmark comprises multiple task types under complementary perspective settings, targeting both camera-perspective and person-perspective spatial reasoning. This design makes it a challenging benchmark for evaluating cross-viewpoint spatial understanding.
Although the paper also describes a larger multi-perspective spatial training dataset generated by the same automated 3D annotation pipeline, the associated training set is not publicly released. Accordingly, in our experiments, we use only the single-image evaluation subset of ViewSpatial-Bench, which contains 4,607 samples in total. Detailed task definitions and per-task sample counts are summarized in Table 4.
| Task Name | Description | #Samples |
|---|---|---|
| Camera perspective – Relative Direction | Under the camera view, determine the relative position of one queried point with respect to another. | 1773 |
| Camera perspective – Object View Orientation | Given an image, determine the facing orientation of the object. | 996 |
| Person perspective – Object View Orientation | Determine the facing orientation of the person appearing in the image. | 996 |
| Person perspective – Relative Direction | Given a specified person-centered perspective, determine the relative position of an object. | 842 |
| Total | 4607 |
3pt
Main results in the paper using grouped abbreviations instead of listing every original task separately. Table 5 summarizes the mapping from each grouped metric to its corresponding original tasks. In all cases, each grouped metric is computed as the average over the corresponding original tasks.
| Benchmark | Abbr. | Original Task |
|---|---|---|
| SPAR-Bench | Dep. | depth_prediction_oc, depth_prediction_oo |
| Dis. | distance_prediction_oc, distance_prediction_oo | |
| Prox. | distance_infer_center_oo | |
| Rel. | obj_spatial_relation | |
| View. | spatial_imagination_oc, spatial_imagination_oo | |
| SPBench | Rel. | object_rel_direction, object_rel_distance |
| Abs. | object_abs_distance, object_size_estimation | |
| ViewSpatial | Cam. | Camera perspective – Relative Direction |
| Camera perspective – Object View Orientation | ||
| Per. | Person perspective – Object View Orientation | |
| Person perspective – Relative Direction |
We construct two training datasets for GeoAnchor. The first is a 3D grounding dataset used in Stage 1 to initialize local spatial perception. The second is a spatial reasoning training dataset used in Stages 2 and 3 as well as the subsequent RL stage. In this section, we describe the construction pipeline, annotation procedure, and summary statistics of these datasets.
Figure 8:
.
Our 3D grounding dataset is constructed based on ScanNet. We sample 10k scenes from more than 2.5 million views spanning over 1,500 scans. For each image, we first employ Qwen3-VL-32B to identify up to five salient objects, together with their textual descriptions and 2D bounding boxes. The prompt used for object extraction is shown in Figure. 8.
To ensure the consistency between the generated referring expressions and the localized regions, we perform an additional verification step after object extraction. Concretely, for each object, we take the generated referring expression and use it as a grounding query on the same image to obtain a second predicted bounding box. We then measure the overlap between the original box and the re-grounded box using intersection-over-union (IoU). An annotation is retained only if the IoU exceeds 0.5; otherwise, it is discarded. This self-consistency check effectively filters out noisy cases in which the generated textual description does not faithfully correspond to the extracted object region. The prompt used for verification is shown in Figure. [fig:verification95prompt].
Table ¿tbl:tab:grounding95verification95stats? reports the verification results. Across 10,000 images, 47,663 object annotations are automatically extracted, and 44,998 pass the self-consistency check, giving an object-level retention rate of 94.41%. The mean and median IoU between the original and re-grounded boxes are 0.787 and 0.856, showing that most generated referring expression–bounding box pairs are well aligned.
| Question Type | Count |
|---|---|
| Single Position Question | 224,990 |
| Multiple Position Question | 159,530 |
| Direction Question | 164,920 |
| Total | 549,440 |
We estimate depth maps and camera poses using Depth Anything v3, and recover object 3D coordinates through depth-guided back-projection. These recovered coordinates are further converted into object-level position annotations and pairwise direction annotations for supervising local latent learning. For an object with bounding box \(b_i=[x_i^{(1)}, y_i^{(1)}, x_i^{(2)}, y_i^{(2)}]\), we define its 2D reference point as the box center, i.e., \(u_i=(x_i^{(1)}+x_i^{(2)})/2\) and \(v_i=(y_i^{(1)}+y_i^{(2)})/2\). Let \(\Omega_i\) denote the set of pixels covered by the box. The object depth is estimated by averaging the predicted depth values inside the box: \[z_i = \frac{1}{|\Omega_i|}\sum_{(u,v)\in\Omega_i} D(u,v).\]
Given the camera intrinsic matrix \(K\) and pose \((R,\mathbf{t})\), where \(R\in\mathbb{R}^{3\times3}\) and \(\mathbf{t}\in\mathbb{R}^3\), the object position in the camera and world coordinates is computed as: \[\mathbf{p}_i^{c}=z_iK^{-1} \begin{bmatrix} u_i, v_i, 1 \end{bmatrix}^T, \qquad \mathbf{p}_i^{w}=R\mathbf{p}_i^{c}+\mathbf{t}.\] We use \(\mathbf{p}_i^{w}\) as the object-level position annotation. For an object pair \((i,j)\), the relative displacement \(\Delta_{ij}=\mathbf{p}_j^{w}-\mathbf{p}_i^{w}\) is used to derive pairwise direction annotations.
Table [tbl:tab:pseudo95depth953d] shows that 3D coordinates recovered from pseudo depth closely match those from ground-truth depth. The mean and median Euclidean errors are only 0.09,m and 0.04,m, with most deviation concentrated on the depth axis. Meanwhile, 98.0% and 92.9% of recovered points fall within 0.5,m and 0.2,m, respectively. This confirms that pseudo depth introduces only limited noise and serves as reliable 3D supervision for subsequent local latent learning.
To increase the diversity of supervision, we construct training queries in three forms: point-coordinate queries, bounding-box-based queries, and natural-language queries. Each sample involves one to three objects depending on the task type. After filtering invalid detections and geometrically unreliable samples, the final pipeline yields 550k 3D grounding samples. Table 6 summarizes the statistics for grounding dataset.
| Source Dataset | Type | Question Type | Count |
|---|---|---|---|
| ScanNet | depth_prediction_oc | numeric | 4000 |
| depth_prediction_oo | numeric | 4000 | |
| distance_infer_center_oo | multiple choice | 8000 | |
| distance_prediction_oc | numeric | 4000 | |
| distance_prediction_oo | numeric | 4000 | |
| obj_spatial_relation_oo | multiple choice | 8000 | |
| spatial_imagination_oc | multiple choice | 4000 | |
| spatial_imagination_oo | multiple choice | 4000 | |
| ScanNet++ | depth_prediction_oc | numeric | 4000 |
| depth_prediction_oo | numeric | 4000 | |
| distance_infer_center_oo | multiple choice | 4000 | |
| distance_prediction_oc | numeric | 4000 | |
| distance_prediction_oo | numeric | 4000 | |
| obj_spatial_relation_oo | multiple choice | 4000 | |
| spatial_imagination_oc | multiple choice | 4000 | |
| spatial_imagination_oo | multiple choice | 4000 | |
| Structured3D | depth_prediction_oc | numeric | 4000 |
| depth_prediction_oo | numeric | 4000 | |
| distance_prediction_oc | numeric | 4000 | |
| distance_prediction_oo | numeric | 8000 | |
| spatial_imagination_oc | numeric | 4000 | |
| spatial_imagination_oo | numeric | 4000 | |
| Total | 100000 |
| Question Category | Question Type | Count |
|---|---|---|
| Relative Direction | multiple choice | 2253 |
| Absolute Distance | numeric | 1127 |
| Object Size | numeric | 1514 |
| Relative Distance | multiple choice | 1034 |
| Total | 5928 |
The spatial reasoning training dataset is mainly built from SPAR and SpatialLadder-26K. We sample 100k questions from SPAR and 5k from SpatialLadder-26K. Basic dataset statistics are summarized in Table 7 and Table 8.
The complete implementation details are summarized in Table 9. The table covers the key experimental settings for our 4-stage collaborative training.
| Category | Item | Setting |
|---|---|---|
| General | Base MLLM | Qwen3-VL-2B |
| Hardware | 8 \(\times\) NVIDIA A800 | |
| SFT | Local token length | \(l_{\mathrm{pos}} = l_{\mathrm{dir}} = 2\) |
| Global token length | \(l_{\mathrm{geo}} = 8\) | |
| S1 training epochs | 1 | |
| S1 batch size | 64 | |
| S1 learning rate | \(1 \times 10^{-4}\) | |
| S2/3 training epochs | 1 | |
| S2/3 batch size | 32 | |
| S2/3 learning rate | \(2 \times 10^{-5}\) | |
| Text loss weight | \(\lambda_t = 1\) | |
| Local loss weight | \(\lambda_l = 1\) | |
| Global loss weight | \(\lambda_g = 0.1\) | |
| VGGT pooling levels | \(L = 3\) | |
| VGGT pooling resolutions | \(\{r_1, r_2, r_3\} = \{1, 2, 4\}\) | |
| Geometry balance coefficient | \(\lambda_{\mathrm{bal}} = 0.05\) | |
| RL | Rollouts per question | \(N = 8\) |
| Sampling temperature | 1.0 | |
| Pattern reward | \(r_{\mathrm{pattern}} = 0.5\) | |
| KL-divergence coefficient | \(\beta = 0.01\) | |
| Learning rate | \(5 \times 10^{-7}\) | |
| EMA smoothing coefficient | \(\kappa = 8\) | |
| EMA update rate | \(\mu = 0.2\) |
As described in Section 4.3, the tasks from the three benchmarks are regrouped into three categories: Position, Direction, and Mixed, according to the type of local spatial evidence they require. Table 10 further provides an explicit mapping from these three categories to the original task types in each benchmark.
| Category | Task | Benchmark |
|---|---|---|
| Position | depth_prediction_oc | SPAR-Bench |
| depth_prediction_oo | SPAR-Bench | |
| distance_prediction_oc | SPAR-Bench | |
| distance_prediction_oo | SPAR-Bench | |
| distance_infer_center_oo | SPAR-Bench | |
| obj_spatial_relation | SPBench | |
| object_rel_direction | SPBench | |
| obj_abs_distance | SPBench | |
| object_size_estimation | SPBench | |
| object_rel_distance | SPBench | |
| Camera - Relative Direction | ViewSpatial | |
| Direction | Camera - Object View Orientation | ViewSpatial |
| Person - Object View Orientation | ViewSpatial | |
| Mixed | spatial_imagination_oc | SPAR-Bench |
| spatial_imagination_oo | SPAR-Bench | |
| Person - Relative Direction | ViewSpatial |
This subsection describes the three dense alignment baselines presented in ablation study of VGGT alignment strategy. All three methods use the same projected geometry tokens as in Section 3.2, and differ only in how the dense VGGT feature sequence is compressed to match the length of the global token. Following Section 3.2, let the final VGGT feature map be of size \(H \times W \times D\). After flattening the spatial dimensions in raster order, the VGGT feature sequence is denoted as: \[F = \{f_1, \ldots, f_{l_f}\} \in \mathbb{R}^{l_f \times D}, \qquad l_f = H \times W,\] and the global latent \(z^{\mathrm{geo}}\) is projected into geometry tokens: \[G = \{g_1, \ldots, g_{l_{\mathrm{geo}}}\} \in \mathbb{R}^{l_{\mathrm{geo}} \times D},\] where \(l_{\mathrm{geo}}\) is the number of geometry tokens.
| SPAR-Bench | SPBench | ViewSpatial | ||||||||||||||||||||||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 2-7(lr)8-10(lr)11-13 Model | Avg. | Dep. | Dis. | Prox. | Rel. | View | Avg. | Rel. | Abs. | Avg. | Cam. | Per. | ||||||||||||||||||||||||
| Qwen2.5-VL-3B | 28.7 | 25.5 | 25.7 | 55.6 | 29.4 | 21.7 | 32.9 | 42.8 | 26.5 | 37.3 | 39.7 | 33.6 | ||||||||||||||||||||||||
| + vanilla SFT | 59.8 | 42.5 | 56.8 | 76.8 | 75.2 | 65.1 | 59.2 | 68.0 | 53.5 | 41.9 | 36.1 | 46.8 | ||||||||||||||||||||||||
| + text CoT SFT | 60.7 | 45.2 | 60.6 | 74.2 | 76.1 | 62.5 | 52.5 | 56.2 | 50.1 | 39.3 | 36.8 | 43.1 | ||||||||||||||||||||||||
| GeoAnchor | 66.0 | 49.7 | 62.9 | 81.5 | 81.0 | 71.4 | 66.1 | 76.6 | 59.3 | 47.0 | 46.8 | 47.4 | ||||||||||||||||||||||||
| Improvement | +37.3 | +24.2 | +37.2 | +25.9 | +51.6 | +49.7 | +33.2 | +33.8 | +32.8 | +9.7 | +7.1 | +13.8 | ||||||||||||||||||||||||
Mean Pooling. This baseline performs the strongest compression. It averages the dense VGGT sequence and the geometry-token sequence along the sequence dimension, producing one global vector for each side, and then directly aligns the two pooled vectors with Smooth L1 loss. In other words, the entire feature sequence is reduced to a single scene-level representation before alignment.
Adaptive Pooling. This baseline first resamples the VGGT sequence from length \(l_f\) to length \(l_{\mathrm{geo}}\) by adaptive average pooling. Concretely, the sequence is partitioned into \(l_{\mathrm{geo}}\) contiguous bins of nearly equal size, and the features inside each bin are averaged to obtain a resampled VGGT target sequence. This produces one target vector for each geometry token, enabling token-wise dense supervision after length matching.
Linear Interpolation. This baseline also converts the VGGT sequence to length \(l_{\mathrm{geo}}\), but uses 2D linear interpolation instead of average pooling. Compared with adaptive pooling, it preserves the sequence order through continuous interpolation between neighboring VGGT features, and then aligns the interpolated sequence with the geometry tokens in a token-wise manner.
For the latter two baselines, let \(\tilde{F}=\{\tilde{f}_1,\ldots,\tilde{f}_{l_{\mathrm{geo}}}\}\in\mathbb{R}^{l_{\mathrm{geo}}\times D}\) denote the resampled VGGT sequence obtained by either adaptive pooling or linear interpolation. Their alignment loss is defined uniformly as: \[\mathcal{L}_{\mathrm{dense}}= \frac{1}{l_{\mathrm{geo}}} \sum_{j=1}^{l_{\mathrm{geo}}} \mathrm{SmoothL1}(g_j,\tilde{f}_j).\]
We implement our latent reasoning framework on Qwen2.5-VL-3B, and Table 11 shows the corresponding results. GeoAnchor consistently outperforms conventional SFT methods such as vanilla SFT and text CoT SFT, exhibiting the same improvement trend as on the Qwen3-VL-2B illustratede in the main paper. These results demonstrate the effectiveness and generalizability of our method, as it can be seamlessly integrated into different model backbones and consistently improve their spatial reasoning capabilities.
We carefully tune the lengths of the local and global tokens to ensure that the latent space remains both expressive and compact. Table ¿tbl:tab:ablation95latent95length? shows that the best overall performance is achieved when the local token length is 2 and the global token length is 8. This indicates that effective spatial reasoning depends on a latent space that remains compact while preserving sufficient spatial information. Overly short latents impose an excessive information bottleneck, while overly long latents weaken the compactness and functional specialization of the latent space. The balanced setting therefore provides the most effective trade-off between preserving spatial information and supporting downstream reasoning.
Table [tbl:tab:ablation95sampleing] shows that the multi-scale setting \(\{1,2,4\}\) achieves the best overall performance, obtaining the best results on all three benchmarks. This indicates that combining coarse-to-fine spatial resolutions provides more effective geometry supervision than either a smaller or an overly fine scale set. Moreover, although the single-scale \(5\times5\) setting uses a comparable number of tokens, it still underperforms \(\{1,2,4\}\). This suggests that the advantage of multi-scale pooling lies in its ability to capture complementary geometric cues at different spatial resolutions, jointly preserving global scene layout and finer local structure, which leads to more effective geometry alignment.