Abstract: Flow-matching-based vision-language-action (VLA) models have emerged as powerful policies for robotic manipulation, yet a critical capability remains underexplored: fine-grained behavioral control, the ability to govern how a robot performs a task by intervening on its internal representations. Representation steering is a well-established interpretability tool for language and vision-language models, where behavioral features are typically encoded as linear directions, but we show that these classic methods fall short in VLAs. We propose DiMaS, a Distribution-Matching Steering strategy tailored to flow-matching VLAs, which transports between representation distributions rather than shifting along a fixed direction, and show that it effectively controls behavior across two state-of-the-art VLAs. We further examine the generalizability of this strategy as the tasks it is learned from and evaluated on grow increasingly dissimilar, characterizing where behavioral control transfers and where it weakens. Finally, through an analysis of the representation structure of the action expert, we explain why classical linear steering falls short in the visuomotor setting: behavioral features are linearly decodable but not linearly steerable, which motivates the distribution-matching design of DiMaS. Our code is publicly available.1\(^,\) 2

Keywords: Vision-Language-Action Models; Behavioral Control; Representation Steering; Robot Manipulation

1 Introduction↩︎

The recent rise of foundation models, spanning Large Language Models (LLMs) [1], [2] and Vision-Language Models (VLMs) [3], [4], has driven rapid progress in robot policies, now commonly framed as Vision-Language-Action (VLA) models [5][9]. These policies aim to transfer the rich visual and linguistic priors encoded in LLMs and VLMs to robotic control.

Concretely, a VLA takes images from visual sensors together with a natural-language task instruction as multimodal inputs, and predicts low-level robot actions; these inputs are typically processed by a pretrained or finetuned VLM backbone. Early VLAs generated actions autoregressively, treating them as discrete tokens [5], [6]. More recent state-of-the-art systems instead pair the VLM with a flow-matching [10] action expert that predicts continuous actions conditioned on the VLM representations [7], [8].

Despite this progress, steering VLAs remains underexplored: that is, developing post-hoc methods that control specific attributes of the predicted actions by intervening on internal representations. For any deployed generative model, such control carries substantial application value, since it gives a user a direct mechanism to shape the output. Representation steering is now a well-established and effective tool for LLMs and VLMs [11][14], yet it is unclear how these recipes carry over to VLAs, where a flow-matching action expert produces continuous trajectories rather than discrete tokens. Our goal in this paper is to develop steering methods that intervene on VLA representations to control targeted features of the predicted trajectory, and to verify that this control generalizes to settings beyond those used to calibrate the intervention. Further, our analysis of the action expert’s representation structure yields broad insights into why classical steering strategies fall short in VLAs.

Through this work, we make key contributions in the following directions:

  • We propose a simple and novel Distribution Matching based Steering (DiMaS) mechanism, a steering strategy specifically adapted to recent flow-matching VLAs. It intervenes on latent representations and applies across models of different scales, which we validate on SmolVLA [7] and \(\pi_{0.5}\) [8] across multiple steering tasks.

  • We examine the generalizability of steering as the tasks it is learned from and evaluated on grow increasingly dissimilar, characterizing where behavioral control transfers and where it weakens, and yielding insight into the robustness of representation-level control in VLAs.

  • Through an investigation of representation structure inside a flow-matching VLA, with a particular focus on the action expert, we provide insights into why classical linear steering strategies that succeed for LLMs and VLMs fall short in the visuomotor setting, motivating the design of DiMaS.

2 Related Works↩︎

Representation steering in generative models. Representation steering refers to performing inference-time interventions on the internal representations of generative models to control specific attributes or behaviors of their output. With roots in latent-space editing of VAEs and GANs [15], [16], the idea has gained significant traction with the rise of large language models (LLMs) and multimodal large language models (MLLMs), where it has proven effective for addressing alignment problems [11], [17][19] and controlling generation style in multimodal models [13], [14], [20].
Most such approaches are grounded in the linear representation hypothesis (LRH), which posits that semantic features are encoded as linear directions in activation space [21], and that representations can be expressed as sparse linear combinations of such directions [22], enabling simple additive interventions at inference time. While recent work has questioned the LRH [23], [24] and proposed richer models of concept representation [25], linear steering remains the dominant paradigm for controlling LLMs and VLMs. We find that a fixed linear shift is inadequate for controlling behavioral features in VLAs, motivating a strategy based on distribution matching. In spirit, our approach is most related to the optimal transport-based steering of Rodriguez et al. [26], which they develop for LLMs and diffusion models; our focus differs, targeting behavioral control in the visuomotor representations of flow-matching VLAs.

Interpretability and steering for VLAs. [27] study mechanistic interpretability in autoregressive VLAs that generate discrete action tokens. They identify clusters of FFN value vectors that can be steered at inference time by overriding the corresponding neuron activations, and give early evidence that such interventions can modulate motion speed and trajectory height. Their framework applies only to autoregressive VLAs, whereas we target the current state-of-the-art flow-matching VLAs.

Closest to our work, [28] formalize feature-observability and feature-controllability in VLAs. They learn a regressor to predict behavioral features from VLM hidden states and use its coefficients as a steering vector to shift the VLM representations at inference time. A key limitation is that they must pause the intervention to preserve task success. Our method preserves task success without pausing the intervention, and replaces the linear shift with a steering strategy based on distribution matching.

3 Method↩︎

We first introduce the notation and the general VLA architecture we consider in 3.1, and then present our method, DiMaS, in [method:dimas].

3.1 Notation and VLA architecture↩︎

We assume a generic VLA architecture that covers the paradigm of current SOTA VLAs based on joint VLM and flow-matching networks for action prediction. At any timestep \(t\), the VLA \(f\) receives an observation \(x_t = (I_t, T, x_t^s)\) consisting of an image, text instruction, and proprioceptive state and predicts low-level actions for each robot joint, denoted as \(y_t\). The VLA \(f = (f_V, f_A)\) consists of two main components, \(f_V\), which is the VLM backbone and \(f_A\), the action expert network. The VLM representations are used to condition the action expert \(f_A\) that generates the actions. It is trained with a flow-matching objective [10] to iteratively refine noisy action tokens \(a_{0,t} \sim \mathcal{N}(0, \mathbf{I})\) across \(M\) denoising steps. \[a_{m+1, t} = a_{m,t} + \frac{1}{M}\, f_A\!\left(a_{m,t},\, m,\, f_V(x_t)\right), \, \, a_{0,t} \sim \mathcal{N}(0, \mathbf{I}),\;m \in \{0, \ldots, M-1\}\] The final predicted action chunk \(y_t\) is obtained directly from \(a_{M,t}\) through an environment-specific transformation.

Throughout this paper, we analyze or intervene on residual stream representations that can be sourced from different parts of \(f\). In the LLM backbone of the VLM, the residual stream representation at layer \(l\) and token position \(p\) is denoted as \(h^p_l(x_t) \in \mathbb{R}^{d_V}\). For the action expert \(f_A\), analogously, they are denoted as \(h^{p,m}_l(x_t) \in \mathbb{R}^{d_A}\), where additionally, \(m\) is the denoising step.

3.2 DiMaS: Distribution matching for steering VLA representations↩︎

Figure 1: DiMaS Method Overview: A VLM backbone encodes multimodal robot inputs and conditions a flow-matching based action expert to predict the final action chunk. Training: We extract residual-stream representations h^{p,m}_{l} from the action expert at a specific layer l^*, and flow-matching step m. After obtaining a continuous behavioral feature (e.g.speed); the representations are grouped by feature value to create source and target distributions \mathcal{D}^{-}, \mathcal{D}^{+} (feature absent, present respectively). The steering intervention is learnt as a transport map between the two distributions. Inference: During test-time, steering is gated via a binary classifier g. To achieve feature control and high success rate we interpolate between the source/target distributions through the transport map \mathcal{T} and interpolation knob \alpha \in[0, 1].

We introduce a novel steering method grounded in optimal transport that learns a mapping between distributions of internal representations. A complete overview of the method is illustrated in 1.
Representations and behavioral feature. The first step of our method sources representations from action expert \(f_A\), and pairs each with a continuous scalar feature/concept measuring a target behavioral property we wish to control.
For a given network layer \(\ell\) and flow-matching denoising step \(m\), we extract representations \(h^{p,m}_{\ell}(x_t) \in \mathbb{R}^d\) for each \(x_t\) across all rollouts. Each position \(p\) yields a representation. For simplicity of notation, we will denote the \(i\)’th extracted representation as \(h_i\). To extract the corresponding scalar feature in our experiments, we use the predicted action to extract the feature of interest (although not necessary for our method). Concretely, each predicted action is a vector \(a_i = (\Delta x, \Delta y, \Delta z, \dots, \texttt{gripper})\), encoding the end-effector displacement together with the gripper state. From these actions we derive a scalar behavioral feature; for instance, when the targeted property is the end-effector speed, the feature is the magnitude of the translational displacement \(\phi(a_i) = \sqrt{\Delta x^2 + \Delta y^2 + \Delta z^2}\). Each representation \(h_i\) is then associated with the feature value of its corresponding action \(\phi_i=\phi(a_i)\), giving a continuous score for the behavioral property we wish to control.
Source and target distributions. To partition the representations into a source distribution \(\mathcal{D}^{-}\) and a target distribution \(\mathcal{D}^{+}\), we threshold on the behavioral feature using its empirical quantiles:

\[\mathcal{D}^{-} = \{h_i\,:\, \phi_i \leq q_{\tau}\}, \, \, \mathcal{D}^{+} = \{h_i\,:\, \phi_i > q_{1-\tau}\}\] Representations whose feature value falls below the lower quantile \(q_{\tau}\) form the source set \(\mathcal{D}^{-}\), while those above the upper quantile \(q_{1-\tau}\) form the target set \(\mathcal{D}^{+}\). Restricting each distribution to a tail, rather than splitting at the median, yields cleaner, well-separated populations of the feature-absent and feature-present behavior, which sharpens the learned mapping.

Steering as a transport map. At its core, steering a behavioral feature amounts to learning a map \(\mathcal{T}\) that transports a representation from the source distribution \(\mathcal{D}^{-}\) to the target distribution \(\mathcal{D}^{+}\), thereby inducing the feature in representations that lack it. Different steering strategies correspond to different choices of \(\mathcal{T}\). The simplest and most widely used is linear steering, where \(\mathcal{T}\) is an additive shift in a fixed direction. This fixed direction can be extracted in multiple ways. For instance, one popular form is the mean-difference steering, where this direction is equal to the difference of the distribution means, \(\mathcal{T}(h) = h + (\mu^{+} - \mu^{-})\), with \(\mu^{\pm}= \mathbb{E}_{h \sim \mathcal{D}^\pm}[h]\). A more flexible variant of linear steering is based on regression, where \(\mathcal{T}\) displaces the representation along the direction of a regressor trained to predict the feature value, so that the shifted representation attains a desired target value (refer to Appendix 9 for more details). Both are linear by construction, which, as we empirically show in 5, limits their ability to control behavioral features in VLAs with a denoising action expert. DiMaS instead instantiates \(\mathcal{T}\) as an optimal transport map, which respects the full geometry of \(\mathcal{D}^{-}\) and \(\mathcal{D}^{+}\) rather than a single direction.

Concretely, we learn the transport map \(\mathcal{T}^{(m)}_{\ell}(\cdot)\) from \(\mathcal{D}^{-}\) to \(\mathcal{D}^{+}\) (see Figure 1, left) by minimizing the Kantorovich optimal transport objective, \[\mathcal{W}_2^2(\mathcal{D}^{-}, \mathcal{D}^{+}) = \min_{\gamma \in \Pi(\mathcal{D}^{-}, \mathcal{D}^{+})} \int_{\mathcal{Z} \times \mathcal{Z}} \|z^{-} - z^{+}\|^2 \, \mathrm{d}\gamma(z^{-}, z^{+}),\]

where \(\mathcal{Z}\) denotes the support of the distributions and \(\|\cdot\|\) is the standard Euclidean norm acting as the ground cost. In practice, \(\mathcal{D}^{-}\) and \(\mathcal{D}^{+}\) are only accessible through finite empirical samples, so this reduces to a discrete optimal transport problem, which we solve efficiently via low-rank Sinkhorn [29]; full details are provided in 8.1.
Test-time intervention strategy. Given a representation \(h = h^{p,m}_{\ell}(x_t)\) at test time, we first project \(h\) onto \(\mathcal{D}^{-}\) via a nearest-neighbor projection \(P(h) = \arg\min_{z \in \mathcal{D}^{-}} \|z - h\|\), and then apply the learned transport map to obtain the steered representation \(\mathcal{T}^{(m)}_{\ell} \circ P(h)\). To avoid unnecessary interventions, steering is gated by a feature filter: a binary classifier \(g_{\ell}^{(m)} : \mathbb{R}^{d} \to \{0, 1\}\), implemented as a linear probe trained to separate \(\mathcal{D}^{-}\) from \(\mathcal{D}^{+}\). The gate becomes active, i.e., \(g_{\ell}^{(m)}(h) = 1\) when the feature is absent, i.e.\(h\) is predicted to lie in \(\mathcal{D}^{-}\), and \(g_{\ell}^{(m)}(h) = 0\) when it is already present, so that representations already exhibiting the targeted feature are left unchanged.
Interpolation as the secret sauce. A crucial goal behind our steering methodology is that a user should be able to control the feature of interest while simultaneously succeeding at the task. In other words, task success should not be achieved by needing to stop the steering intervention. In our experiments, we observed that while completely transporting \(h\) to \(\mathcal{D}^+\) via \(\mathcal{T}^{(m)}_{\ell} \circ P(h)\) controls the feature of interest, it introduces deviation from the original \(h\) that can lead to task failures. Given that we perform evaluations in a closed-loop setup and tasks/models can often be fragile, this sometimes causes significant drop in success rate. To ensure a minimal drop in success rate while still controlling the feature, we interpolate between the distributions \(\mathcal{D}^-\) and \(\mathcal{D}^+\), via a parameter \(\alpha \in [0, 1]\). This intervention is summarized below and illustrated on the right of 1:

\[h \;\leftarrow\; \begin{cases} (1 - \alpha)\, h \;+\; \alpha \, \big(\mathcal{T}^{(m)}_{\ell} \circ P(h)\big), & \text{if } g_{\ell}^{(m)}(h) = 1, \\[4pt] h, & \text{otherwise.} \end{cases}\] Here \(\alpha = 0\) recovers the unsteered representation and \(\alpha = 1\) applies the full transport, with intermediate values yielding a continuous modulation of the target behavior. From a user perspective, \(\alpha\) provides a normalized, interpretable interpolation knob to control steering strength while maintaining success rate as much as possible. Interestingly, for OT based LLMs/VLM steering [26], \(\alpha\) plays a similar role of a normalized knob controlling steering strength. For DiMaS, it serves an even more important purpose by also affecting the success rate.

4 Experiments↩︎

We evaluate our method on two flow-matching-based VLA models of different scales, SmolVLA [7] and \(\pi_{0.5}\) [8] and LIBERO benchmark [30] for controlled, axis-isolated evaluation. This simulator consists of four suites, namely LIBERO-Object, LIBERO-Spatial, LIBERO-Goal, and LIBERO-10, where the first three comprise shorter-horizon tasks and LIBERO-10 comprises longer-horizon ones. Each suite contains 10 tasks, and each task provides 50 initializations that vary initial placement of the objects. More details are provided in Appendix 8.3.
Target feature. We consider two target behavioural features: motion speed, measured as the mean norm of end-effector velocity, and end-effector vertical displacement, measured as the accumulated absolute displacement along the vertical axis.
Baselines. We compare against three categories of baselines. (i) Mean-difference steering applies a constant additive shift equal to the difference of the source and target means, \(\mathcal{T}(h) = h + (\mu^{+} - \mu^{-})\), where \(h\) represents a hidden representation of the VLM component. (ii) Regression-based steering, based on [28], fits a linear regressor to predict \(\phi\) from \(h\) and shifts the representation along the regressor’s direction so that its predicted feature attains a target quantile \(q^\star\). We evaluate this baseline both in the VLM and the flow-matching (FM) action expert. Its closed form is given in Appendix 9. Both representation-level baselines are linear steering methods by construction, with the shift gated by a linear classifier and applied only to representations that do not yet exhibit the target behavior. (iii) Prompt injection modulates behavior purely through instruction rephrasing (e.g.”go faster,” “slower”), with no activation-level intervention.
Metrics. We report three metrics: Success Rate (SR), which measures task completion; feature value, the value of the targeted behavioral feature (speed or \(z\)-displacement); and the statistical significance of the feature shift relative to the unsteered baseline, assessed via paired \(t\)-tests across episodes.

4.1 Can representation steering control behavior?↩︎

We first ask whether the targeted feature can be modulated/controlled, and how DiMaS compares to the classic steering and instruction-level baselines introduced above. For each method we measure two quantities relative to the unsteered policy: whether it shifts the mean feature value (paired \(t\)-test) and how much its success rate changes. 2 and 3 present these results for decreasing (blue, H\(\to\)L, high to low) and increasing (red, L\(\to\)H, low to high) end-effector speed and vertical displacement, evaluated across all baselines. Statistically significant shifts are shown in solid fill, while non-significant ones are striped.
The linear steering and prompt baselines perform inconsistently. They leave the target feature unchanged, and where they do move it, they do so unreliably: the increasing and decreasing interventions often shift the feature in the same direction rather than opposite ones. DiMaS, by contrast, modulates the target feature in both directions: speed on both models, and vertical displacement on \(\pi_{0.5}\). We largely preserve the success rate when modulating speed, but incur a drop when modulating vertical displacement. This is an expected result, since this feature is often tied to task completion. On LIBERO-Object, for example, decreasing vertical displacement keeps the end-effector too low to lift the object into the basket, so the placing step fails. Finally, for both features the decreasing intervention produces the more statistically significant shifts (e.g. color-filled marks).

While \(\pi_{0.5}\) achieves the strongest results overall, our steering approach proves effective across both models despite their substantial difference in scale, underscoring its generality. SmolVLA shows comparatively more sensitivity to action intervention, which we attribute in part to its action chunking strategy: SmolVLA predicts a new chunk at every timestep, producing more dynamic, less smoothed trajectories, whereas \(\pi_{0.5}\) predicts 10 actions at a time, yielding smoother motion that is inherently more robust to intervention.

Figure 2: Speed steering. Change in mean end-effector speed (\Deltaspeed) versus change in success rate (\DeltaSR), relative to the unsteered baseline, across the three LIBERO suites, for (a) \pi_{0.5} and (b) SmolVLA. Left of the vertical dashed line corresponds to high-to-low (H→L) steering, right to low-to-high (L→H). Marker shape denotes method; solid fill indicates a statistically significant change relative to baseline (t-test, p<0.01).
Figure 3: Z-displacement steering. Change in accumulated vertical displacement (\Deltaz-displacement) versus change in success rate (\DeltaSR), relative to the unsteered baseline, across the three LIBERO suites. Conventions as in 2.

4.2 Understanding steering generalization↩︎

We next investigate what makes a learned steering effective, and how well it transfers, using LIBERO’s axis-isolated task structure. We vary two axes: (1) task diversity, the number and variety of tasks the steering vector is learned from, and (2) test distribution shift severity, i.e. how much the evaluation tasks differ from the training tasks.

  1. Setting 1: Held-out initial states of the same task. For each task individually, we learn a steering from a subset of initial states and evaluate it on the remaining, held-out initial states of the same task.

  2. Setting 2: Held-out initial states across tasks in a suite. Considering all tasks jointly, we learn a single steering from a subset of initial states and evaluate its transfer to the remaining initial states. This differs from Setting 1 only in the diversity of the tasks used to learn the steering, not in the evaluation protocol: both evaluate on held-out initial states, so the comparison isolates the effect of learning from a single task versus from many.

  3. Setting 3: Held-out tasks within a suite. Within a given LIBERO suite (e.g.LIBERO-Object), we learn a steering from a subset of tasks and evaluate it on the held-out tasks of the same suite. This is the first setting that tests transfer to tasks unseen at training time.

  4. Setting 4: Transfer across disjoint suites. We learn a steering from one LIBERO suite (e.g.LIBERO-Object) and evaluate it on a different, disjoint suite (e.g.LIBERO-Goal), testing transfer across distinct task distributions.

Settings 1 and 2 differ in the diversity of the tasks the steering is learned from, asking how much that diversity matters. Settings 3 evaluates transferability of steering within the same suite (e.g. LIBERO-object), while Settings 4 further stress tests the transfer to other suites (e.g. LIBERO-goal). Together, these settings provide fine-grained evidence about both what the steering needs in order to work and how robustly it generalizes. It is worth emphasizing that most steering and even VLA evaluation on LIBERO is typically done in Setting 2. Our settings 3 and 4 validate increasingly out-of-distribution capabilities of the steering method. We compare these steering settings for decreasing the end-effector speed and illustrate the results in 4 (comparison for other tasks is provided in Appendix 10).
Across both models, all three suites, and all four settings, DiMaS shifts speed in the intended direction, with a significant decrease in nearly every case: an intermediate-layer intervention reliably slows the executed action. We now compare the four settings, which progressively increase the gap between the tasks the steering is learned from and those it is evaluated on.
Aggregating tasks helps (Setting 1 vs.Setting 2). Settings 1 and 2 share the same evaluation protocol and differ only in the diversity of the tasks used to learn the steering. Learning from many tasks jointly (Setting 2) yields a more consistent effect than learning from a single task (Setting 1): the induced shift is more stable and the effect is significant across suites on both models. We attribute this to the breadth of the learning set: a single task samples a narrow region of representation space, so the learned transport partly reflects features specific to that task, whereas aggregating tasks lets it capture the structure associated with the target behavior across a broader region of the representation space.
Steering persists on held-out and disjoint tasks (Settings 3–4). Importantly, DiMaS control effectiveness does not collapse when the steering is applied to tasks unseen at training time (Setting 3) or transferred across disjoint suites (Setting 4). For \(\pi_{0.5}\), decreasing speed remains significant across suites at the held-out-task setting (Setting 3), and the effect persists in the intended direction even under the hardest cross-suite transfer to Goal, where the mean shifts in the intended direction although the change does not reach significance at \(p<0.01\) (Setting 4, \(p=0.032\)). The learned transport therefore captures structure tied to the target behavior rather than to the specific tasks it was learned from, and remains usable beyond its training setting.
Effectiveness is weakest on the most diverse suite. LIBERO-Goal is the hardest suite. Its tasks are behaviorally diverse, spanning pushing objects and operating a stove in addition to pick-and-place, whereas Object and Spatial share a pick-and-place behavior across all tasks. This diversity coincides with weaker steering: Goal is the only suite with cells that do not reach significance at \(p<0.01\) (SmolVLA Setting 1, \(p=0.09\); \(\pi_{0.5}\) Setting 4, \(p=0.032\)), and its shifts are generally smaller and noisier than on Object or Spatial, which are significant at every setting on both models. Both Goal’s task heterogeneity and the smaller per-task sample available at Setting 1 likely contribute to this effect.

a
b

c

Figure 4: Steering across evaluation levels on \(\pi_{0.5}\) (left) and SmolVLA (right). We compare the unsteered baseline against DiMaS for decreasing speed. Panels show, left to right, the baseline and the settings discussed in 4.2; columns are the three LIBERO suites (Object, Spatial, Goal). Box plots show the per-episode feature distribution at each level, annotated with success rate and \(p\)-value vs.baseline; box opacity is higher when the shift is statistically significant (\(p<0.01\)). Arrow length is proportional to the induced \(\Delta\)mean.. a — \(\pi_{0.5}\), b — SmolVLA

4.3 Steering on longer-horizon tasks.↩︎

Figure 5: Decreasing speed on LIBERO-10 (\pi_{0.5}).

As shown in 5, we extend our evaluation to LIBERO-10, whose tasks have substantially longer horizons than the suites above, decreasing speed. DiMaS produces a downward shift (\(\Delta\text{mean}=-0.024\), \(p<0.01\) for \(\alpha=0.5\)) at a small success cost (\(96\% \to 90\%\)), showing that behavioral control remains effective even over extended horizons.

5 Investigating internal structure in VLAs↩︎

We now analyze representation structure inside the VLM \(f_V\) and the action expert \(f_A\) to motivate some of DiMaS’s design choices and explain why linear steering methods in our experiments are very inconsistent and ineffective. We do so through the lens of output action/trajectory features we would like to control, asking whether each can be represented as a fixed direction in latent space. As a running example throughout this section, we investigate SmolVLA [7] on LIBERO-Object for the "speed" feature.

Linear separability. To determine whether a vector/direction in the hidden representation space is a “good” candidate for representing a given VLA concept, we first quantify the separability between representations in \(\mathcal{D}^-, \mathcal{D}^+\) across different layers \(l\) in the VLM \(f_V\), and across different layers \(l\) and flow-matching steps \(m\) in \(f_A\). For \(f_V\), we extract the last-token representation \(h^{-1}_l(x)\), which summarizes the full input context. To quantify the linear separability, we fit a linear classifier (SVM) on representations \(h\) drawn from \(\mathcal{D}^-\) and \(\mathcal{D}^+\).

We report the separability accuracies in 6. The representations in the action expert are almost completely linearly separable (accuracy near 100%) at any layer in the later flow-matching steps, as the action tokens are progressively denoised. Interestingly, even at flow-matching step 0, the deeper layers show higher linear separability (\(>93\%\)) than any layer in the VLM (max 87%). This is one of the main reasons we choose to intervene in the action expert rather than the VLM, and why we prioritize the deeper layers in the action expert.

Is a linear shift enough to match distributions? Even though a linear classifier in the action expert can separate high- and low-speed representations, this alone is not sufficient to validate whether the normal vector to the hyperplane (i.e., the steering vector) represents the feature/concept. To do so, we visualize whether shifting representations along the steering vector transports them from \(\mathcal{D}^-\) to \(\mathcal{D}^+\).

We qualitatively visualize this shift in the 2D PCA projection space for layer \(l=0\) and flow-matching step \(m=8\) in 6 (separability is 100\(\%\)). We shift the high-speed representations (red points) by the unit-norm steering vector scaled by a factor \(\beta \in \{2, 50, 300\}\) that controls the strength of linear steering. Note that the representations are steered in the original high-dimensional space. If \(\mathcal{D}^-, \mathcal{D}^+\) can be matched in the original high-dimensional space through a shift in a fixed steering direction, the match will be preserved in the 2D PCA projection space (since PCA projection is a linear operation).
As one might expect, linear steering only shifts the source distribution (red points) in a fixed direction, which leads to the shifted (orange/green) distributions. Since the "shape" of original red (high-speed) and blue (low-speed) point distributions differs, this operation is unable to match the two distributions. It is also worth noting that if \(\beta\) is chosen as a small value or as the minimum positive shift needed to flip the classifier’s decision (similar to our regression baseline [28]), the shifted distributions nearly coincide with the original red points (purple points mixed with red). These observations indicate that using a steering vector as the concept representation fails to capture the difference between the distributions of high and low speed representations and that we need more complex transformations to move from one to the other.

a

b

c

Figure 6: (Left) Linear separability inside SmolVLA for speed feature in different VLM layers and (Middle) different action expert layers, flow-matching steps. Reported as linear SVM classification accuracy (\(\%\)) between \(\mathcal{D}^-, \mathcal{D}^+\). Higher is better. (Right) 2D PCA visualization for representations in \(l=1, m=8\) in action expert. Original high (low) speed representations are in red (blue). Shifting high speed representations along steering vector leads to shifted distributions in orange and green. Linear steering fails to match red and blue distributions..

6 Analyzing DiMaS↩︎

We present two further analyses of DiMaS: the effect of the interpolation factor \(\alpha\) on the trade-off between steering strength and task success, and qualitative examples of steered trajectories for reducing vertical displacement.

Figure 7: image.

Figure 8: image.

Figure 9: image.

7 Conclusion and future work↩︎

We study behavioral control in Vision-Language-Action (VLA) models: not only what a robot does but how it does it. Across two state-of-the-art VLAs with a flow-matching action expert, we show that classical linear steering is too restrictive to control behavior reliably, and our analysis of the hidden representation space, in both the vision-language backbone and the action expert, explains why. We address this with DiMaS, which steers behavior by transporting hidden representations from a source distribution, where the target feature is absent, to a target distribution where it is present, using optimal transport, regularized by an interpolation knob. DiMaS substantially outperforms linear and prompt-based baselines in controlling speed and height while better preserving task success. Finally, we evaluate its generalizability across four settings that vary both the diversity of the tasks the steering is learned from and how far the evaluation tasks depart from them, and distill from this what an effective steering needs in order to be constructed and applied.
Our results open several natural extensions. First, our intervention currently treats all timesteps alike; a next step is to make when to intervene as much a target of learning as how, steering selectively (e.g.only while the end-effector is translating rather than purely rotating) for even smoother control. Second, for policies that emit a chunk of actions, such as \(\pi_{0.5}\), we associate each hidden representation with its corresponding output action; yet an output action is shaped by more than one representation in the chunk, so grouping representations by the feature of their exact corresponding action is only an approximation, and a finer account of this correspondence could sharpen the learned transport. Third, because DiMaS applies to any feature observable from the representations, the same mechanism extends readily to new tasks, embodiments, and more abstract behavioral features, pointing toward representation-level control as a general interface for shaping robot behavior.

Acknowledgements↩︎

This work has been partially supported by ANR grant VISA DEEP (ANR-20-CHIA-0022), HPC resources of IDRIS under the file A0191016602 allocated by GENCI, and Cluster PostGenAI@Paris (ANR-23-IACL-0007, FRANCE 2030). We also thank Mustafa Shukor for his insightful discussions and valuable feedback.

We provide details of our proposed method, DiMaS, in 8, and describe the baseline methods in 9. 10 presents additional experimental results omitted from the main paper. In 11 we analyze DiMaS further, studying the contribution of the interpolation factor \(\alpha\) (11.1) and showing how increasing and decreasing vertical displacement affect the end-effector height (11.2). Finally, we ablate the choice of steering layer in 12.

8 Implementation Details of DiMaS↩︎

8.1 Steering as a transport map: details↩︎

We provide additional details on the optimal transport formulation introduced in Section 3.2. The optimization ranges over \(\Pi(\mathcal{D}^{-}, \mathcal{D}^{+})\), the set of all joint distributions (transport plans) on \(\mathcal{Z} \times \mathcal{Z}\) whose marginals coincide with \(\mathcal{D}^{-}\) and \(\mathcal{D}^{+}\). Intuitively, \(\gamma(z^{-}, z^{+})\) specifies the allocation of mass from \(\mathcal{D}^{-}\) to \(\mathcal{D}^{+}\) that minimizes the total quadratic transport cost.

In practice, \(\mathcal{D}^{-}\) and \(\mathcal{D}^{+}\) are only accessible through finite empirical samples \(\mathbf{X}^{-} = \{z^{-}_i\}_{i=1}^{n} \sim \mathcal{D}^{-}\) and \(\mathbf{X}^{+} = \{z^{+}_j\}_{j=1}^{m} \sim \mathcal{D}^{+}\). The continuous objective thus reduces to the discrete optimal transport problem, \[\min_{\mathcal{T} \in \,\Pi(\mathbf{a},\, \mathbf{b})} \left\langle \mathbf{C},\, \mathcal{T} \right\rangle_{\!F} , \label{eq:ot95discrete}\tag{1}\] where \(\mathbf{C}_{ij} = \|z^{-}_i - z^{+}_j\|^2\) is the squared Euclidean cost matrix, \(\mathbf{a} = \tfrac{1}{n}\mathbf{1}_n\) and \(\mathbf{b} = \tfrac{1}{m}\mathbf{1}_m\) are uniform marginal weights, and \[\Pi(\mathbf{a}, \mathbf{b}) \;=\; \Bigl\{ \mathcal{T} \in \mathbb{R}^{n \times m}_{+} \;\Big|\; \mathcal{T}\,\mathbf{1}_m = \mathbf{a},\; \mathcal{T}^{\!\top}\mathbf{1}_n = \mathbf{b} \Bigr\}\] is the transport polytope. To scale to the high-dimensional hidden representations encountered in VLA models, we solve Eq. 1 using the low-rank Sinkhorn algorithm [29], which factorises the transport plan as \[\mathcal{T} \;=\; \mathbf{Q}\; \operatorname{diag}(\mathbf{g}^{-1})\; \mathbf{R}^{\!\top} \label{eq:lowrank}\tag{2}\]

where \(\mathbf{Q} \in \mathbb{R}^{n \times r}_{+}\), \(\mathbf{R} \in \mathbb{R}^{m \times r}_{+}\), and \(\mathbf{g} \in \mathbb{R}^{r}_{+}\) are low-rank factors of rank \(r \ll \min(n,m)\).

Moreover, a regularization term is added to the discrete OT formulation \[\min_{\mathcal{T} \in \,\Pi(\mathbf{a},\, \mathbf{b})} \left\langle \mathbf{C},\, \mathcal{T} \right\rangle_{\!F} \;-\; \varepsilon \,H(\mathcal{T}), \label{eq:ot95regularised}\tag{3}\]

where \(H(\mathcal{T}) = -\sum_{ij} \mathcal{T}_{ij} \log \mathcal{T}_{ij}\) is the entropy of the transport plan and \(\varepsilon > 0\) controls the degree of regularisation. Smaller \(\varepsilon\) gives a sparser and more faithful approximation of the exact OT solution. We used the implementation of this algorithm in the Python Optimal Transport (POT) package [31].

8.2 Hyperparameters and design choices↩︎

The complete set of hyperparameters and design choices involved in DiMaS are given below in 1. The overall design is robust in the sense that all these choices are frozen for all our experiments, which includes all the various model, suite and steering task choices.

Note that the steering intervention is applied only at the output of a single layer in the action expert, but for all flow-matching steps. \(\alpha\) is the most influential choice in terms of balancing control strength with success rate. While \(\alpha=0.5\) tends to be a safe default choice, it could be tuned for each task individually for optimal balance. The only hyperparameter that depends on the underlying VLA model, and may not generalize to an arbitrary new base VLA, is the layer choice \(\ell\).Empirically, we find that late layers work well in general: we use the second-to-last layer, but observe little variation in steering effectiveness across the last few layers. We provide an analysis of this choice for \(\pi_{0.5}\) in 12. A more systematic study of layer selection, and of how the optimal layer depends on the base VLA, remains a valuable direction for future work.

Table 1: Hyperparameters of DiMaS.
Component Hyperparameter Value
Distributions \(\bmD^-, \bmD^+\) Lower quantile \(q^-\) \(0.25\)
Upper quantile \(q^+\) \(0.75\)
Distribution classifier \(g\) Model SVM (linear kernel)
Regularisation \(C\) \(0.1\)
Learning transport map \(\mathcal{T}_l^{(m)}\) Algorithm Low-rank Sinkhorn
Regularisation \(\varepsilon\) \(10^{-4}\)
Max.iterations \(5\,000\)
Rank \(r\) \(\min(n, m)\) (default)
Steering intervention Interpolation coefficient \(\alpha\) \(0.5\)
Layer \(\ell\) Second last
FM steps \(m\) All steps

6pt

8.3 Models and datasets↩︎

We evaluate our methods on two state-of-the-art VLA models. SmolVLA  [7] is a compact vision-language-action model built on a 256M-parameter SmolVLM backbone with a flow-matching action expert. We use the publicly available checkpoint lerobot/smolvla_libero, fine-tuned on the LIBERO benchmark. \(\pi_{0.5}\) [8] is a larger VLA model built on a PaliGemma-3B vision-language backbone combined with a 300M-parameter flow-matching action expert. We use the checkpoint lerobot/pi05-libero, fine-tuned on the LIBERO benchmark. Both models generate actions via iterative denoising over 10 flow-matching steps. One main difference between these two checkpoints is that the LIBERO-SmolVLA one is trained to predict a chunk of 50 at each timestep and use only the first one at inference time, whereas for \(\pi_{0.5}\) it is able to predict and execute 10 actions directly, which makes its trajectories smoother and faster to run.

To evaluate our proposed steering strategy we use LIBERO [30], a benchmark for lifelong robot learning comprising five task suites: LIBERO-Object, LIBERO-Spatial, LIBERO-Goal, LIBERO-10 (also referred to as LIBERO-Long), each containing 10 tasks, and LIBERO-90, which contains 90 short-horizon tasks, for a total of 130 manipulation tasks. We evaluate on four suites:

  • LIBERO-Object focuses on object-centric manipulation, requiring the robot to pick up and place various household objects.

  • LIBERO-Spatial introduces spatial reasoning, requiring the robot to reason about object relationships.

  • LIBERO-Goal contains goal-conditioned tasks with more diverse objectives, including articulated object manipulation.

  • LIBERO-10 (LIBERO-Long) consists of long-horizon tasks that require composing multiple sub-goals in sequence.

8.4 Computational Cost↩︎

Computing the OT transport map for a single suite (LIBERO-Object) across 50 training episodes represents approximately \(\sim\)​3,500 samples per distribution. Solving the Low-Rank Sinkhorn problem across all 10 flow-matching steps takes approximately 85 minutes. This cost is paid once offline and does not affect inference latency. Theoretically, the time complexity of the Low-Rank Sinkhorn algorithm is \(\mathcal{O}\bigl((n + m) \cdot r \cdot K\bigr)\), where \(n\) and \(m\) are the number of source and target samples, \(r\) is the rank of the factorisation, and \(K\) is the number of iterations.

Regarding inference-time intervention, all experiments were conducted on NVIDIA Tesla V100, A100, and H100 GPUs. Our steering mechanism is lightweight: the overhead (SVM classification + transport) averages only \(0.62 \pm 0.05\) ms per flow-matching step where an intervention occurs (calculated on \(n=98\) intervention steps). Even in the worst case, where the classifier triggers steering at all 10 flow-matching steps, this amounts to just \(6.2\) ms of overhead per robot timestep. Since each robot timestep corresponds to an action budget of approximately \(50\) ms, this worst-case overhead consumes only \({\sim}12\%\) of the available budget, leaving substantial headroom for real-time deployment.

9 Baseline Details↩︎

The representation-level baselines intervene either in the VLM backbone or in the action expert \(f_A\). Following prior work [28], the VLM interventions are applied at every layer of the backbone, using the token-averaged residual stream \(\bar{h}_\ell(x_t) = \mathbb{E}_p\!\left[h^p_\ell(x_t)\right] \in \mathbb{R}^{d_V}\) at each layer \(\ell\); the action-expert intervention acts at a single layer \(\ell\), using the per-token representations \(h^{p,m}_\ell(x_t) \in \mathbb{R}^{d_A}\), where \(m\) indexes the denoising step. In all cases the additive shift is added to every token position \(p\) (and, in the action expert, every denoising step \(m\)) at the intervened layer(s).

Following the main paper, we split representations into a source distribution \(\mathcal{D}^{-}\) (feature value below \(q_\tau\)) and a target distribution \(\mathcal{D}^{+}\) (feature value above \(q_{1-\tau}\)), referred to as the source and target tails of the feature distribution. We write \(h \in \mathbb{R}^d\) for a generic representation, with \(d \in \{d_V, d_A\}\), and all quantities below are fit and applied per intervened layer.

All representation-level baselines share the feature filter of the main paper: the shift \(\mathcal{T}(h)\) is applied only to representations predicted to lack the targeted feature, gated by the binary classifier \(g_\ell^{(m)}(h) \in \{0, 1\}\) that is active (\(g_\ell^{(m)}(h) = 1\)) when \(h\) is assigned to \(\mathcal{D}^{-}\). The effective update is therefore \[h \;\mapsto\; \begin{cases} \mathcal{T}(h) & \text{if } g_\ell^{(m)}(h) = 1, \\[4pt] h & \text{otherwise,} \end{cases}\] so that representations already exhibiting the feature are left unchanged. The baselines differ only in the choice of transport map \(\mathcal{T}\).

Mean-difference steering (VLM Mean). This baseline applies a constant additive shift equal to the difference of the tail means, \[\mathcal{T}(h) = h + (\mu^{+} - \mu^{-}), \qquad \mu^{\pm} = \mathbb{E}_{h \sim \mathcal{D}^{\pm}}[h],\] with \(\mu^{+}\) and \(\mu^{-}\) estimated over the target and source tails.

Regression-based steering (VLM Regression, FM Regression). This baseline replaces the constant shift with an adaptive, representation-dependent one, and is applied identically to the VLM and action-expert representations. We fit a ridge regressor on the representations to predict the scalar feature, \[\hat{\phi}_{\ell}(h) = w_\ell^\top h + b_\ell, \qquad w_\ell \in \mathbb{R}^{d},\]

and set a target value \(s^\star\) equal to a tail quantile of the feature distribution, \(s^\star \in \{q_\tau, q_{1-\tau}\}\), depending on the steering direction. The representation is then displaced by the minimum-norm perturbation that moves the predicted feature from \(\hat{s} = \hat{\phi}_{\ell}(h)\) to \(s^\star\), \[\mathcal{T}(h) = h + (s^\star - \hat{s})\,\frac{w_\ell}{\lVert w_\ell \rVert_2^{2}}.\]

VLM Regression intervenes on the token-averaged VLM representation at every layer, and FM Regression on the action-expert representation at a single layer.

Prompt steering. As a non-representational baseline, we steer the target feature through the language instruction, appending a directive to the task prompt rather than modifying any hidden representation. For speed we append “Do this quickly.” (low-to-high) or “Do this slowly.” (high-to-low), and for end-effector height “Do this at high height.” or “Do this at low height.

10 Additional experimental results↩︎

The main paper reports the generalization study for decreasing end-effector speed. Here we provide the remaining three feature and direction combinations across all four settings and both models: increasing speed (10), decreasing vertical displacement (11), and increasing vertical displacement (12).

Decreasing vertical displacement. This direction closely mirrors the speed-decrease result of the main paper (11). On both models, the intervention yields a significant downward shift at every setting on Object and Spatial (\(p<0.01\)), including transfer to held-out tasks (Setting 3) and, for Spatial, to a disjoint suite (Setting 4). As with speed, LIBERO-Goal is the harder case: the shift stays in the intended direction but reaches significance only at some settings (\(\pi_{0.5}\) Setting 3, \(p<0.01\); SmolVLA Setting 4). The strong effect on Object and Spatial comes with a larger success-rate cost than for speed, consistent with vertical displacement being more tightly coupled to the lifting and placing phases needed to complete the task.

The increasing directions are more constrained by the task. Steering the two increasing directions is more demanding than the decreasing ones. For increasing speed (10), \(\pi_{0.5}\) shifts the feature in the intended direction throughout and is significant at the intermediate settings; the effect weakens only under the largest shift (Setting 4 on Spatial and Goal, \(p=0.494\) and \(p=0.567\)), and on SmolVLA it holds in a subset of cells. Increasing vertical displacement (12) is the most constrained combination, with a significant effect concentrated in fewer cells (e.g.\(\pi_{0.5}\) Object, Setting 1, \(p<0.01\)). We read this asymmetry as reflecting the policy’s default behavior: successful trajectories naturally decelerate and lower the end-effector to grasp and place, so steering that reinforces these tendencies (slower, lower) aligns with the direction the representations already encode, whereas steering the opposite way (faster, higher) pushes against it. Steering is thus most effective when it amplifies a behavior the policy is already inclined toward, rather than reversing it.

a
b

c

Figure 10: Steering across evaluation levels on \(\pi_{0.5}\) (left) and SmolVLA (right). We compare the unsteered baseline against DiMaS for increasing speed. Panels show, left to right, the baseline and the settings discussed in Section 4.2; columns are the three LIBERO suites (Object, Spatial, Goal). Box plots show the per-episode feature distribution at each level, annotated with success rate and \(p\)-value vs.baseline; box opacity is higher when the shift is statistically significant (\(p<0.01\)). Arrow length is proportional to the induced \(\Delta\)mean.. a — \(\pi_{0.5}\), b — SmolVLA

a
b

c

Figure 11: Steering across evaluation levels on \(\pi_{0.5}\) (left) and SmolVLA (right). We compare the unsteered baseline against DiMaS for decreasing vertical displacement. Panels show, left to right, the baseline and the settings discussed in Section 4.2; columns are the three LIBERO suites (Object, Spatial, Goal). Box plots show the per-episode feature distribution at each level, annotated with success rate and \(p\)-value vs.baseline; box opacity is higher when the shift is statistically significant (\(p<0.01\)). Arrow length is proportional to the induced \(\Delta\)mean.. a — \(\pi_{0.5}\), b — SmolVLA

a
b

c

Figure 12: Steering across evaluation levels on \(\pi_{0.5}\) (left) and SmolVLA (right). We compare the unsteered baseline against DiMaS for increasing vertical displacement. Panels show, left to right, the baseline and the settings discussed in Section 4.2; columns are the three LIBERO suites (Object, Spatial, Goal). Box plots show the per-episode feature distribution at each level, annotated with success rate and \(p\)-value vs.baseline; box opacity is higher when the shift is statistically significant (\(p<0.01\)). Arrow length is proportional to the induced \(\Delta\)mean.. a — \(\pi_{0.5}\), b — SmolVLA

11 Additional DiMaS analysis↩︎

We report additional results complementing the main paper: speed density plots for the \(\alpha\) ablation across additional task suites for both \(\pi_{0.5}\) and SmolVLA in 11.1, and examples of end-effector height under increasing and decreasing vertical-displacement steering in 11.2.

11.1 Interpolation factor: per-suite density↩︎

Figure 13 highlights the clear distribution shift introduced by DiMaS for speed H\(\to\)L steering. Across all three suites, increasing \(\alpha\) consistently shifts the distribution towards lower speed values. However, the success rate also drops gradually for \(\alpha \in \{0.3, 0.5\}\) and more sharply for \(\alpha \in \{0.7, 1.0\}\).

a
b

Figure 13: Effect of the interpolation coefficient \(\alpha\) on speed and task success. Speed distributions when steering towards lower speeds for \(\pi_{0.5}\) (top) and SmolVLA (bottom). Different curves show the resulting speed distributions when changing the interpolation factor of DiMaS.. a — \(\pi_{0.5}\), b — SmolVLA

11.2 Additional qualitative results↩︎

Figure 14: End-effector height trajectories under DiMaS height steering on SmolVLA. Evolution of the observed EEF z position over time for two successful episodes of LIBERO-Object tasks (task 0 and task 6). Curves show the unsteered baseline (black), H\toL steering (blue), and L\toH steering (red).

Figure 14 shows the evolution of the vertical end-effector position over time for two episodes of the LIBERO-Object suite, illustrating the effect of DiMaS height steering on the trajectories. On the left, the red curve (L\(\to\)H) is consistently higher than the baseline during the transport phase, while the blue curve (H\(\to\)L) remains close to or below it. On the right, H\(\to\)L produces a noticeably lower trajectory while L\(\to\)H follows the baseline closely. This suggests that the steering effect is task-dependent.

12 Steering-layer ablation↩︎

Figure 15 shows that the choice of steering layer strongly determines the induced speed shift and, to a lesser extent, the preserved success rate. For decreasing speed (top row), steering at later layers (L15–L17) most effectively concentrates mass at lower speeds, moving the distribution below the baseline while largely preserving task success. Earlier and middle layers (L3–L9) are markedly less effective: they shift mass toward higher speeds, opposite to the intended direction. The accompanying drop in success rate (e.g.SR \(=76\%\) on Spatial at L3) partly reflects this loss of control, but also the faster execution itself, since moving more quickly leaves less margin to grasp and place the object reliably. The same layer dependence holds for increasing speed (bottom row), where later layers again produce the cleanest shift in the intended direction. Across both directions and all three suites, the behavioral feature is most cleanly and controllably represented in the later layers of the network, which motivates our choice of a late intervention layer in the main experiments.

a
b

Figure 15: Effect of steering layer on speed and task success. Results for \(\pi_{0.5}\) across the three LIBERO suites. Each curve shows the resulting speed distribution when steering at a different layer (L3–L17), with the corresponding success rate (SR) in the legend; the shaded region is the unsteered baseline.. a — Steering to decrease speed (H\(\to\)L)., b — Steering to increase speed (L\(\to\)H).

References↩︎

[1]
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. Language models are few-shot learners. Advances in Neural Information Processing Systems (NeurIPS), 33: 1877–1901, 2020.
[2]
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971, 2023.
[3]
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervision. In International conference on machine learning, pages 8748–8763. PMLR, 2021.
[4]
Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katherine Millican, Malcolm Reynolds, et al. Flamingo: a visual language model for few-shot learning. Advances in Neural Information Processing Systems (NeurIPS), 35: 23716–23736, 2022.
[5]
Brianna Zitkovich, Tianhe Yu, Sichun Xu, Peng Xu, Ted Xiao, Fei Xia, Jialin Wu, Paul Wohlhart, Stefan Welker, Ayzaan Wahid, et al. Rt-2: Vision-language-action models transfer web knowledge to robotic control. In Conference on Robot Learning, pages 2165–2183. PMLR, 2023.
[6]
Moo Jin Kim, Karl Pertsch, Siddharth Karamcheti, Ted Xiao, Ashwin Balakrishna, Suraj Nair, Rafael Rafailov, Ethan Foster, Grace Lam, Pannag Sanketi, et al. Openvla: An open-source vision-language-action model. arXiv preprint arXiv:2406.09246, 2024.
[7]
Mustafa Shukor, Dana Aubakirova, Francesco Capuano, Pepijn Kooijmans, Steven Palma, Adil Zouitine, Michel Aractingi, Caroline Pascal, Martino Russi, Andres Marafioti, et al. Smolvla: A vision-language-action model for affordable and efficient robotics. arXiv preprint arXiv:2506.01844, 2025.
[8]
Physical Intelligence, Kevin Black, Noah Brown, James Darpinian, Karan Dhabalia, Danny Driess, Adnan Esmail, Michael Equi, Chelsea Finn, Niccolo Fusai, et al. : A vision-language-action model with open-world generalization. arXiv preprint arXiv:2504.16054, 2025.
[9]
Haoquan Fang, Jiafei Duan, Donovan Clay, Sam Wang, Shuo Liu, Weikai Huang, Xiang Fan, Wei-Chuan Tsai, Shirui Chen, Yi Ru Wang, et al. Molmoact2: Action reasoning models for real-world deployment. arXiv preprint arXiv:2605.02881, 2026.
[10]
Yaron Lipman, Ricky TQ Chen, Heli Ben-Hamu, Maximilian Nickel, and Matt Le. Flow matching for generative modeling. arXiv preprint arXiv:2210.02747, 2022.
[11]
Alexander Matt Turner, Lisa Thiergart, Gavin Leech, David Udell, Juan J Vazquez, Ulisse Mini, and Monte MacDiarmid. Steering language models with activation engineering. arXiv preprint arXiv:2308.10248, 2023.
[12]
Andy Zou, Long Phan, Sarah Chen, James Campbell, Phillip Guo, Richard Ren, Alexander Pan, Xuwang Yin, Mantas Mazeika, Ann-Kathrin Dombrowski, et al. Representation engineering: A top-down approach to ai transparency. arXiv preprint arXiv:2310.01405, 2023.
[13]
Pegah Khayatan, Mustafa Shukor, Jayneel Parekh, Arnaud Dapogny, and Matthieu Cord. Analyzing fine-tuning representation shift for multimodal llms steering alignment. International Conference on Computer Vision, 2025.
[14]
Jayneel Parekh, Pegah Khayatan, Mustafa Shukor, Arnaud Dapogny, Alasdair Newson, and Matthieu Cord. Learning to steer: Input-dependent steering for multimodal llms. Advances in Neural Information Processing Systems, 38: 159799–159834, 2026.
[15]
Diederik P Kingma and Max Welling. Auto-encoding variational bayes. arXiv preprint arXiv:1312.6114, 2013.
[16]
Yujun Shen, Ceyuan Yang, Xiaoou Tang, and Bolei Zhou. Interfacegan: Interpreting the disentangled face representation learned by gans. IEEE transactions on pattern analysis and machine intelligence, 44 (4): 2004–2018, 2020.
[17]
Nina Panickssery, Nick Gabrieli, Julian Schulz, Meg Tong, Evan Hubinger, and Alexander Matt Turner. Steering llama 2 via contrastive activation addition. arXiv preprint arXiv:2312.06681, 2023.
[18]
Andy Arditi, Oscar Obeso, Aaquib Syed, Daniel Paleka, Nina Panickssery, Wes Gurnee, and Neel Nanda. Refusal in language models is mediated by a single direction. arXiv preprint arXiv:2406.11717, 2024.
[19]
Xinchi Qiu, Lei Yu, Yuchen Zhang, Aobo Yang, Narine Kokhlikyan, Nicola Cancedda, Diego Garcia-Olano, et al. Hallucination reduction with casal: Contrastive activation steering for amortized learning. arXiv preprint arXiv:2510.02324, 2025.
[20]
Mateusz Pach, Shyamgopal Karthik, Quentin Bouniot, Serge Belongie, and Zeynep Akata. Sparse autoencoders learn monosemantic features in vision-language models. Advances in Neural Information Processing Systems, 38: 95706–95742, 2026.
[21]
Kiho Park, Yo Joong Choe, and Victor Veitch. The linear representation hypothesis and the geometry of large language models. arXiv preprint arXiv:2311.03658, 2023.
[22]
Nelson Elhage, Tristan Hume, Catherine Olsson, Nicholas Schiefer, Tom Henighan, Shauna Kravec, Zac Hatfield-Dodds, Robert Lasenby, Dawn Drain, Carol Chen, et al. Toy models of superposition. arXiv preprint arXiv:2209.10652, 2022.
[23]
Josh Engels, Eric Michaud, Isaac Liao, Wes Gurnee, and Max Tegmark. Not all language model features are one-dimensionally linear. In International Conference on Learning Representations, volume 2025, pages 84591–84622, 2025.
[24]
Thomas Fel, Binxu Wang, Michael A Lepori, Matthew Kowal, Andrew Lee, Randall Balestriero, Sonia Joseph, Ekdeep S Lubana, Talia Konkle, Demba Ba, et al. Into the rabbit hull: From task-relevant concepts in dino to minkowski geometry. arXiv preprint arXiv:2510.08638, 2025.
[25]
Usha Bhalla, Thomas Fel, Can Rager, Sheridan Feucht, Tal Haklay, Daniel Wurgaft, Siddharth Boppana, Matthew Kowal, Vasudev Shyam, Jack Merullo, et al. Do sparse autoencoders capture concept manifolds? arXiv preprint arXiv:2604.28119, 2026.
[26]
Pau Rodriguez, Arno Blaas, Michal Klein, Luca Zappella, Nicholas Apostoloff, Xavier Suau, et al. Controlling language and diffusion models by transporting activations. In International Conference on Learning Representations, volume 2025, pages 89812–89855, 2025.
[27]
Bear Häon, Kaylene Stocking, Ian Chuang, and Claire Tomlin. Mechanistic interpretability for steering vision-language-action models. Conference on Robot Learning (CoRL), 2025.
[28]
Hugo Buurmeijer, Carmen Amo Alonso, Aiden Swann, and Marco Pavone. Observing and controlling features in vision-language-action models. arXiv preprint arXiv:2603.05487, 2026.
[29]
Meyer Scetbon, Marco Cuturi, and Gabriel Peyré. Low-rank sinkhorn factorization, 2021.
[30]
Bo Liu, Yifeng Zhu, Chongkai Gao, Yihao Feng, Qiang Liu, Yuke Zhu, and Peter Stone. Libero: Benchmarking knowledge transfer for lifelong robot learning. Advances in Neural Information Processing Systems, 36: 44776–44791, 2023.
[31]
Rémi Flamary, Nicolas Courty, Alexandre Gramfort, Mokhtar Z Alaya, Aurélie Boisbunon, Stanislas Chambon, Laetitia Chapel, Adrien Corenflos, Kilian Fatras, Nemo Fournier, et al. Pot: Python optimal transport. Journal of Machine Learning Research, 22 (78): 1–8, 2021.

  1. Github page: https://github.com/pegah-kh/dimas↩︎

  2. Blog/Project page: https://pegah-kh.github.io/dimas/↩︎