Robustness of Robotic Manipulation:
Foundations and Frontiers


Abstract

Humans and animals exhibit remarkable robustness in physical manipulation, yet robots remain far behind. Progress toward human-level manipulation robustness is hindered by the absence of a unified and systematic understanding: different subfields frame robustness in distinct ways, often leaving the concept ambiguous and limiting deeper analysis as well as communication across research areas. This paper presents a systematic study of manipulation robustness. We begin with a formal definition, characterizing robustness as the degree to which a manipulation system can achieve its goal in the presence of uncertainty and variation. Building on this definition, we introduce general formulations of manipulation robustness from probabilistic and control-theoretic perspectives. We then synthesize the guiding principles and concrete mechanisms of manipulation robustness across perception, planning, control, policy learning, and hardware, illustrating each mechanism through representative works, including foundational and recent studies. In addition, we revisit existing metrics and evaluation methods for quantifying manipulation robustness. Finally, we distill broader lessons for designing robust manipulation systems and discuss open problems and future directions toward achieving human-level robustness in robotic manipulation.

1 Introduction↩︎

Manipulation in the real world is inherently uncertain [1]. Object pose may be partially observed, friction and compliance may be poorly modeled, contact events are discontinuous and difficult to predict, and task instances may vary substantially across episodes. Robust manipulation is therefore not simply a matter of executing a nominal plan accurately in an idealized environment. Instead, it is the ability to achieve task goals despite such uncertainty and variation.

Biological systems provide compelling evidence that robust manipulation is possible. Humans often handle these uncertainties and variations without explicit deliberation, relying instead on sensorimotor control strategies that combine prediction with rapid feedback correction [2]. For instance, humans regulate grip forces to prevent objects from slipping in the hand. When grasping objects such as fruits (Fig. 1), humans apply grip forces with a safety margin above the expected slip threshold. Under varying or dynamic conditions, this margin is adjusted accordingly [3], preventing grasp failure even in uncertain environments.

Robust behaviors are also generally observed in animals. Sea otters, for example, juggle pebbles on their bellies [4]. They maintain this behavior despite turbulent disturbances and across wide variations in object shape, mass, and body posture. Even far less complex organisms exhibit robust interaction with their environment. For example, coordinated ciliary motion in paramecium generates fluid flows that draw and capture food particles [5]. Without complex visual perception, the organism can handle uncertainty in particle location using only mechanosensory cilia in highly dynamic fluid environments. These examples suggest that robustness in physical interaction is a fundamental property of biological systems [6].

Figure 1: Overview of the robustness of robotic manipulation, illustrated using a pear grasping task. Robotic manipulation faces three core challenges: epistemic uncertainty, aleatoric uncertainty, and episodic variation. The paper is organized around two robustness principles: uncertainty and variation regulation, which concerns reducing epistemic uncertainty or tolerating uncertainty and variation, and failure management, which concerns preventing failure, recovering from temporary failure, or tolerating local failure. We discuss robustness mechanisms across five submodules of a manipulation system: perception, planning, control, policy learning, and hardware. Several icons or illustrations are original and conceptually inspired by [7]–[12].

In contrast, today’s robotic systems manipulate objects with a level of robustness still below that of humans and many animals [1], [13]. Even routine behaviors such as pick-and-place, which humans perform effortlessly, remain challenging for robots to execute reliably in the presence of contact uncertainty [14]. Although impressive demonstrations from academia and industry have shown remarkable capabilities under controlled conditions, many of these systems degrade significantly when deployed outside carefully engineered environments. Ultimately, robotic manipulation must operate in open, time-varying environments characterized by significant uncertainty.

This evident gap between the robustness of robotic and human manipulation motivates this paper. A major obstacle to progress is the absence of a deep understanding and overarching framework for manipulation robustness. Each subfield of robotics discusses robustness in its own terms, often restricted to a narrow domain and leaving the concept implicit or underspecified [15][17]. To address this issue, we propose a shared conceptual language and a taxonomy to connect existing research threads and enable the accumulation and accessibility of insights across subfields. We also identify several foundational but unanswered questions: What explains the extraordinary robustness observed in human and animal manipulation? What mechanisms and principles enable robust manipulation? How should robustness be quantified and measured? This paper calls attention to these questions and encourages the field to pursue a deeper and more systematic understanding of robust manipulation skills.

Robust manipulation skills appear not to arise from a single mechanism, but rather from the coordination of mechanisms across perception, planning, control, policy learning, and hardware, calling for joint research across subfields [13]. This perspective distinguishes the present paper from prior reviews on robotic manipulation, such as data-driven grasp synthesis [18], robot learning for manipulation [19], foundation models for robotics [20], etc. While these works have made significant contributions to broad areas of manipulation, robustness is typically treated only incidentally rather than as the central object of study. To our knowledge, there has not yet been a comprehensive review that systematically examines the principles, mechanisms, and evaluation methods underlying robust manipulation under real-world uncertainty and variation.

This paper makes the following contributions: (i) it proposes a task-centered definition of manipulation robustness; (ii) it formulates manipulation robustness from both probabilistic and control-theoretic perspectives; (iii) it organizes robustness mechanisms across perception, planning, control, policy learning, and hardware under a set of guiding principles; and (iv) it synthesizes evaluation protocols and open challenges for future research.

In the following, we begin by introducing a definition (Section 2) and a mathematical formulation (Section 3) of manipulation robustness. We then present robustness principles (Section 4), followed by a systematic summary of existing robustness mechanisms (Section 5). Next, we analyze current metrics and evaluation protocols (Section 6). We then synthesize key insights and observations (Section 7). Finally, we highlight open challenges and suggest future directions toward achieving human-level manipulation robustness (Section 8).

2 Definition of Manipulation Robustness↩︎

This section introduces the concept of robustness progressively. We begin with general definitions of robustness across disciplines, narrow the scope to robotics, and finally formalize the notion of manipulation robustness.

2.1 Robustness Definitions↩︎

Despite its frequent use, the notion of robustness is often interpreted differently across contexts and disciplines. Several domain-specific efforts have attempted to formalize robustness in biology [6], machine learning [21], and robotics [22]. We provide representative definitions below for illustration.

Definition 1 (Biological robustness). Robustness is a property that allows a system to maintain its functions despite external and internal perturbations [6].

Definition 2 (Machine learning robustness). The robustness of machine learning models denotes the capacity of a model to sustain stable predictive performance in the face of variations and changes in the input data [21].

Definition 3 (Robotic robustness). Robustness specifies the degree to which a behavior can fulfill its task despite being challenged by an adversity [22].

Synthesizing these perspectives reveals a common theme: robustness is inherently context-dependent. It is not an absolute property of a system, but a relational one. It is defined only with respect to what the system is intended to achieve and the conditions under which it operates. Without context, claims of robustness become ill-posed and difficult to compare across domains or methods.

To make this dependence explicit, we identify four foundational dimensions that together define the context of robustness: goal, challenge, mechanism, and evaluation. With these dimensions, we define robustness as follows:

Definition 4 (Robustness). Robustness refers to a system’s ability to achieve a given goal under specified challenges, enabled by particular mechanisms, and assessed through quantitative evaluation.

This structure applies broadly across disciplines. For example, in machine learning, robustness typically refers to maintaining predictive performance (goal[21] under distribution shifts or adversarial perturbations (challenge[23], achieved through techniques such as adversarial training (mechanism), and evaluated using quantitative metrics such as worst-case accuracy, or performance degradation under controlled perturbations (evaluation[24].

2.2 Robustness in Robotics↩︎

What distinguishes robotics is the embodied physical interactions and tight integration of subsystems. Goals: Because robotic systems integrate sensing, mechanics, planning, control, actuation, and learning, robustness must often be addressed both within and across the components. Thus, goals may be defined at different levels of abstraction, ranging from high-level task completion (e.g., object transport or assembly) to component-level objectives such as state estimation accuracy, collision avoidance, or trajectory tracking performance. Challenges: Direct interaction with the physical world exposes robots to uncertainty in perception and contact dynamics, sensing noise and external disturbances, as well as variations in system parameters and environmental conditions. Mechanisms: Robustness is realized through complementary passive and active strategies spanning hardware and algorithms. Passive strategies, such as compliant hardware and soft materials, physically tolerate disturbances. Active strategies, such as motion planning, feedback control, and learning methods, proactively mitigate model mismatch or compensate for noise or disturbances. Evaluation: Robustness in robotics is commonly evaluated on task execution success rates, complemented by more fine-grained metrics such as stability measure, control convergence, or theoretical guarantees, defined at the level of individual tasks or system components.

2.3 Robotic Manipulation Robustness↩︎

Manipulation is characterized as “an agent’s control of its environment through selective contact” [1], which makes manipulation robustness a distinct subclass of robotic robustness, with unique challenges arising from contact interactions among the robot, manipulated objects, and the environment.

Manipulation is dominated by multi-body contact interactions, which induce dynamics that are high-dimensional, hybrid, discontinuous, constrained, and therefore difficult to model accurately. These interaction dynamics constitute a major challenge, as small variations in geometry, contact conditions, or material properties can lead to qualitatively different outcomes. Perception introduces additional challenges in manipulation because effective contact interaction depends on estimating interaction-relevant states that are largely external to the robot, such as object pose, insertion depth, local contact geometry, material properties, and deformation. Many of these quantities are only partially observable, visually ambiguous, or occluded during contact. Therefore, it has substantially higher demands on perception than tasks such as locomotion, where the robot primarily regulates its own state. Together, these factors complicate the design of robust mechanisms in manipulation planning, control, and learning.

Moreover, real-world manipulation encompasses a broad spectrum of task specifications, in which robots must achieve diverse objectives across domestic, industrial, and open-world environments. Unlike locomotion, manipulation relies on multi-body contact interactions among the robot, manipulated objects, and the environment that actively change the state of the world, leading to a combinatorial diversity of manipulation goals. Finally, the discontinuous and task-dependent nature of contact interactions makes the definition of general evaluation criteria and the establishment of formal robustness guarantees particularly difficult.

We refine Definition 4 for manipulation by specifying the four foundational dimensions as follows:

  • Goals: Robotic manipulation robustness is defined relative to manipulation goals, which involve achieving desired object states through contact interactions, potentially within task-specific tolerance levels.

  • Challenges: Challenges arise from uncertainty and variation, which we categorize into epistemic uncertainty, aleatoric uncertainty, and episodic variation.

  • Mechanisms: Robustness mechanisms improve task goal achievement by tolerating or mitigating uncertainty and variation in contact interactions, preventing failures, or enabling recovery when failures occur.

  • Evaluation: Evaluation methods define quantitative criteria for measuring manipulation performance and robustness, enabling principled comparison across approaches.

Among these dimensions, challenges play a central role. The importance of uncertainty has been recognized since the earliest days of robotics in the 1950s [25], which remains a central theme in robotics research [26]. Specifically, epistemic uncertainty stems from incomplete or inaccurate knowledge of system properties and can, in principle, be reduced through additional data or modeling. Aleatoric uncertainty reflects irreducible randomness, such as observation noise or stochastic transition disturbances. Episodic variation refers to systematic differences in the physical world that remain fixed within a single task execution but change across task episodes, such as variations in object identity, environment configuration, or robot embodiment [27]. An intuitive analogy is that of a robot searching for a light switch in a dark room. Episodic variation: across different rooms, the position or shape of the switch may vary. Epistemic uncertainty: within a given room, the initial belief about the switch location is uncertain but can be refined through interaction, such as touching the wall. Aleatoric uncertainty: even with perfect knowledge, however, randomness such as sensor noise, slips of the hand, or minor actuation errors may still occur.

3 Formulation of Manipulation Robustness↩︎

Given the definition in Section 2, we formulate the general problem of manipulation robustness using a unified partially observable stochastic control model. We then instantiate this model through probabilistic and control-theoretic perspectives, respectively. The unified model serves as an analytical framework to structure the discussion in the remainder of the paper. It is intended as a conceptual tool for analysis rather than a prescriptive model; it does not imply that a robot must explicitly represent these processes to achieve robust manipulation.

3.1 Unified Formulation↩︎

We formalize a manipulation episode as a discrete-time control process over a finite horizon \(T\), denoted by the sequence \(\{(x_t, u_t, y_t)\}_{t=0}^T\). Here, \(x_t \in \mathcal{X}\) is the system state (encompassing both robot and object generalized coordinates), \(u_t \in \mathcal{U}\) is the control input, and \(y_t \in \mathcal{Y}\) is the partial observation. Note that we present here a time-discretized model, though in principle a continuous-time formulation can also be adopted.

The evolution of this system is driven by underlying deterministic functions subjected to noise: \[\begin{align} \label{eq:unified95dynamics} x_{t+1} &= f(x_t, u_t, w_t ; \theta), \\ y_t &= g(x_t, v_t ; \theta). \end{align}\tag{1}\] In this framework, manipulation robustness is defined by how well a system achieves a target goal \(\mathcal{G}\) despite three fundamental challenges:

Aleatoric Uncertainty

The variables \(w_t \in \mathcal{W}\) and \({v_t \in \mathcal{V}}\) represent exogenous process and measurement noise. This captures inherent, irreducible stochasticity in the physical world, such as thermal sensor noise, unmodeled micro-impacts, or random slip during contact.

Epistemic Uncertainty

The parameter \(\theta \in \Theta\) dictates the true physical and perceptual properties of the task, such as object mass, surface friction, and camera calibration. However, the robot rarely has perfect knowledge of \(\theta\) or the true state \(x_t\). It operates using internal estimates \(\hat{\theta}\) and \(\hat{x}_t\). The discrepancy between reality (\(\theta, x_t\)) and the robot’s belief (\(\hat{\theta}, \hat{x}_t\)) constitutes epistemic uncertainty. Unlike aleatoric noise, this uncertainty can theoretically be reduced through exploration or learning.

Episodic Variation

While epistemic uncertainty concerns what the robot does not know, episodic variation describes the objective physical differences between separate executions of a task. We define an episodic variation \(\nu\) as a specific draw of the environment parameters, initial state, and goal region: \[\nu = (\theta, x_0, \mathcal{G}) \in \mathcal{N}, \label{eq:variation-def}\tag{2}\] where \(\mathcal{N}\) denotes the space of all possible task episodes. Note that the goal \(\mathcal{G} \subset \mathcal{X}\) represents a desired set of states, especially the desired state of the manipulated object. This compact notation covers many terminal manipulation objectives, but it does not by itself encode temporal ordering among subgoals. Ordered long-horizon tasks, such as opening a door, entering a room, and then closing the door, would require a richer formulation, e.g., a sequence of goal sets \((\mathcal{G}_1,\ldots,\mathcal{G}_K)\) or a temporal task specification.

Table 1: Overview of the manipulation robustness mechanisms (Learning mechanisms in Table [tbl:tab:learning95robustness95mechanisms]).
Submodule Mechanism Principle Section
Perception Active and Interactive Perception Uncertainty Reduction [sec:sec:perception:interactive]
Invariance in Perceptual Representations Uncertainty & Variation Toleration [sec:sec:perception:invariance]
Multimodality Local Failure Toleration [sec:sec:perception:multimodality]
Planning Reasoning over Uncertainties Uncertainty Reduction [sec:subsec:planning-uncertainty]
Motion Strategy and Funneling Uncertainty Reduction [sec:subsec:planning-funnel]
Robustness Margin Uncertainty & Variation Toleration [sec:subsec:planning-manifold]
Closure Uncertainty & Variation Toleration [sec:subsec:planning-closure]
Control Predictivity Failure Prevention [sec:control:predictivity]
Active Compliance Uncertainty & Variation Toleration [sec:control:compliance]
Control Invariance Uncertainty & Variation Toleration [sec:control:invariance]
Hardware Passive Compliance Uncertainty & Variation Toleration [sec:mechanical:passive-compliance]
Adhesion Uncertainty & Variation Toleration [sec:mechanical:adhesion]
Morphology Adaptation Temporary Failure Recovery [sec:mechanical:morphology-adaptation]

3.2 Probabilistic View↩︎

Learning-based and probabilistic planning methods treat the noise variables (\(w_t, v_t\)) and episodic variations (\(\nu\)) as probability distributions. Consequently, the dynamics and observation models are viewed as stochastic transitions, \[\begin{align} x_{t+1} \sim \mathcal{T}(\cdot \mid x_t, u_t; \theta), \;y_t \sim \mathcal{O}(\cdot \mid x_t; \theta). \end{align}\] Robustness in this paradigm is typically formulated as a Partially Observable Markov Decision Process (POMDP) and evaluated as the expected performance of a policy \(\pi\) across a distribution of episodes \(\rho_\mathcal{N}\): \[\Gamma(\pi) = \mathbb{E}_{\nu \sim \rho_\mathcal{N}} \bigl[ J_{\nu}(\pi) \bigr],\] where \(J_\nu(\pi)\) is the expected return of the trajectory under the true parameters \(\theta\), despite the policy acting on its imperfect belief \(\hat{\theta}\). A policy is defined as a mapping \(\pi : \mathcal{B} \to \mathcal{U}\) from beliefs to actions, where a belief \(b_t \in \mathcal{B}(\mathcal{X})\) is a probability distribution over the state space \(\mathcal{X}\) conditioned on the history of observations and actions.

3.3 Control-Theoretic View↩︎

In contrast, we follow robust control frameworks (e.g., \(H_\infty\) or Robust MPC) and model aleatoric noise \(w_t, v_t\) and episodic variations \(\nu\) as unknown, deterministic variables confined within bounded sets \(\mathcal{W}, \mathcal{V}\), and \(\mathcal{N}\). Robustness here is framed as a min-max optimization problem. The goal is to synthesize a policy that guarantees constraint satisfaction and minimizes a cost function \(c(x_t, u_t)\) against the worst-case possible realizations of the uncertainties: \[\min_{\pi} \max_{\nu \in \mathcal{N}, w_t \in \mathcal{W}, v_t \in \mathcal{V}} \sum_{t=0}^{T} c(x_t, u_t). \label{eq:robust-policy-maxmin}\tag{3}\] While the probabilistic view maximizes average reliability across a distribution of environments, the control-theoretic view provides strict performance and safety guarantees in a worst-case robustness sense.

4 Principles of Robust Manipulation↩︎

There are diverse mechanisms for achieving robustness in robotic manipulation, spanning perception, planning, control, learning, and hardware. While these mechanisms are often studied separately, many can be understood through several recurring principles. At a high level, robustness mechanisms can be organized along two axes. The first concerns how systems cope with uncertainty and variation: either by reducing epistemic uncertainty or by tolerating uncertainty and variation. The second concerns failure management: whether robustness is achieved primarily through preventing failures from occurring, recovering from temporary failures or tolerating local failures.

4.1 Uncertainty and Variation Regulation↩︎

One route to robustness is to reduce epistemic uncertainty by improving the robot’s knowledge of the environment, task, or system state. This principle underlies active and interactive perception (Section 5.1.1), belief-space planning methods that explicitly reason over uncertainty (5.2.1), and world-model-based learning approaches that learn predictive representations of environment dynamics (5.4.2).

In contrast, a second route is based on tolerating the challenges: it accepts that aleatoric uncertainty and episodic variation cannot be eliminated, and instead seeks to make task execution insensitive to them. This principle is perhaps more ubiquitously employed in robotic manipulation. One common strategy is compliance, realized both actively through control (Section 5.3.1) and passively through hardware and materials (5.5.1), which reduces sensitivity to contact uncertainty and modeling errors. Another is invariance, where behaviors or representations are designed to remain unaffected by uncertainty or variation, as exemplified by closure-based grasping (5.2.4), invariant perceptual representations (5.1.2), and control invariance (5.3.2). Robustness can also arise from mechanics-driven strategies that exploit task and environmental structure, such as motion strategies and funnels (5.2.2). Toleration also appears in learning-based methods that generalize across scenes, tasks, and environments through data diversity or policy design (5.4.1, 5.4.2), as well as in hardware mechanisms such as adhesion (5.5.2), which leverage favorable contact physics to tolerate uncertainty and variation.

These two strategies are not mutually exclusive. Robust manipulation systems typically combine them, reducing uncertainty where possible while remaining effective in the presence of irreducible uncertainty.

4.2 Failure Management↩︎

Along the second axis, robustness can be understood through failure management. Here, two complementary strategies emerge: failure prevention, which seeks to avoid failure a priori, and failure recovery or tolerance, which assumes that failures may occur and emphasizes recovery from them or tolerance of local failure. Tasks involving safety-critical operations or irreversible consequences (e.g., an object’s fragility) often demand preventative mechanisms, as failures are unacceptable. In contrast, for tasks where temporary failures can be tolerated, such as regrasping a rigid object after a drop, recovery mechanisms may be sufficient. Local failure tolerance further complements these strategies by allowing failures confined to specific components, sensing modalities, or subtasks to be absorbed without causing overall task failure.

Failure prevention aims to avoid entering failure-prone states altogether. This principle appears in a variety of robustness mechanisms. Examples include predictive control methods that down-weight risky rollouts in receding-horizon optimization (Section 5.3.3), closure-based grasp synthesis that remains stable under disturbance wrenches (5.2.4), and compliance-based control and hardware design (5.3.1, 5.5.1), which reduce the likelihood of failure arising from contact modeling errors. Many learning-based manipulation policies similarly emphasize failure avoidance, either by training predominantly on successful demonstrations or by incorporating conservative and safety-aware objectives that discourage risky behaviors and constraint violations (5.4.3).

By contrast, a complementary strategy focuses on restoring task execution once failure has occurred. From this perspective, failures need not be catastrophic; they may be temporary and recoverable, or local and confined to particular components. Recent learning-based systems have demonstrated the ability to acquire recovery behaviors from imperfect or failed demonstrations (Section 5.4.1), or to explicitly learn recovery policies that resume task execution following failure. Temporary failures may also arise from distribution shifts or changing task conditions, motivating adaptation mechanisms, either through morphological adaptation (5.5.3) or learning-based adaptation (5.4.4).

Local system failures, on the other hand, refer to failures confined to a particular component, sensing modality, actuator, module, or subtask, without immediately implying failure of the overall manipulation task. Such failures can be mitigated through modularity and redundancy. Modular robotic systems [28], [29] can localize failure and prevent its propagation to the rest of the system. Redundancy provides multiple pathways to task success and allows degradation in one component to be compensated for by others. Examples include multi-sensor fusion (Section 5.1.3), dual end-effectors [30], and alternative task-level plans [31]. Both modularity and redundancy are also common principles in biological systems; for example, ants dynamically reassign roles during cooperative transport, using redundant individuals to maintain collective performance despite local failures [32].

Guided by these two axes of robustness principles—uncertainty and variation regulation, and failure management—we next revisit the mechanisms that enable robust manipulation.

5 Mechanisms of Robust Manipulation↩︎

In this section, we revisit existing approaches to robustness in robotic manipulation. We organize the literature into five submodules of a robotic system—perception, planning, control, learning, and hardware (Fig. 1)—and summarize the mechanisms in Table 1 and Table 2. Within each submodule, we discuss the mechanisms through the lens of the robustness principles introduced in Section 4, while noting that each mechanism is listed under its dominant principle even though it may reflect multiple principles in practice. We also highlight the challenges they address, representative works, and illustrative manipulation tasks. Owing to the rapid recent progress and growing attention to robot learning, the policy learning section (5.4) covers a comparatively larger body of recent work. At the same time, we emphasize foundational contributions, particularly in planning and control, to modern robotic manipulation.

5.1 Perception↩︎

In the context of manipulation, the goal of perception is to acquire sufficient information for a robot to interact effectively with objects in its environment. Perception is particularly challenged by epistemic uncertainty, since task-relevant properties and system dynamics are often initially unknown. Important information, such as the number of objects in a scene, their motion constraints, or friction properties, may only become available through interaction with the environment. Consequently, active and interactive perception (Section 5.1.1) constitute important mechanisms for reducing uncertainty. At the same time, perception should maintain beliefs only over variables that are relevant to the task while remaining invariant to irrelevant features (5.1.2); such invariance is an effective mechanism for tolerating episodic variation across scenes. From a failure-management perspective, multimodal perception provides redundancy across sensing modalities and can mitigate local failures within the perception system (5.1.3).

Figure 2: Examples of robustness mechanisms in perception. (a) Interactive perception: the robot interacts with the environment to improve perceptual understanding [33]. (b) Invariance in perceptual representation: the perception of the toy remains invariant to changes in camera pose, object pose or deformation, and lighting conditions [34]. (c) Multimodal sensing: perception combining tactile, visual, and acoustic sensing modalities [35]. Figures are used with permission from the authors, or adapted under CC BY.

5.1.1 Active and Interactive Perception↩︎

Epistemic uncertainty in robot perception can be reduced by exploiting a robot’s ability to act in and interact with its environment [36]. A major challenge is occlusion, which can often be addressed simply by moving a camera to acquire a different viewpoint [37], for example, to support grasp execution [38] (Fig. 2-a). By actively selecting viewpoints that reveal previously hidden information, perceptual uncertainty about the environment is reduced, facilitating subsequent manipulation. Beyond contact-free viewpoint selection, robots can also physically interact with their environment to acquire information. This paradigm, known as interactive perception [39], [40], is likewise observed in animals that employ exploratory behaviors to gather information about their surroundings [41]. Many task-relevant object properties are difficult or impossible to infer from passive observation alone. For example, lifting a milk carton reveals its mass distribution, while removing the lid of a box may reveal its contents. In cluttered environments, robots may need to manipulate obstructing objects to expose a target object [42]. Interactive perception can also reveal motion constraints and kinematic structures of articulated objects [43]. Across these examples, the common principle is the active reduction of epistemic uncertainty through action and interaction.

5.1.2 Invariance in Perceptual Representations↩︎

Perception for manipulation is challenging because objects need to be reliably detected from different viewing angles and lighting conditions (Fig. 2-b). In this sense, invariant perceptual representations help tolerate episodic variations in viewpoint, appearance, and environmental conditions. Before deep learning approaches started to dominate computer vision [44], computer vision pipelines were often based on features such as SIFT [45], SURF [46], or ORB [47], which were designed to be invariant to rotation, scale, and possibly other transforms. In deep learning based approaches, such invariance is often achieved using data-augmentation [44] and pooling mechanisms [48]. More recently, foundation models such as DINO [49] have demonstrated the ability to learn task-agnostic visual representations that are robust to clutter, occlusion, and illumination changes. At an even higher level of abstraction, language-based scene representations exhibit strong invariance to visual variations, contributing to the robustness and generalization capabilities observed in vision-language-action models [50].

5.1.3 Multimodality↩︎

Robotic manipulation can leverage multiple sensing modalities, including visual, tactile, auditory, and proprioceptive information, to perceive and model the environment, thereby enabling more effective interaction with it (Fig. 2-c). These modalities are both complementary and partially redundant in their functions. As a result, failure in one modality need not lead to failure of the overall perception system; other modalities can still provide sufficient information to support task execution. Humans, for example, often rely on touch when vision is unavailable. In robotics, vision provides rich global scene awareness but is susceptible to occlusion, whereas tactile sensing offers detailed local information about properties such as stiffness, texture, friction, and contact geometry. Audition can capture transient interaction events, such as the sound of a snap-fit assembly [35]. These modalities have been successfully combined in prior work to improve perception robustness [51], [52]. Beyond multimodal sensing, redundancy can also arise within a single sensing modality, as exemplified by the large number of whiskers in rodents or the compound eyes of insects. From a failure-management perspective, such redundancy helps mitigate local failures and prevents them from compromising overall perceptual functionality.

5.2 Planning↩︎

Manipulation planning operates on information provided by perception and is therefore inevitably affected by perceptual imperfections. In particular, imperfect perception and state estimation introduce epistemic uncertainty: a mismatch between the true robot–environment parameters \(\theta\) and the internal model \(\hat{\theta}\) (dynamics uncertainty), and consequently between the true state \(x_t\) and its estimate \(\hat{x}_t\) (state uncertainty). Under such mismatches, trajectories computed by the planner may fail to achieve the desired goal \(\mathcal{G}\). Manipulation planning must also contend with aleatoric uncertainty in the dynamics and observation models, as well as episodic variation across task instances.

Existing robust planning mechanisms improve manipulation performance through two complementary strategies. On the one hand, epistemic uncertainty can be reduced by explicitly reasoning about it (Section 5.2.1) or by exploiting task mechanics through motion strategies and funnels (5.2.2). On the other hand, planning can tolerate uncertainty and disturbances by incorporating robustness margins (5.2.3) or by establishing forms of closure around target objects (5.2.4).

5.2.1 Reasoning over Uncertainties↩︎

Even high-quality sensors and advanced perception algorithms cannot fully eliminate epistemic uncertainty. A direct way to mitigate its effect is to represent uncertainty explicitly as a probability distribution over possible states—commonly referred to as a belief—and to update this belief online using new observations (Fig. 3-a). Formally, the agent maintains a belief \(b_t \in \mathcal{B}(\mathcal{X})\) over the state \(x_t\), conditioned on the history \(h_t = (y_{0:t}, u_{0:t-1})\) under the internal model \(\hat{\theta}\), i.e., \(b_t(x) = \mathbb{P}(x_t = x \mid h_t, \hat{\theta})\). Planning then operates in belief space through a policy \(u_t = \pi(b_t)\), aiming to improve expected task performance \(\Gamma(\pi)\) across episodic variations \(\nu\).

Belief representations enable exploration that actively reduces uncertainty by trading off immediate task performance against information gathering. They also support cautious control, allowing robot policies to account for the reliability of available information. For instance, [53] applied belief dynamics to grasping, synthesizing locally optimal feedback policies in belief space with replanning. Similarly, [54] incorporated pose and contact-dynamics uncertainty into a belief-space formulation and validated it on planar pushing without sensory feedback. At longer horizons, belief reasoning extends to uncertainty-aware task and motion planning; [55] integrated belief propagation with symbolic decision making to address incomplete knowledge of object properties and locations in mobile manipulation.

5.2.2 Motion Strategy and Funneling↩︎

Besides reasoning over uncertainties in belief space, another way to reduce uncertainty is to exploit task mechanics. In motion strategies, also termed sensorless manipulation by [56], contact dynamics and gravity provide passive negative feedback that drives objects toward desired configurations, often without exteroceptive sensing. This contrasts with sensor strategies, which rely on rich observations but are susceptible to perceptual errors. Motion strategies are effective in diverse settings, including tray-tilting to orient objects of unknown pose using only gravity and contact [56], push–grasp behaviors in clutter that funnel the target into the gripper while displacing distractors [57], and industrial part feeders whose fences enforce consistent orientation under pose and contact variability [58]. More generally, environmental contact and gravity as task mechanics can reduce uncertainties, or collapse episodic variations into a predictable, smaller set of outcomes (often termed extrinsic dexterity [59]). For example, pressing a loose pair of chopsticks against a tabletop passively aligns their tips through the collision.

A representative abstraction of motion strategy is funneling [60]. A funnel is a region of attraction \(\mathcal{F}\subset\mathcal{X}\) such that, for \(x_0\in\mathcal{F}\) and bounded disturbances \(w_t \in \mathcal{W}, v_t \in \mathcal{V}\), the resulting execution converges to the goal set \(\mathcal{G}\) (i.e., \(x_T \in \mathcal{G}\)), in other words reducing the pose or geometry uncertainty and model mismatch. Such strategies exploit repeated contacts and environmental constraints to steer outcomes without precise sensing, as in vibrating bowl feeders used for industrial part feeding. The funneling concept traces to pre-image backchaining [61], which recursively constructs sets of states from which motions succeed under bounded execution errors. It was later extended by LQR-Tree methods [62], where verified funnels cover the reachable state space to enable robust trajectory planning under model uncertainty. More recently, composing simple in-hand manipulation funnels has shown surprising robustness to variations in object mass and pose [63] (Fig. 3-b).

Figure 3: Examples of robust planning mechanisms. (a) Reasoning over uncertainties: planning in belief space, where the belief of the box position b(q^o) is updated to b(q^o_+) after it is pushed by the robot hand [64]. (b) Funneling with in-hand reorientation primitives, where primitives are chained to reach the target cube orientation [63]; (c) Force closure (c1), form closure (c2), and caging (c3). Photos in (a) and (b) are adapted under CC BY; figure (c) is inspired by [65], [66].

5.2.3 Robustness Margin↩︎

Beyond explicitly estimating or compensating for uncertainty, an alternative strategy is to tolerate uncertainty by planning with a robustness margin. The central idea is to generate motions that remain sufficiently far from constraint boundaries, so that moderate disturbances or model mismatch do not immediately lead to failure. Many manipulation tasks require changing the object’s state through interaction while satisfying state and input constraints, \(x_t\in\mathcal{X}\) and \(u_t\in\mathcal{U}\). These constraints encode collision avoidance, contact consistency, kinematic limits, torque bounds, and related feasibility conditions. Classical methods in this category primarily guarantee feasibility under a nominal model [31], [67], [68].

However, under epistemic mismatch in \(\theta\) and stochastic disturbances, executions may deviate from the nominal constraint manifold and violate constraints. Robust manipulation planning addresses this gap by constructing robustified feasible subsets, thereby maintaining a “margin” from constraint violation to tolerate uncertainty. For example, chance-constrained formulations enforce the probabilistic satisfaction of contact and state constraints under stochastic dynamics [69], [70], while set-based approaches employ tightened constraint sets or disturbance-invariant tubes to guarantee feasibility under bounded uncertainty [71]. These approaches ensure that constraint satisfaction is preserved within a specified disturbance or risk budget, thereby tolerating uncertainty at the planning level. We will discuss several more works along this line in Section 6.2.2 and how they can be used as robustness evaluation protocols. The rich body of work in robust control theory provides a strong foundation for future developments of this mechanism.

5.2.4 Closure↩︎

In the spirit of uncertainty tolerance, robustness can also be achieved through closure, which imposes geometric or force constraints on object motion and thereby makes the motion resistant to, for example, external disturbances to manipulated objects. In many grasping and fixturing tasks, the goal is to maintain the object within an in-hand stable set. Therefore, closure can be viewed as shaping the robot-object interactions such that the object’s reachable configurations remain within a bounded subset \({\mathcal{G} \subset \mathcal{X}}\) despite bounded disturbances \({w_t\in\mathcal{W}}\) in the dynamics \({x_{t+1}=f(x_t,u_t,w_t ; \theta)}\). Force closure captures equilibrium grasps that can resist arbitrary external wrenches via internal contact forces [72] (Fig. 3-c1). Form closure is a geometric condition in which contact constraints eliminate all object motions even without friction [73] (Fig. 3-c2). Caging (also referred to as “object closure” in some contexts) further relaxes contact requirements by trapping the object within a bounded region of its configuration space without precise force regulation [74] (Fig. 3-c3). Across these closure methods, the shared principle is to restrict the configuration space so that the object cannot move freely or escape, thereby making the manipulation outcome tolerant to perceptual errors, control imprecision, and external disturbances.

5.3 Control↩︎

Whereas planning operates at a higher level, often through open-loop decision making, control closes the loop in real time by responding to sensory feedback and correcting deviations that no plan can fully anticipate. A large class of robotic control methods addresses motion control in free space, such as trajectory tracking or point-to-point reaching, where the manipulator operates without contact. Manipulation, however, is fundamentally characterized by the making and breaking of unilateral frictional contacts, which introduce discontinuities and uncertainties. As a result, manipulation control must regulate not only motion but also forces and contact events. Controllers in contact-rich settings must therefore contend with external disturbances and actuation errors, as well as epistemic uncertainty arising from imperfect contact models. Poorly designed controllers can amplify these challenges, leading to instability or excessive contact forces, particularly in interactions with rigid environments. From the perspective of robustness, compliance (5.3.1) and invariance (5.3.2) primarily serve to tolerate uncertainty and variation, whereas predictive control (5.3.3) contributes to failure prevention by anticipating future contact events and system evolution.

5.3.1 Compliance↩︎

Compliance balances the control of position and force; higher compliance allows larger position deviations under external forces. It can be realized passively through mechanical design (Section 5.5.1) or actively within the control loop [75]. Active compliance control regulates how a robot responds to external forces during contact, allowing it to comply rather than rigidly enforcing a motion. This property makes compliance a toleration strategy that bridges the gap between the robot’s internal model and the complex, often unpredictable, physics of the real world. Impedance and admittance control represent the two classical approaches. Impedance control enforces compliance from motion to force by commanding motion trajectories that indirectly regulate contact forces [76], \[M_d(\ddot x - \ddot x_d) + D_d (\dot{x} - \dot{x}_d) + K_d (x-x_d) = \tau_{\text{ext}}, \label{eq:impedance}\tag{4}\] where \(x \in \mathbb{R}^m\) denotes the task-space position (or pose) of the end-effector, \(x_d\) is the desired trajectory, and \(\tau_{\text{ext}}\) is the external wrench arising from contact. The matrices \(M_d\), \(D_d\), and \(K_d\) are the desired inertia, damping, and stiffness parameters that define a virtual mass–spring–damper system. This equation specifies the closed-loop interaction dynamics: external forces induce bounded motion deviations governed by \((M_d,D_d,K_d)\). On the other hand, admittance control does so from force to motion by computing the motion response to a commanded force [77].

Classical active compliance methods typically rely on accurate models or restrictive assumptions. More recently, model-free approaches have emerged, learning compliant manipulation behavior from human demonstrations or through reinforcement learning [78][80]. For example, an adaptive compliance policy can adjust compliance both spatially and temporally from demonstrations for contact-rich tasks [79]. Future directions may explore combining active compliance with passive mechanical compliance to enhance adaptability under contact uncertainty.

5.3.2 Invariance↩︎

Invariance ensures that a system executes the behavior prescribed by a nominal controller while remaining within a certified safe set, such as bounds on force tracking error [81] or admissible state regions. Like other active compliance control, invariance-based methods do not reduce epistemic uncertainty in states or parameters. Rather, they achieve robustness by constraining system evolution so that bounded disturbances or modeling errors cannot drive the state outside a prescribed safe set; in this way, disturbances are tolerated rather than eliminated. Formally, consider the closed-loop dynamics \(x_{t+1} = f(x_t, u_t, w_t; \theta)\) under an output-feedback policy \(u_t=\pi_t(b_t)\), where \(b_t(x) = \mathbb{P}(x_t = x \mid h_t, \hat{\theta})\). A set \(\mathcal{C}\subset\mathcal{X}\) is robustly forward invariant if \[x_t \in \mathcal{C} \;\Rightarrow\; x_{t+1}\in \mathcal{C}, \quad \forall w_t\in\mathcal{W}.\] Invariance-based control enforces this condition by modifying or projecting control inputs when the boundary of \(\mathcal{C}\) is at risk of being violated, thereby guaranteeing constraint satisfaction despite aleatoric uncertainty and bounded modeling errors.

Practical approaches include reachability analysis, which conservatively overapproximates all possible evolutions to guarantee safety [82], and control barrier functions (CBFs), which enforce invariance through real-time inequality constraints. When combined with control Lyapunov functions, CBF-based methods unify safety and goal achievement in an optimization framework, as demonstrated in tasks such as ball balancing [83].

5.3.3 Predictivity↩︎

Neuroscience suggests that human manipulation relies on the integration of long-term prediction with short-horizon reactivity [2]. Reactivity provides rapid sensory feedback, particularly tactile and visual, to detect mismatches between expected and actual outcomes and correct errors. Predictivity, on the other hand, enables the anticipation of collisions or contact events. In this way, predictive control primarily tolerates uncertainty and variation by proactively compensating for their future effects, thereby helping prevent failures before they occur. Purely local reactive control may fail in contact-rich settings, where abrupt changes in dynamics due to contacts can drive the system into unsafe states. To address this, model predictive control (MPC) incorporates a lookahead structure with reactive schemes. It optimizes actions over a finite horizon and executes the first control input. For example, MPC-based hybrid force–motion control is employed to simultaneously regulate end-effector trajectories and contact forces during interaction, explicitly accounting for contact mode transitions within the prediction horizon [84]; a contact-implicit MPC is utilized that predicts motion and contact forces within a trust region, enabling stable local control under unilateral contact and friction constraints [85]; tube-based MPC guarantees that the real trajectory remains within a bounded “tube” around the nominal one, ensuring performance under external disturbances and model mismatch [86].

Table 2: Overview of learning-based robustness mechanisms for robotic manipulation (Complementary to Table [tbl:tab:robustness95mechanisms]).
Mechanism Method Principle Section
Training Distribution Domain Randomization [sec:learning-distribution-centric]
Data Augmentation Variation Toleration
Data Scaling and Task Diversity
Policy Architecture Perceptual and Structural Inductive Biases Uncertainty & Variation Toleration [sec:architecture-distribution-centric]
Generative Policy Parameterizations Uncertainty & Variation Toleration
Temporal Abstraction and Hierarchy Failure Prevention
World-model-based Policies Uncertainty Reduction
Learning Objective Adversarial Training Objectives Uncertainty & Variation Toleration [sec:objective-distribution-centric]
Safety- and Constraint-based Objectives Failure Prevention
Regularized and Conservative Objectives Failure Prevention
Policy Adaptation Interactive Human-in-the-loop Imitation Temporary Failure Recovery [sec:adaptation-distribution-centric]
Autonomous Continual Improvement

5.4 Policy Learning↩︎

In parallel with the classical perception–planning–control pipeline, end-to-end policy learning has emerged as an alternative paradigm for manipulation. Learned policies can adapt to unstructured variations in the open world that are difficult to address analytically. In this sense, policy learning mainly improves robustness by tolerating episodic variation and disturbances. Depending on the policy design, it can both prevent failure through robust action selection and recover after temporary task failure.

Within the formalism of Section 3, these methods instantiate a parameterized policy \(\pi_\phi\) and train it from data so that the robust performance functionals \(\Gamma(\pi)\) remain high across an episodic distribution \(\rho_\mathcal{N}(\nu)\). In the rest of this section, we group policy-learning approaches according to four complementary ways (Table 2) in which they achieve robustness: some methods shape the tasks, environments, and perturbations that \(\pi_\phi\) is trained on (5.4.1); some methods design policy classes and architecture whose inductive biases and temporal structure reduce sensitivity to nuisance variation and partial observability (5.4.2); some methods modify the learning objective to better align training with robust performance (5.4.3); and others allow policies to adapt at deployment through residual corrections, online updates, or test-time adaptation (5.4.4).

5.4.1 Training Distribution↩︎

In the notation of Section 3, robustness is measured by the functional \(\Gamma(\pi)\) evaluated over an episodic distribution \(\rho_\mathcal{N}(\nu)\). A branch of approaches acts directly on this distribution: it shapes how training episodes are generated while keeping the policy class \(\Pi\) and the learning objective fixed. These methods aim to tolerate variations across episodes by broadening the training distribution over dynamics, initial states, environments, and goals.

A first family of methods uses domain randomization, widely used, particularly in reinforcement learning. Instead of a single nominal environment, one defines a family of “worlds” \(\theta\) (and sometimes \(x_0\) and \(\mathcal{G}\)) from a broad distribution. Policies trained under such variation can transfer to real manipulation with minimal tuning, as demonstrated for vision-based grasping appearance [87] and for dexterous in-hand manipulation [88], [89]. By expanding the training distribution \(\rho_\mathcal{N}\) to encompass diverse scenarios, these methods enable the policy to tolerate variations in dynamics and appearance.

A second line achieves robustness via data augmentation on training trajectories. These methods keep the generator of \(\nu\) fixed, but expand the empirical distribution \(\hat{\rho}_\mathcal{N}\) by transforming demonstrations according to task priors: prior work applies SE(2)/SE(3) transformations to scenes and actions to exploit invariance and equivariance [90][92], re-renders fixed trajectories from novel viewpoints to vary appearance [93], [94], and retargets demonstrated trajectories to new object and scene configurations [95][97]. Beyond such structured priors, a related class injects calibrated noise into states and actions to mitigate covariate shift and compounding error, making the policy more robust [98][100].

A third class relies on data scaling and task diversity, as most visibly demonstrated in recent vision–language–action (VLA) models. Rather than explicitly parameterizing randomization or augmentations, these methods assemble very large, heterogeneous datasets of manipulation episodes spanning many robots, scenes, tasks, and language-specified goals, and train a single policy across this mixture [50], [101][103]. Robustness here is treated as generalization over a very broad empirical \(\hat{\rho}_\mathcal{N}\) that pools demonstrations and deployments across embodiments, tasks, and environments, with a single \(\pi_\phi\) expected to maintain high \(\Gamma(\pi_\phi)\) throughout this mixture. An important direction in data scaling extends beyond successful task executions to include recovery behaviors following temporary failures, capturing the adaptive strategies required in unstructured environments. Policies trained only on ideal trajectories may overfit to nominal conditions. For example, Large Behavior Models incorporate hybrid sim-to-real datasets containing both successful executions and recovery episodes [104].

5.4.2 Policy Architecture↩︎

A policy is a mapping \(\pi : \mathcal{B} \to \mathcal{U}\) from beliefs to actions. A line of approaches pursues robustness by designing policy architecture and data representation: they constrain how information about observations is encoded and how actions are parameterized, without changing the training distribution \(\rho_\mathcal{N}\) or the learning objective. The goal is to make the mapping \(b_t \mapsto u_t\) inherently tolerant to nuisance variation, partial observability, and long horizons.

First, perceptual and structural inductive biases build a geometric state from raw observations. Keypoint-based representations encode objects and goals as a small set of 2D/3D landmarks and express tasks as geometric relations among these keypoints, so the policy operates in a low-dimensional space aligned with contacts and affordances rather than raw pixels [105][107]. SE(3)-equivariant representations constrain internal features (and sometimes actions) to transform consistently under rigid motions of the workspace, so that rotating or translating the scene induces the same transformation in the encoded state and allows one controller to reuse a strategy across pose-shifted instances of a task [108][111]. Graph-structured encoders treat objects, robot links, and goals as nodes with edges capturing spatial or relational constraints, and use message passing to learn local interaction rules that generalize to scenes with more objects, new goal configurations, and multi-object rearrangement [112][114]. Across these designs, robustness comes from aligning the learned representation with the geometry and relational structure of the task, so that pose- and relation-preserving changes in the scene tend to yield similar decisions, resulting in stable performance under these variations.

Second, generative policy parameterizations implement \(\pi_\phi\) as a conditional generative process over short action sequences, most commonly using diffusion or flow-matching models for manipulation tasks [115], [116]. The policy maps beliefs \(b_t\) to actions by starting from injected noise and iteratively refining a sample via a supervised denoising or flow-integration process. Empirically, this iterative computation with noise injection induces an inductive bias that improves closed-loop robustness to covariate shift [117].

Temporal abstraction and hierarchy represent actions as skills or options that span multiple timesteps. The policy acts in an extended action space of options: a high-level policy selects a skill, and a low-level controller executes it until termination, so the decision-making problem becomes a semi-MDP with a shorter effective horizon [118]. Hierarchical RL methods learn both the skills and the high-level policy, often using subgoal states as the interface between levels and off-policy updates to remain sample efficient in continuous control [119]. In manipulation, skill-based approaches learn and plan with such temporally extended behaviors, such as predicting long-horizon discrete action sequences from a scene image for sequential physical reasoning [120], discovering reusable closed-loop skills from unsegmented demonstrations and training a meta-controller to compose them [121], or chaining separate diffusion models for individual skills to solve unseen long-horizon tasks under constraints [122]. Robustness in this class arises from horizon reduction: the high-level policy makes skill-level decisions per episode, which reduces compounding error, supports reuse of skills across task variations, and thereby helps prevent failures before they occur.

Lastly, world-model-based policies couple the controller with a learned dynamics model in a latent space. They maintain a latent state \(z_t = e(o_{0:t}; \theta)\) together with a predictive model \(\hat{p}_\psi(z_{t+1} \mid z_t, u_t)\), and use short imagined rollouts in this latent space to reason about future states before committing to an action [123][126]. World-model-based policies can reduce effective epistemic uncertainty by learning predictive structure over how the environment evolves under the robot’s actions. Rather than eliminating uncertainty, the learned model provides an internal approximation of otherwise unknown or partially observed dynamics that can be queried during decision making. In manipulation, such latent models are used for planning and runtime monitoring to score candidate actions under many predicted futures and reject those that would enter unsafe, constraint-violating, or high-risk regions [127][129].

5.4.3 Learning Objective↩︎

While distribution- and architecture-centric approaches shape what the policy sees and how it represents it, in this part we introduce methods that instead shape what the policy optimizes for, embedding worst-case performance, safety, or conservatism directly into the learning objective.

Adversarial training objectives make robustness explicit by optimizing the min–max criteria of Eq. 3 rather than a nominal expected return. The policy is trained against an adversary that chooses disturbances, environment parameters, or competing behaviors to minimize task performance while the protagonist maximizes it. In practice, the adversary injects destabilizing forces or physical perturbations to grasps and object poses, so policies are trained on deliberate worst cases rather than randomly sampled disturbances [130][132]. By repeatedly optimizing against these hard cases, the learned policy becomes less sensitive to uncertainty and variation, thereby improving its ability to tolerate them at deployment.

Safety- and constraint-based objectives restrict which policies are admissible when we evaluate robustness. In the POMDP view of Section 3, one augments the control problem with constraint costs and searches for policies that achieve high task performance \(\Gamma(\pi)\) while keeping these costs below a threshold, as in constrained MDP methods such as Constrained Policy Optimization [133]. Related work uses safety layers or control barrier functions that project or override the learned action when it would violate certified constraints [134], [135], thereby preventing failures by keeping execution within admissible regions.

Regularized and conservative objectives aim to improve robustness to epistemic uncertainty and limited data coverage by constraining how far the learned policy and value function extrapolate beyond the empirical episodic distribution \(\hat{\rho}_\mathcal{N}(\nu)\). Offline RL work frames this as “extrapolation error” and shows that robust performance typically requires (i) behavior regularization, which keeps \(\pi\) close to actions well supported by the data, and (ii) pessimistic value estimates that down-weight out-of-distribution actions [136][139]. In real-robot manipulation, these ideas appear in conservative offline RL that improves over imitation on tabletop tasks using logs from safe operation [140], in large-scale deployments that rely on heavily regularized off-policy updates for stable long-term waste-sorting with mobile manipulators [141], and in regularized imitation methods that combine optimal-transport trajectory matching with a pull toward demonstration behavior to achieve few-shot real visual manipulation [142]. Across these examples, robustness comes from biasing learning toward regions of state–action space that are well covered and conservatively valued, thereby reducing out-of-distribution actions and helping prevent failures during execution.

5.4.4 Policy Adaptation↩︎

This line of work assumes a base policy \(\pi_{\text{base}}\) has been trained under some episodic distribution \(\rho_\mathcal{N}^{\text{train}}(\nu)\), and focuses on how its behavior is modified using feedback from the deployed environment when the test-time distribution \(\rho^{\text{test}}(\nu)\) differs from \(\rho_\mathcal{N}^{\text{train}}(\nu)\). Unlike distribution-centric approaches, they do not redesign \(\rho_\mathcal{N}\) up front, but instead close the loop between deployment performance and further adaptation.

First, interactive human-in-the-loop imitation refines a pretrained policy at deployment using online expert feedback. Humans can intuitively identify and correct failure-imminent situations or recover from temporary failure, providing targeted supervision where pre-trained policies are weakest [143]. In DAgger-style methods, imitation learning is reduced to online no-regret learning: at each iteration, the current policy is rolled out to generate states, the expert labels those visited states with preferred actions, these pairs are added to an aggregate dataset, and a new policy is trained on the union [144]. This ensures that the learned stationary policy performs well under the state distribution it induces, rather than only on states seen in expert demonstrations, which directly addresses covariate shift in sequential decision making. Subsequent work adapts this idea to human experts in safety-critical or latency-limited settings [145], to remote teleoperation for contact-rich manipulation [146], to “on-the-job” deployment where autonomy and learning run continuously with humans stepping in on hard cases [143], and to kinesthetic delta corrections with force-aware residual policy via compliance [78].

Second, autonomous continual improvement adapts policies during deployment without relying on explicit expert labels. One branch uses online RL on the real robot: a warm-start policy (often from demonstrations or offline data) continues to collect rollouts on the hardware, and policy updates \(\pi_\phi\) under the actual test-time episodic distribution, with mechanisms such as safety constraints, automatic or cheap resets, and simple reward specification to make this practical for contact-rich manipulation [147][152]. A second branch uses in-context and prompt-based adaptation: sequence models are trained on multi-task data so that, at test time, a short prompt of teleoperated demonstrations or recent trajectories for the current variation \(\nu\) conditions the model, enabling few-shot adaptation to new objects, layouts, and goals without gradient updates [153][155].

5.5 Hardware↩︎

Manipulation robustness can also be achieved through the robot’s embodiment, complementing the contributions of perception, planning, control, and learning. This perspective is grounded in the concept of morphological computation, where physical structures perform functions that would otherwise require explicit control [156]. Through the design of materials, geometry, and actuation, the body can shape the system dynamics \(f\) such that a wider range of states \(x_t\) and actions \(u_t\) lead to successful outcomes, while reducing the consequences of errors in perception, modeling, and execution.

A recurring theme of hardware-based robustness is the toleration of object variation and external disturbances through physical design. This can be achieved through passive compliance (Section 5.5.1), which absorbs disturbances and accommodates contact uncertainty, or through adhesive contact mechanisms (5.5.2), which remain effective despite geometric variation and perturbations. In addition, some robotic systems achieve robustness through morphological adaptation (5.5.3), reconfiguring their physical structure in response to changing conditions and thereby restoring performance when an existing morphology becomes ineffective.

5.5.1 Passive Compliance↩︎

Passive compliance enhances manipulation robustness by physically tolerating disturbances and contact uncertainty through compliant, soft surface material, before they propagate through the control loop. It can be viewed as a mechanical counterpart to active compliance (Section 5.3.1) that shapes the interaction dynamics prior to control. By allowing bounded motion in response to external forces, compliant structures reduce the sensitivity of contact interactions to disturbances, geometric variation, and modeling errors. As a result, reliable and adaptive manipulation can often be achieved with reduced reliance on precise sensing, accurate models, and high-bandwidth control.

This principle is widely observed in both biological and robotic systems. Dolphins, for example, use marine sponges as protective tools while foraging on rocky seabeds (Fig. 4-a), thereby tolerating environmental uncertainty and avoiding damaging contacts [157]. In robotic manipulation, passive compliance is commonly used to tolerate misalignment and reduce impact forces in assembly tasks [158]. Similar principles underlie soft grippers based on compliant materials or tendon-driven structures, which passively conform to object geometry under uncertainty [159]. Hybrid designs further combine compliance with controllable stiffening, as exemplified by granular-jamming grippers, enabling compliant exploration followed by adaptive stiffening to securely grasp the object [160], [161].

Figure 4: Examples of hardware mechanisms for robust manipulation. Hardware intelligence in robotics is often inspired by biological strategies. (a) Passive compliance: Dolphins wrap marine sponges around their beaks for protection [157], while robotic hands benefit from compliant surface material. (b) Adhesion: gecko feet inspire adhesive grippers, such as [162]. Figure (a) is adapted under CC BY; image courtesy of (b): Prof. Kellar Autumn, Lewis and Clark College [163].

5.5.2 Adhesion↩︎

Adhesive material modifies robot-object contact surfaces to sustain strong tangential (shear) forces with minimal normal force, which is particularly suitable for delicate, thin, or flat objects. By enabling surface-based attachment rather than relying solely on friction-limited point contacts, adhesion alters the contact dynamics and makes manipulation more tolerant to uncertainty in contact location, object pose, and contact modeling. Common adhesion mechanisms include gecko-inspired dry adhesion and vacuum-based suction [160]. Gecko-inspired adhesives exploit van der Waals forces to generate strong shear interactions across a wide range of surface textures without requiring active control (Fig. 4-b). Such systems have been used to grasp bulky objects in microgravity [164], demonstrating robustness to object misalignment and uncertain contact conditions. Vacuum-based suction is another widely adopted adhesion mechanism. Besides adhesion, many suction cups incorporate bellows that provide passive compliance, allowing the system to accommodate pose errors and conform to local surface curvature. Multi-affordance manipulators that combine suction cups with parallel-jaw grippers [30] further exploit mechanical redundancy, improving robustness and extending applicability across a broader range of objects and environments.

5.5.3 Morphology Adaptation↩︎

The environment and task requirements encountered by a robot are often not fixed. As a result, a morphology that performs well under one set of conditions may become ineffective when the environment, object properties, or task requirements change. Some robotic systems address this challenge by adapting their embodiment, modifying geometry, stiffness distribution, or kinematic structure to suit new conditions. In this way, morphological adaptation can be viewed as a mechanism for recovering from temporary failure caused by changing environments or task requirements [165]. Complementary to passive compliance and adhesion, which shape interactions locally at the contact interface, morphological adaptation reshapes the embodiment at a global level.

In soft robotic hands, for example, the bending location and range can be adjusted by inserting a stiff rod into the center of each finger [166]. Such reconfiguration enables the end-effector to switch between grasping modes and maintain performance across different object types without requiring complex sensing or control. PolyBot [28] further illustrates modular reconfiguration, where detachable modules autonomously rearrange to accommodate new workspace constraints. By adapting their morphology to changing conditions, such systems improve robustness, support fault tolerance, and extend manipulation capabilities in unstructured environments.

6 Evaluation Methods of Manipulation Robustness↩︎

Evaluating manipulation robustness requires quantitative criteria that reflect a system’s capability to achieve its task objectives under challenges. In this section, we first present empirical evaluation protocols that assess how effectively a robot performs physical manipulation tasks in simulation or in the real world. We then revisit analytical evaluation protocols that assess robustness using grasp-quality measures, margin-based metrics, and other abstract mathematical constructs rather than empirical test-time task performance.

6.1 Empirical Evaluation Protocols↩︎

Based on task performance, these evaluation protocols empirically assess whether and how effectively a manipulator achieves the task goal \(\mathcal{G}\) despite varying challenges. They provide more direct comparisons for end users [167] and are widely adopted by both empirical (learning-based) and analytical methods.

6.1.1 Goal-Oriented Measures↩︎

Given a manipulation task specification together with domains of epistemic uncertainty, aleatoric uncertainty, and episodic variation, the most widely used robustness evaluation protocol is to measure empirical task success or failure [1]. This approach, commonly referred to as Monte Carlo robustness tests, executes the system repeatedly under a sampled model mismatch \((\hat{\theta},\theta)\), stochastic disturbances \(w_t\) and \(v_t\), and episodic variations \(\nu\), and estimates the probability of achieving the task goal. Concretely, robustness is the expected probability of goal satisfaction under the challenges, \[\mathbb{E}_{\nu \sim \rho_\mathcal{N}, w_{0:T} \sim \rho_\mathcal{W}, v_{0:T} \sim \rho_\mathcal{V}} \left[\mathbf{1}\{x_T \in \mathcal{G}\} \mid (\hat{\theta}, \theta) \right], \label{eq:evaluation-target}\tag{5}\] often approximated by empirical success rates \({n_{\text{success}}}/{n_{\text{total}}}\) computed over multiple trials. This evaluation directly aligns with the robustness formulations introduced in Section 3.

Following this evaluation protocol, examples include success rates in object pose estimation under perception noise and occlusion [91]; success rates or expected returns in reinforcement learning, which often reduce to goal satisfaction indicators, evaluated across episodic variations such as different initial object placements or object instances [89], [152]; and success under adversarial perturbations, including external disturbances applied to manipulated objects or injected joint-level perturbations during execution [132]. Variants of this approach include systematic parameter sweeps, where parameters in \(\theta\) (e.g., inference latency, friction coefficients, lighting conditions) are varied across controlled ranges and performance is plotted as a function of the parameter [115].

6.1.2 Stage-wise Measures↩︎

While the formulation in Eq. 5 focuses on terminal goal satisfaction, there are many manipulation tasks more naturally characterized by trajectory-level subgoals. In such cases, success depends not only on reaching a final state but also on correctly executing a sequence of intermediate stages. Consequently, several works adopt stage-wise evaluation metrics, in which a task is decomposed into semantically meaningful phases and success is assessed at each stage. For example, [168] evaluate fragile-object manipulation policies across three stages: approach, stable grasp, and task completion. [169] decompose a paper-cup lifting task into clamping and lifting stages. Evaluators may also design task-specific semantic rubrics to assess policy behavior beyond aggregate success rates [170]. For instance, in a Flip-and-Serve Pancake task, relevant criteria may include: “Robot collided with anything?”, “Robot flipped pancake?”, “Robot picked up pancake?”, in addition to overall task success. The stage-wise protocols provide a more fine-grained assessment of policy performance, enabling the identification of failure modes that may be obscured by a single binary success measure.

6.2 Analytical Evaluation Protocols↩︎

While success probability under challenges provides an empirical, task-level notion of robustness, it is often insufficiently fine-grained for answering more specific questions about how a manipulation system withstands challenges. In many scenarios, one is not only interested in whether a task succeeds, but also in the structure and degree of tolerance to disturbances, uncertainty, or variation. This is where analytical evaluation protocols become valuable: they aim for potential capability, rather than only realized capability, through explicit mathematical constructs. They are structured and interpretable, tailored to specific manipulation mechanisms and failure modes. They are often approximated by the success probability in Eq. 5 . Unlike performance-based protocols, which primarily serve as empirical diagnostic tools for robustness assessment, analytical protocols can also function as optimization objectives embedded within planning, control, or learning methods for the synthesis of robust manipulation behavior. As follows, we introduce several analytical evaluation protocols.

6.2.1 Closure-based Grasp Quality Measures↩︎

Closure properties characterize how contact forces and geometric constraints prevent an object from deviating from a desired configuration or set of configurations. These properties are typically analyzed from two different perspectives: local conditions at the contact and global conditions within the object’s configuration space.

Local measures derived from force closure or form closure quantify the instantaneous robustness of quasistatic manipulation via wrench-space or geometric analysis. These measures can be evaluated either analytically or empirically. Analytical approaches, commonly referred to as grasp quality measures [171], compute robustness using physics-based models to determine the system’s ability to resist external disturbances. In contrast, empirical approaches approximate these analytic measures by learning from data. For instance, grasp datasets labeled with analytic quality values are used to train neural networks that predict robustness directly from visual or sensory inputs [172], enabling scalable estimation in unstructured environments.

While local measures assess stability via instantaneous force or geometric conditions, many complex manipulation tasks depend on the global characteristics of the configuration and contact spaces—specifically, how reachable, connected, or “trapped” an object remains within its environment. Topological analysis captures such global invariants, providing a perspective complementary to purely local evaluations [173], [174]. For example, energy-bounded caging [175] defines robustness in terms of the size or persistence of the configuration-space region that keeps the object contained. In this context, a larger connected component implies a greater structural tolerance to uncertainty or disturbances.

6.2.2 Margin-based Measures↩︎

“Margin” to failure quantifies robustness by measuring the distance between a system’s operating state and the boundary of task failure. Intuitively, systems maintaining larger margins are more robust, as they can tolerate greater disturbances, modeling errors, or stochastic noise before a failure occurs. Consider the task of placing a cup of coffee on a coaster: precise positioning is secondary as long as the liquid level remains safely below the rim. If the goal region \(\mathcal{G}\) is defined as the set of states where all liquid remains inside the cup, its complement \(\mathcal{G}^{c}\) corresponds to failure (spillage). While a binary reward only detects the occurrence of failure, a safety margin quantifies the resilience of the current state: \[\sigma_{\text{margin}} = \operatorname{dist}(x, \partial \mathcal{G}),\] where \(\operatorname{dist}(\cdot)\) denotes a task-appropriate distance metric and \(\partial \mathcal{G}\) is the boundary of the goal region. Positive values indicate the system is within \(\mathcal{G}\), with larger values implying greater robustness.

In neurophysiology, this distance is often defined through energy, i.e., the energy required for the system to transition from a safe state \(x\) to failure \(\mathcal{G}^c\) [176]. Humans intuitively maintain such margins by relying on simplified internal models rather than exact dynamics, selecting control strategies that preserve safety margins [177]. Similar concepts have been applied to evaluate robotic manipulation robustness [178]. Analogous ideas appear in force- and wrench-based robustness metrics, such as grasp stability measures that quantify the maximum external wrench a grasp can resist before slippage [179]. Extending the concept over time leads to the notion of a “safety tube”, a region surrounding a nominal trajectory within which the system can remain despite bounded disturbances [86], [180].

6.2.3 Signal Temporal Logic Measures↩︎

Signal Temporal Logic (STL) is a formal language for describing time-dependent task requirements over real-valued signals, such as contact force, object position, or velocity — examples include conditions such as maintaining a grasp force below a threshold, reaching a target region within a given time, or keeping an object stable throughout execution. STL further provides quantitative semantics that assign a robustness score measuring how strongly a trajectory satisfies or violates the specified requirements [181], [182]. In robotics manipulation, STL robustness has been predominantly used as an optimization target — either as a reward signal in reinforcement learning [183], [184], a loss function for neural predictive control [185], or a planning objective in task-and-motion planning [186]. Its use as a post-hoc evaluation metric remains rare: [170] have explicitly leveraged STL robustness for grading learned manipulation policies after training, demonstrating how the robustness score reveals not only whether a policy succeeds but how close it is to failure across complex, temporally structured objectives. The STL measures are conceptually related to the margin-to-failure measures discussed above, but generalize them to composite temporal specifications involving sequencing, timing, and conditional constraints. A practical limitation is that STL robustness is scale-dependent: predicates expressed in different physical units (e.g., millimeters versus Newtons) produce robustness values on incomparable scales, so that the min/max aggregation inherent in the semantics can be dominated by the choice of units rather than by task-relevant difficulty [187], [188]. Despite these challenges, the broader use of STL measures as an evaluation tool for manipulation robustness remains a promising direction.

6.2.4 Convergence and Divergence Measures↩︎

In dynamically complex manipulation tasks such as tossing, catching, or balancing, robustness often arises from the intrinsic structure of the system dynamics. Humans routinely exploit such dynamics by inducing self-correcting behaviors in which small perturbations are naturally compensated, allowing motion to remain stable under bounded aleatoric disturbances \(w \in \mathcal{W}\). From an evaluation perspective, this form of robustness can be captured by convergence or divergence measures [189], which quantify how trajectories evolve under perturbations. For example, the largest eigenvalue of the symmetric part of the system Jacobian measures the local rate of trajectory divergence or contraction, with negative values indicating stabilizing, disturbance-rejecting dynamics. Such measures can be incorporated into constrained optimization or planning formulations, e.g., as constraints on contact dynamics, to favor actions that induce convergent behavior and naturally funnel executions back toward desired outcomes.

7 Discussions↩︎

In this section, we discuss several complementary insights and implications that emerge from the robustness mechanisms.

7.1 Robustness and Performance Trade-offs↩︎

Manipulation performance is inherently context-dependent. A system may exhibit high performance in controlled laboratory settings, yet experience sharp degradation in unstructured real-world environments. Systems optimized for narrow conditions can achieve exceptional speed or precision, but often do so at the expense of robustness when those conditions change. In nature, star-nosed moles exemplify this trade-off: their specialized nasal appendages enable rapid prey detection in wetland tunnels, but such specialization may be fragile under environmental shifts [190]. By contrast, some systems with moderate peak performance may trade efficiency or accuracy for robustness. Compliant robotic grippers, for example, adapt well to variations in object shape or mass, but typically sacrifice positional precision [159]. These observations suggest that progress in manipulation should be evaluated not only by peak performance under ideal conditions, but by the ability to sustain reliable function under imperfect and unpredictable in-the-wild challenges.

Robustness often intersects with but differs from other fundamental concepts, including safety, stability, generalizability, etc. Safety concerns preventing harm to humans, the environment, or the robot itself, and is critical in domains such as autonomous driving, navigation, and human–robot interaction [191]. Stability, common in control and grasping, refers to maintaining or converging to a desired state under perturbations. Robustness extends beyond these notions, encompassing not just safe or stable operation but the broader pursuit of general task goals under uncertainty or variation.

Robustness and generalizability are also frequently discussed interchangeably, especially in robot learning, but they capture different aspects of system behavior. Generalizability concerns the scope over which a learned policy or model can be applied beyond its training conditions, such as new objects, scenes, embodiments, task instructions, or simulation-to-real transfer [19], [192], [193]. Robustness, in contrast, concerns the reliability of task achievement under a specified set or distribution of uncertainty and variation. Thus, generalization may serve as one mechanism for robustness when the deployment variations are covered by the generalized competence of the policy, but it is neither necessary nor sufficient: a controller may be robust within a narrow operational condition without generalizing broadly, while a generalist policy may cover many tasks yet remain sensitive to small perturbations, contact uncertainty, or distributional shifts not represented in its evaluation protocol.

7.3 Insights from beyond Manipulation Robustness↩︎

Robustness research in domains beyond manipulation offers valuable insights. As mentioned in Section 2, locomotion robustness is particularly relevant, as locomotion can be viewed as a form of “self-manipulation” [1], [194]. Techniques such as adversarial training for controller robustness [195], widely explored in locomotion, remain comparatively underutilized in manipulation and represent a promising direction. Similarly, biological inspirations such as gecko-inspired adhesives, originally studied in climbing and locomotion, have motivated gripper designs [164]. Despite these overlaps, manipulation presents unique challenges: unlike locomotion, which often reduces interaction to discrete foot–ground contacts, manipulation involves rich, multi-body interactions among the hand, object, and environment, with diverse contact types and combinatorial complexity that significantly complicate robustness analysis and design.

7.4 Beneficial Role of Perturbations↩︎

Finally, perturbations, typically viewed as detrimental, can in some cases enhance robustness. In granular media, small vibrations can break clogging arches when pouring sand into a funnel [196]. The vibrating bowl feeder discussed in Section 5.2.2 provides another example, where controlled vibrations guide parts toward consistent configurations. Analogously in manipulation, slight, intentional perturbations can dislodge unstable contacts, reduce excessive contact forces, or guide objects out of shallow local minima. Everyday actions such as wiggling a Jenga block or flicking an omelette exemplify how controlled perturbations stabilize interactions rather than destabilize them. These observations highlight that robustness is not always achieved by tackling uncertainties or variations, but sometimes by strategically exploiting them.

8 Challenges and Open Problems↩︎

The frameworks and principles presented thus far provide a basis for understanding how robustness arises, yet they also make clear that essential scientific and engineering questions remain unanswered. This section identifies some of the open problems and discusses the challenges that must be overcome to achieve manipulation robustness under real-world challenges.

8.1 Benchmarking and Evaluation of Manipulation Robustness↩︎

Effective comparison of manipulation robustness across different methods requires well-designed benchmarks and consistent evaluation methods. Simulation-based benchmarking [197][199] allows experiments under identical conditions and facilitates large-scale testing, yet faces the persistent challenge of the sim-to-real gap that limits physical realism and transferability, as well as being confined to a very narrow distribution of tasks and environments. Real-world benchmarking efforts, including standardized object sets [200], manipulation competitions [201], [202], cloud-based robotic platforms [203], and community-run distributed infrastructures [204], [205], aim to address this gap, yet each makes trade-offs among task diversity, scalability, accessibility, authenticity, or consistency. Future progress depends on balancing these aspects through advances in computation, communication, and standardized cross-community protocols. Most importantly, the existing simulation or real-world benchmarks rarely treat robustness as an explicit evaluation criterion, instead evaluating performance primarily under relatively fixed conditions. Fair comparison further requires robustness evaluation methods that quantify system performance. While analytical and empirical methods have been developed for specific manipulation tasks (Section 6), a unified framework that systematically captures real-world uncertainty or variation remains an open challenge.

8.2 Analytical and Empirical Robustness↩︎

Consistent with the analytical and empirical evaluation protocols discussed in Section 6, these two methodological traditions also appear broadly in robustness research. Analytical robustness methods reflect Plato’s philosophy of abstraction, seeking general principles and guarantees through formal modeling and analysis; empirical methods follow Aristotle’s practical spirit, emphasizing performance validated through data and experience. The field continues to debate whether progress in robotics lies in one paradigm or their synthesis [206], and robustness as a core robotic topic follows the same divide. Data offer robots experiential knowledge that models alone cannot provide, enabling robustness to emerge from interaction, much as humans develop know-how before know-why through sensorimotor learning. Yet unlike human toddlers, robots are not bound to start from scratch. They can inherit models and priors designed by humans, effectively knowing why before knowing how. Such models allow robots to adapt behavior in a data-efficient way while mitigating the limitations of purely data-driven methods, which are often sensitive to aleatoric uncertainty and distribution shift. Consider the game of Jenga: humans learn to play robustly not only through trial and error but also by leveraging internal models of perception, planning, and control. Similarly, the future of manipulation robustness likely lies in unifying data-driven learning and model-based reasoning—integrating know-how and know-why into robot systems.

8.3 Towards Human-Level Manipulation Robustness↩︎

Achieving human-level manipulation robustness remains a grand challenge in robotics. Evidence suggests that such robustness emerges from the integration of multiple complementary strategies, such as extensive lifetime experience, accurate world models grounded in physical reasoning, domain-specific embodiment, and rapid adaptation. For example, crows bend hooks to retrieve food, effectively co-designing hardware and control in the wild [207]; such behavior cannot be explained by mechanics or control alone. Humans similarly integrate rich lifetime priors with online exploration and adaptation, inspiring approaches such as reinforcement learning bootstrapped with prior data [152]. Yet, synthesizing all these strategies in a single robotic system remains rare and challenging. Other principles evident in animals and humans, such as redundancy in planning and sensing, are gaining attention as essential components for manipulation robustness. Reaching animal-level manipulation robustness already exceeds current robotic capabilities, but is prospective; attaining human-level robustness represents a far greater leap and the ultimate aspiration for robotic manipulation.

9 Conclusion↩︎

This paper makes advances towards a systematic understanding of manipulation robustness by defining the concept, formulating the problem, and revisiting prior work across subfields. We revisited representative mechanisms and evaluation methods, highlighting the core principles that enable robustness in robotic manipulation and providing insights for future practitioners interested in this topic. Despite decades of progress, achieving human-level manipulation robustness on robotic systems under real-world uncertainties and variations remains a central challenge. Bridging this gap will require not only improved algorithms and hardware, but also clearer formulations of robustness and more consistent evaluation methods. We hope this paper will help organize future research efforts toward robotic manipulation systems that routinely approach the robustness demonstrated by humans and animals in the physical world.

Acknowledgements↩︎

The project is partially funded by the European Commission under the Horizon Europe Framework Program project SoftEnable, grant number 101070600. The authors thank Oliver Brock, Matthew T. Mason, Yan Zhang, Yunke Ao, Rafael I. Cabral Muchacho, Shaohang Han, Haoyu Li, Zizhe Zhang, and Jinda Cui for helpful discussions, comments, or proofreading. Part of Fig. 13-c were created using generative AI tools (ChatGPT/OpenAI). These figures are original representative illustrations intended for conceptual explanation. The authors verified their technical accuracy and take full responsibility for the final content.

Declaration of conflicting interests↩︎

The authors declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.

References↩︎

[1]
M. T. Mason, “Toward robotic manipulation,” Annual Review of Control, Robotics, and Autonomous Systems, vol. 1, no. 1, pp. 1–28, 2018.
[2]
J. R. Flanagan, M. C. Bowman, and R. S. Johansson, “Control strategies in object manipulation tasks,” Current opinion in neurobiology, vol. 16, no. 6, pp. 650–659, 2006.
[3]
A. M. Hadjiosif and M. A. Smith, “Flexible control of safety margins for action based on environmental variability,” Journal of Neuroscience, vol. 35, no. 24, pp. 9106–9121, 2015.
[4]
P. Foster-Turley and H. Markowitz, “A captive behavioral enrichment study with asian small-clawed river otters (aonyx cinerea),” Zoo biology, vol. 1, no. 1, pp. 29–43, 1982.
[5]
T. Fenchel, “Suspension feeding in ciliated protozoa: Functional response and particle size selection,” Microbial Ecology, vol. 6, no. 1, pp. 1–11, 1980.
[6]
H. Kitano, “Biological robustness,” Nature Reviews Genetics, vol. 5, no. 11, pp. 826–837, 2004.
[7]
H. Lu, Y. Dong, Z. Weng, F. Pokorny, J. Lundell, and D. Kragic, “Grasping a handful: Sequential multi-object dexterous grasp generation,” IEEE Robotics and Automation Letters, 2025.
[8]
Fruit Hunters, Accessed: 2026-03-13“Banana variety box.” https://fruithunters.com/products/banana-variety-box, 2026.
[9]
OnlineDelivery.in, Accessed: 2026-03-13“Mixed fruits with basket (2 kg).” https://www.onlinedelivery.in/mixed-fruits-with-basket-2-kg, 2024.
[10]
Berkeley AI Research, Image from BAIR blog on dexterous manipulation. Accessed: 2026-03-13“Dexterous manipulation blog image.” https://bair.berkeley.edu/static/blog/dex_manip/missing_img1.png, 2019.
[11]
B. Cao, B. Zhang, W. Zheng, J. Zhou, Y. Lin, and Y. Chen, “Real-time, highly accurate robotic grasp detection utilizing transfer learning for robots manipulating fragile fruits with widely variable sizes and shapes,” Computers and electronics in agriculture, vol. 200, p. 107254, 2022.
[12]
S. Sankar et al., “A natural biomimetic prosthetic hand with neuromorphic tactile sensing for precise and compliant grasping,” Science Advances, vol. 11, no. 10, p. eadr9300, 2025.
[13]
A. Billard and D. Kragic, “Trends and challenges in robot manipulation,” Science, vol. 364, no. 6446, p. eaat8414, 2019.
[14]
A. Rodriguez, “The unstable queen: Uncertainty, mechanics, and tactile feedback,” Science Robotics, vol. 6, no. 54, p. eabi4667, 2021.
[15]
J. Moos, K. Hansel, H. Abdulsamad, S. Stark, D. Clever, and J. Peters, “Robust reinforcement learning: A review of foundations and recent advances,” Machine Learning and Knowledge Extraction, vol. 4, no. 1, pp. 276–315, 2022.
[16]
S. Skogestad and I. Postlethwaite, Multivariable feedback control: Analysis and design. john Wiley & sons, 2005.
[17]
K. Ghazi-Zahedi, R. Deimel, G. Montúfar, V. Wall, and O. Brock, “Morphological computation: The good, the bad, and the ugly,” in 2017 IEEE/RSJ international conference on intelligent robots and systems (IROS), 2017, pp. 464–469.
[18]
J. Bohg, A. Morales, T. Asfour, and D. Kragic, “Data-driven grasp synthesis—a survey,” IEEE Transactions on robotics, vol. 30, no. 2, pp. 289–309, 2013.
[19]
O. Kroemer, S. Niekum, and G. Konidaris, “A review of robot learning for manipulation: Challenges, representations, and algorithms,” Journal of machine learning research, vol. 22, no. 30, pp. 1–82, 2021.
[20]
R. Firoozi et al., “Foundation models in robotics: Applications, challenges, and the future,” The International Journal of Robotics Research, vol. 44, no. 5, pp. 701–739, 2025.
[21]
H. B. Braiek and F. Khomh, “Machine learning robustness: A primer,” in Trustworthy AI in medical imaging, Elsevier, 2025, pp. 37–71.
[22]
M. Baum, Robustness in robotic and biological manipulation. Technische Universitaet Berlin (Germany), 2024.
[23]
A. Madry, A. Makelov, L. Schmidt, D. Tsipras, and A. Vladu, “Towards deep learning models resistant to adversarial attacks,” International Conference on Learning Representations, 2018.
[24]
D. Hendrycks and T. Dietterich, “Benchmarking neural network robustness to common corruptions and perturbations,” in International conference on learning representations, 2019.
[25]
R. C. Goertz, “Fundamentals of general-purpose remote manipulators,” Nucleonics, pp. 36–42, 1952.
[26]
M. T. Mason, “Creation myths: The beginnings of robotics research,” IEEE robotics & automation magazine, vol. 19, no. 2, pp. 72–77, 2012.
[27]
J. Cui and J. Trinkle, “Toward next-generation learned robot manipulation,” Science robotics, vol. 6, no. 54, p. eabd9461, 2021.
[28]
M. Yim, D. G. Duff, and K. D. Roufas, “PolyBot: A modular reconfigurable robot,” in Proceedings 2000 ICRA. Millennium conference. IEEE international conference on robotics and automation. Symposia proceedings (cat. No. 00CH37065), 2000, vol. 1, pp. 514–520.
[29]
H. Xie et al., “Reconfigurable magnetic microrobot swarm: Multimode transformation, locomotion, and manipulation,” Science robotics, vol. 4, no. 28, p. eaav8006, 2019.
[30]
A. Zeng et al., “Robotic pick-and-place of novel objects in clutter with multi-affordance grasping and cross-domain image matching,” The International Journal of Robotics Research, vol. 41, no. 7, pp. 690–705, 2022.
[31]
T. Siméon, J.-P. Laumond, J. Cortés, and A. Sahbani, “Manipulation planning with probabilistic roadmaps,” The International Journal of Robotics Research, vol. 23, no. 7–8, pp. 729–746, 2004.
[32]
A. Gelblum, I. Pinkoviezky, E. Fonio, A. Ghosh, N. Gov, and O. Feinerman, “Ant groups optimally amplify the effect of transiently informed individuals,” Nature communications, vol. 6, no. 1, p. 7729, 2015.
[33]
R. Martı́n-Martı́n, Leveraging problem structure in interactive perception for robot manipulation of constrained mechanisms. Technische Universitaet Berlin (Germany), 2018.
[34]
P. R. Florence, L. Manuelli, and R. Tedrake, “Dense object nets: Learning dense visual object descriptors by and for robotic manipulation,” in Conference on robot learning, 2018, pp. 373–385.
[35]
H. Li et al., “See, hear, and feel: Smart sensory fusion for robotic manipulation,” in Conference on robot learning, 2023, pp. 1368–1378.
[36]
J. Aloimonos, I. Weiss, and A. Bandyopadhyay, “Active vision,” International journal of computer vision, vol. 1, no. 4, pp. 333–356, 1988.
[37]
R. Bajcsy, “Active perception,” Proceedings of the IEEE, vol. 76, no. 8, pp. 966–1005, 1988.
[38]
D. Morrison, P. Corke, and J. Leitner, “Multi-view picking: Next-best-view reaching for improved grasping in clutter,” in 2019 international conference on robotics and automation (ICRA), 2019, pp. 8762–8768.
[39]
J. Bohg et al., “Interactive perception: Leveraging action in perception and perception in action,” IEEE Transactions on Robotics, vol. 33, no. 6, pp. 1273–1291, 2017.
[40]
H. Xiong, X. Xu, J. Wu, Y. Hou, J. Bohg, and S. Song, “Vision in action: Learning active perception from human demonstrations,” in Conference on robot learning, 2025, pp. 5450–5463.
[41]
J. Chappell, Z. P. Demery, V. Arriola-Rios, and A. Sloman, “How to build an information gathering and processing system: Lessons from naturally and artificially intelligent systems,” Behavioural Processes, vol. 89, no. 2, pp. 179–186, 2012.
[42]
K. Xu et al., “Autoscanning for coupled scene reconstruction and proactive object analysis,” ACM Transactions on Graphics (TOG), vol. 34, no. 6, pp. 1–14, 2015.
[43]
D. Katz and O. Brock, “Manipulating articulated objects with interactive perception,” in 2008 IEEE international conference on robotics and automation, 2008, pp. 272–277.
[44]
A. Krizhevsky, I. Sutskever, and G. E. Hinton, “Imagenet classification with deep convolutional neural networks,” Advances in neural information processing systems, vol. 25, 2012.
[45]
D. G. Lowe, “Distinctive image features from scale-invariant keypoints,” International journal of computer vision, vol. 60, no. 2, pp. 91–110, 2004.
[46]
H. Bay, T. Tuytelaars, and L. Van Gool, “Surf: Speeded up robust features,” in European conference on computer vision, 2006, pp. 404–417.
[47]
E. Rublee, V. Rabaud, K. Konolige, and G. Bradski, “ORB: An efficient alternative to SIFT or SURF,” in 2011 international conference on computer vision, 2011, pp. 2564–2571.
[48]
Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner, “Gradient-based learning applied to document recognition,” Proceedings of the IEEE, vol. 86, no. 11, pp. 2278–2324, 2002.
[49]
M. Oquab et al., “DINOv2: Learning robust visual features without supervision,” Transactions on Machine Learning Research Journal, 2024.
[50]
B. Zitkovich et al., “Rt-2: Vision-language-action models transfer web knowledge to robotic control,” in Conference on robot learning, 2023, pp. 2165–2183.
[51]
G. Izatt, G. Mirano, E. Adelson, and R. Tedrake, “Tracking objects with point clouds from vision and touch,” in 2017 IEEE international conference on robotics and automation (ICRA), 2017, pp. 4000–4007.
[52]
B. S. Homberg, R. K. Katzschmann, M. R. Dogar, and D. Rus, “Robust proprioceptive grasping with a soft robot hand,” Autonomous robots, vol. 43, no. 3, pp. 681–696, 2019.
[53]
R. Platt, R. Tedrake, L. Kaelbling, and T. Lozano-Perez, “Belief space planning assuming maximum likelihood observations,” Robotics: Science and Systems VI, 2010.
[54]
J. Jankowski, L. Brudermüller, N. Hawes, and S. Calinon, “Robust pushing: Exploiting quasi-static belief dynamics and contact-informed optimization,” The International Journal of Robotics Research, p. 02783649251318046, 2024.
[55]
L. P. Kaelbling and T. Lozano-Pérez, “Integrated task and motion planning in belief space,” The International Journal of Robotics Research, vol. 32, no. 9–10, pp. 1194–1227, 2013.
[56]
M. A. Erdmann and M. T. Mason, “An exploration of sensorless manipulation,” IEEE Journal on Robotics and Automation, vol. 4, no. 4, pp. 369–379, 2002.
[57]
M. Dogar, K. Hsiao, M. Ciocarlie, and S. Srinivasa, “Physics-based grasp planning through clutter,” Robotics: Science and System, pp. 57–64, 2012.
[58]
M. A. Peshkin and A. C. Sanderson, “Planning robotic manipulation strategies for workpieces that slide,” IEEE Journal on Robotics and Automation, vol. 4, no. 5, pp. 524–531, 2002.
[59]
N. C. Dafle et al., “Extrinsic dexterity: In-hand manipulation with external forces,” in 2014 IEEE international conference on robotics and automation (ICRA), 2014, pp. 1578–1585.
[60]
M. Mason, “The mechanics of manipulation,” in Proceedings. 1985 IEEE international conference on robotics and automation, 1985, vol. 2, pp. 544–548.
[61]
T. Lozano-Perez, M. T. Mason, and R. H. Taylor, “Automatic synthesis of fine-motion strategies for robots,” The International Journal of Robotics Research, vol. 3, no. 1, pp. 3–24, 1984.
[62]
R. Tedrake, “LQR-trees: Feedback motion planning on sparse randomized trees,” Robotics: Science and Systems V, 2009.
[63]
A. Bhatt, A. Sieler, S. Puhlmann, and O. Brock, “Surprisingly robust in-hand manipulation: An empirical study,” Robotics: Science and Systems XVII, 2021.
[64]
M. C. Koval, N. S. Pollard, and S. S. Srinivasa, “Pre-and post-contact policy decomposition for planar contact manipulation under uncertainty,” The International Journal of Robotics Research, vol. 35, no. 1–3, pp. 244–264, 2016.
[65]
D. Fox, “Grasping.” Lecture slides, CSE 478: Robot Learning, University of Washington, 2020, [Online]. Available: https://courses.cs.washington.edu/courses/cse478/20wi/site/resources/lec23_grasping.pdf.
[66]
A. Rodriguez, M. T. Mason, and S. Ferry, “From caging to grasping,” The International Journal of Robotics Research, vol. 31, no. 7, pp. 886–900, 2012.
[67]
D. Berenson, S. S. Srinivasa, D. Ferguson, and J. J. Kuffner, “Manipulation planning on constraint manifolds,” in 2009 IEEE international conference on robotics and automation, 2009, pp. 625–632.
[68]
M. Saha and P. Isto, “Manipulation planning for deformable linear objects,” IEEE Transactions on Robotics, vol. 23, no. 6, pp. 1141–1150, 2007.
[69]
L. Blackmore, M. Ono, and B. C. Williams, “Chance-constrained optimal path planning with obstacles,” IEEE Transactions on Robotics, vol. 27, no. 6, pp. 1080–1094, 2011.
[70]
Y. Shirai, D. K. Jha, A. Raghunathan, and D. Romeres, “Chance-constrained optimization in contact-rich systems for robust manipulation,” arXiv preprint arXiv:2203.02616, 2022.
[71]
D. Q. Mayne, M. M. Seron, and S. V. Raković, “Robust model predictive control of constrained linear systems with bounded disturbances,” Automatica, vol. 41, no. 2, pp. 219–224, 2005.
[72]
W. S. Howard and V. Kumar, “On the stability of grasped objects,” IEEE transactions on robotics and automation, vol. 12, no. 6, pp. 904–917, 1996.
[73]
A. Bicchi, “On the closure properties of robotic grasping,” The International Journal of Robotics Research, vol. 14, no. 4, pp. 319–334, 1995.
[74]
G. A. Pereira, M. F. Campos, and V. Kumar, “Decentralized algorithms for multi-robot manipulation via caging,” The International Journal of Robotics Research, vol. 23, no. 7–8, pp. 783–795, 2004.
[75]
M. T. Mason, “Compliance and force control for computer controlled manipulators,” IEEE Transactions on Systems, Man, and Cybernetics, vol. 11, no. 6, pp. 418–432, 2007.
[76]
N. Hogan, “Impedance control of industrial robots,” Robotics and computer-integrated manufacturing, vol. 1, no. 1, pp. 97–113, 1984.
[77]
F. Dimeas and N. Aspragathos, “Reinforcement learning of variable admittance control for human-robot co-manipulation,” in 2015 IEEE/RSJ international conference on intelligent robots and systems (IROS), 2015, pp. 1011–1016.
[78]
X. Xu, Y. Hou, Z. Liu, and S. Song, “Compliant residual DAgger: Improving real-world contact-rich manipulation with human corrections,” The Thirty-ninth Annual Conference on Neural Information Processing Systems, 2026.
[79]
Y. Hou et al., “Adaptive compliance policy: Learning approximate compliance for diffusion guided control,” in 2025 IEEE international conference on robotics and automation (ICRA), 2025, pp. 4829–4836.
[80]
T. Kamijo, C. C. Beltran-Hernandez, and M. Hamaya, “Learning variable compliance control from a few demonstrations for bimanual robot with haptic feedback teleoperation system,” in 2024 IEEE/RSJ international conference on intelligent robots and systems (IROS), 2024, pp. 12663–12670.
[81]
M. P. Polverini, D. Nicolis, A. M. Zanchettin, and P. Rocco, “Implicit robot force control based on set invariance,” IEEE Robotics and Automation Letters, vol. 2, no. 3, pp. 1288–1295, 2017.
[82]
P. Holmes et al., “Reachable sets for safe, real-time manipulator trajectory design,” Robotics: Science and Systems, 2020.
[83]
G. Wang, K. Ren, A. S. Morgan, and K. Hang, “Caging in time: A framework for robust object manipulation under uncertainties and limited robot perception,” The International Journal of Robotics Research, p. 02783649251343926, 2025.
[84]
Y. Jiang, M. Yu, X. Zhu, M. Tomizuka, and X. Li, “Robust model-based in-hand manipulation with integrated real-time motion-contact planning and tracking,” arXiv preprint arXiv:2505.04978, 2025.
[85]
H. T. Suh, T. Pang, T. Zhao, and R. Tedrake, “Dexterous contact-rich manipulation via the contact trust region,” The International Journal of Robotics Research, p. 02783649251398875, 2025.
[86]
J. Nubert, J. Köhler, V. Berenz, F. Allgöwer, and S. Trimpe, “Safe and fast tracking on a robot manipulator: Robust mpc and neural network control,” IEEE Robotics and Automation Letters, vol. 5, no. 2, pp. 3050–3057, 2020.
[87]
J. Tobin, R. Fong, A. Ray, J. Schneider, W. Zaremba, and P. Abbeel, “Domain randomization for transferring deep neural networks from simulation to the real world,” in 2017 IEEE/RSJ international conference on intelligent robots and systems (IROS), 2017, pp. 23–30.
[88]
OpenAI et al., “Solving rubik’s cube with a robot hand,” arXiv preprint arXiv:1910.07113, 2019.
[89]
O. M. Andrychowicz et al., “Learning dexterous in-hand manipulation,” The International Journal of Robotics Research, vol. 39, no. 1, pp. 3–20, 2020.
[90]
P. Florence, L. Manuelli, and R. Tedrake, “Self-supervised correspondence in visuomotor policy learning,” IEEE Robotics and Automation Letters, vol. 5, no. 2, pp. 492–499, 2019.
[91]
A. Mandlekar et al., “MimicGen: A data generation system for scalable robot learning using human demonstrations,” in Conference on robot learning, 2023, pp. 1820–1864.
[92]
E. Ameperosa, J. A. Collins, M. Jain, and A. Garg, “Rocoda: Counterfactual data augmentation for data-efficient robot learning from demonstrations,” in 2025 IEEE international conference on robotics and automation (ICRA), 2025, pp. 13250–13256.
[93]
A. Zhou, M. J. Kim, L. Wang, P. Florence, and C. Finn, “Nerf in the palm of your hand: Corrective augmentation for robotics via novel-view synthesis,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2023, pp. 17907–17917.
[94]
X. Zhang, M. Chang, P. Kumar, and S. Gupta, “Diffusion meets DAgger: Supercharging eye-in-hand imitation learning,” in Robotics science and systems, 2024.
[95]
T. Yu et al., “Scaling robot learning with semantically imagined experience,” Robotics: Science and Systems, 2023.
[96]
Z. Chen, S. Kiami, A. Gupta, and V. Kumar, “Genaug: Retargeting behaviors to unseen situations via generative augmentation,” Robotics: Science and Systems, 2023.
[97]
Z. Xue, S. Deng, Z. Chen, Y. Wang, Z. Yuan, and H. Xu, “Demogen: Synthetic demonstration generation for data-efficient visuomotor policy learning,” Robotics: Science and Systems, 2025.
[98]
M. Laskey, J. Lee, R. Fox, A. Dragan, and K. Goldberg, “Dart: Noise injection for robust imitation learning,” in Conference on robot learning, 2017, pp. 143–156.
[99]
L. Ke, J. Wang, T. Bhattacharjee, B. Boots, and S. Srinivasa, “Grasping with chopsticks: Combating covariate shift in model-free imitation learning for fine manipulation,” in 2021 IEEE international conference on robotics and automation (ICRA), 2021, pp. 6185–6191.
[100]
M. Simchowitz, D. Pfrommer, and A. Jadbabaie, “The pitfalls of imitation learning when actions are continuous,” arXiv preprint arXiv:2503.09722, 2025.
[101]
A. Brohan et al., “RT-1: Robotics transformer for real-world control at scale,” Robotics: Science and Systems XIX, 2023.
[102]
Open X-Embodiment Collaboration, “Open x-embodiment: Robotic learning datasets and rt-x models: Open x-embodiment collaboration 0,” in 2024 IEEE international conference on robotics and automation (ICRA), 2024, pp. 6892–6903.
[103]
M. J. Kim et al., “OpenVLA: An open-source vision-language-action model,” in Conference on robot learning, 2025, pp. 2679–2713.
[104]
TRI LBM Team et al., “A careful examination of large behavior models for multitask dexterous manipulation,” Science Robotics, vol. 11, no. 113, p. eaea6201, 2026.
[105]
L. Manuelli, W. Gao, P. Florence, and R. Tedrake, “Kpam: Keypoint affordances for category-level robotic manipulation,” in The international symposium of robotics research, 2019, pp. 132–157.
[106]
Z. Qin, K. Fang, Y. Zhu, L. Fei-Fei, and S. Savarese, “Keto: Learning keypoint representations for tool manipulation,” in 2020 IEEE international conference on robotics and automation (ICRA), 2020, pp. 7278–7285.
[107]
W. Huang, C. Wang, Y. Li, R. Zhang, and L. Fei-Fei, “ReKep: Spatio-temporal reasoning of relational keypoint constraints for robotic manipulation,” in Conference on robot learning, 2025, pp. 4573–4602.
[108]
A. Simeonov et al., “Neural descriptor fields: Se (3)-equivariant object representations for manipulation,” in 2022 international conference on robotics and automation (ICRA), 2022, pp. 6394–6400.
[109]
H. Ryu, H. Lee, J.-H. Lee, and J. Choi, “Equivariant descriptor fields: Se (3)-equivariant energy-based models for end-to-end visual robotic manipulation learning,” arXiv preprint arXiv:2206.08321, 2022.
[110]
B. Eisner, Y. Yang, T. Davchev, M. Vecerik, J. Scholz, and D. Held, “Deep SE (3)-equivariant geometric reasoning for precise placement tasks,” in The twelfth international conference on learning representations, 2024.
[111]
D. Wang et al., “Equivariant diffusion policy,” in Conference on robot learning, 2025, pp. 48–69.
[112]
R. Li, A. Jabri, T. Darrell, and P. Agrawal, “Towards practical multi-object manipulation using relational reinforcement learning,” in 2020 ieee international conference on robotics and automation (icra), 2020, pp. 4051–4058.
[113]
Y. Lin, A. S. Wang, E. Undersander, and A. Rai, “Efficient and interpretable robot manipulation with graph neural networks,” IEEE Robotics and Automation Letters, vol. 7, no. 2, pp. 2740–2747, 2022.
[114]
Y. Huang, A. Conkey, and T. Hermans, “Planning for multi-object manipulation with graph neural network relational classifiers,” in 2023 IEEE international conference on robotics and automation (ICRA), 2023, pp. 1822–1829.
[115]
C. Chi et al., “Diffusion policy: Visuomotor policy learning via action diffusion,” The International Journal of Robotics Research, p. 02783649241273668, 2023.
[116]
K. B. A. N. B. A. D. D. A. A. E. A. M. R. E. A. C. F. A. N. F. A. L. G. A. K. H. A. B. I. A. S. J. A. T. J. A. L. K. A. S. L. A. A. L.-B. A. M. M. A. S. N. A. K. P. A. L. X. S. A. L. S. A. J. T. A. Q. V. A. A. W. A. H. W. A. U. Zhilinsky, \(\pi\)0: A Vision-Language-Action Flow Model for General Robot Control,” Proceedings of Robotics: Science and Systems, 2025.
[117]
C. Pan et al., “Much ado about noising: Dispelling the myths of generative robotic control,” International Conference on Learning Representations, 2026.
[118]
R. S. Sutton, D. Precup, and S. Singh, “Intra-option learning about temporally abstract actions.” in ICML, 1998, vol. 98, pp. 556–564.
[119]
O. Nachum, S. S. Gu, H. Lee, and S. Levine, “Data-efficient hierarchical reinforcement learning,” Advances in neural information processing systems, vol. 31, 2018.
[120]
D. Driess, J.-S. Ha, and M. Toussaint, “Learning to solve sequential physical reasoning problems from a scene image,” The International Journal of Robotics Research, vol. 40, no. 12–14, pp. 1435–1466, 2021.
[121]
Y. Zhu, P. Stone, and Y. Zhu, “Bottom-up skill discovery from unsegmented demonstrations for long-horizon robot manipulation,” IEEE Robotics and Automation Letters, vol. 7, no. 2, pp. 4126–4133, 2022.
[122]
U. A. Mishra, S. Xue, Y. Chen, and D. Xu, “Generative skill chaining: Long-horizon skill planning with diffusion models,” in Conference on robot learning, 2023, pp. 2905–2925.
[123]
D. Hafner et al., “Learning latent dynamics for planning from pixels,” in International conference on machine learning, 2019, pp. 2555–2565.
[124]
B. Ichter and M. Pavone, “Robot motion planning in learned latent spaces,” IEEE Robotics and Automation Letters, vol. 4, no. 3, pp. 2407–2414, 2019.
[125]
N. Hansen, X. Wang, and H. Su, “Temporal difference learning for model predictive control,” in International conference on machine learning, PMLR, 2022.
[126]
D. Hafner, J. Pasukonis, J. Ba, and T. Lillicrap, “Mastering diverse domains through world models,” arXiv preprint arXiv:2301.04104, 2023.
[127]
H. Liu, S. Dass, R. Martı́n-Martı́n, and Y. Zhu, “Model-based runtime monitoring with interactive imitation learning,” in 2024 IEEE international conference on robotics and automation (ICRA), 2024, pp. 4154–4161.
[128]
K. Nakamura, L. Peters, and A. Bajcsy, “Generalizing safety beyond collision-avoidance via latent-space reachability analysis,” arXiv preprint arXiv:2502.00935, 2025.
[129]
Z. Sun and S. Song, “Latent policy barrier: Learning robust visuomotor policies by staying in-distribution,” Advances in Neural Information Processing Systems, vol. 38, pp. 174280–174305, 2026.
[130]
L. Pinto, J. Davidson, R. Sukthankar, and A. Gupta, “Robust adversarial reinforcement learning,” in International conference on machine learning, 2017, pp. 2817–2826.
[131]
L. Pinto, J. Davidson, and A. Gupta, “Supervision via competition: Robot adversaries for learning tasks,” in 2017 IEEE international conference on robotics and automation (ICRA), 2017, pp. 1601–1608.
[132]
P. Jian, C. Yang, D. Guo, H. Liu, and F. Sun, “Adversarial skill learning for robust manipulation,” in 2021 IEEE international conference on robotics and automation (ICRA), 2021, pp. 2555–2561.
[133]
J. Achiam, D. Held, A. Tamar, and P. Abbeel, “Constrained policy optimization,” in International conference on machine learning, 2017, pp. 22–31.
[134]
G. Dalal, K. Dvijotham, M. Vecerik, T. Hester, C. Paduraru, and Y. Tassa, “Safe exploration in continuous action spaces,” arXiv preprint arXiv:1801.08757, 2018.
[135]
R. Cheng, G. Orosz, R. M. Murray, and J. W. Burdick, “End-to-end safe reinforcement learning through barrier functions for safety-critical continuous control tasks,” in Proceedings of the AAAI conference on artificial intelligence, 2019, vol. 33, pp. 3387–3395.
[136]
S. Levine, A. Kumar, G. Tucker, and J. Fu, “Offline reinforcement learning: Tutorial, review, and perspectives on open problems,” arXiv preprint arXiv:2005.01643, 2020.
[137]
A. Kumar, A. Zhou, G. Tucker, and S. Levine, “Conservative q-learning for offline reinforcement learning,” Advances in neural information processing systems, vol. 33, pp. 1179–1191, 2020.
[138]
I. Kostrikov, A. Nair, and S. Levine, “Offline reinforcement learning with implicit q-learning,” in International conference on learning representations, 2021.
[139]
S. Fujimoto and S. S. Gu, “A minimalist approach to offline reinforcement learning,” Advances in neural information processing systems, vol. 34, pp. 20132–20145, 2021.
[140]
G. Zhou, L. Ke, S. Srinivasa, A. Gupta, A. Rajeswaran, and V. Kumar, “Real world offline reinforcement learning with realistic data source,” in 2023 IEEE international conference on robotics and automation (ICRA), 2023, pp. 7176–7183.
[141]
A. Herzog et al., “Deep RL at scale: Sorting waste in office buildings with a fleet of mobile manipulators,” Robotics: Science and Systems (RSS), 2023.
[142]
S. Haldar, V. Mathur, D. Yarats, and L. Pinto, “Watch and match: Supercharging imitation with regularized optimal transport,” in Conference on robot learning, 2023, pp. 32–43.
[143]
H. Liu, S. Nasiriany, L. Zhang, Z. Bao, and Y. Zhu, “Robot learning on the job: Human-in-the-loop autonomy and learning during deployment,” The International Journal of Robotics Research, p. 02783649241273901, 2022.
[144]
S. Ross, G. Gordon, and D. Bagnell, “A reduction of imitation learning and structured prediction to no-regret online learning,” in Proceedings of the fourteenth international conference on artificial intelligence and statistics, 2011, pp. 627–635.
[145]
M. Kelly, C. Sidrane, K. Driggs-Campbell, and M. J. Kochenderfer, “Hg-dagger: Interactive imitation learning with human experts,” in 2019 international conference on robotics and automation (ICRA), 2019, pp. 8077–8083.
[146]
A. Mandlekar, D. Xu, R. Martı́n-Martı́n, Y. Zhu, L. Fei-Fei, and S. Savarese, “Human-in-the-loop imitation learning using remote teleoperation,” arXiv preprint arXiv:2012.06733, 2020.
[147]
S. Levine, C. Finn, T. Darrell, and P. Abbeel, “End-to-end training of deep visuomotor policies,” Journal of Machine Learning Research, vol. 17, no. 39, pp. 1–40, 2016.
[148]
A. Rajeswaran et al., “Learning complex dexterous manipulation with deep reinforcement learning and demonstrations,” Robotics: Science and Systems XIV, 2018.
[149]
T. Johannink et al., “Residual reinforcement learning for robot control,” in 2019 international conference on robotics and automation (ICRA), 2019, pp. 6023–6029.
[150]
K. Xu et al., “Dexterous manipulation from images: Autonomous real-world rl via substep guidance,” in 2023 IEEE international conference on robotics and automation (ICRA), 2023, pp. 5938–5945.
[151]
J. Luo et al., “Serl: A software suite for sample-efficient robotic reinforcement learning,” in 2024 IEEE international conference on robotics and automation (ICRA), 2024, pp. 16961–16969.
[152]
J. Luo, C. Xu, J. Wu, and S. Levine, “Precise and dexterous robotic manipulation via human-in-the-loop reinforcement learning,” Science Robotics, vol. 10, no. 105, p. eads5033, 2025.
[153]
Y. Duan et al., “One-shot imitation learning,” Advances in neural information processing systems, vol. 30, 2017.
[154]
E. Valassakis, G. Papagiannis, N. Di Palo, and E. Johns, “Demonstrate once, imitate immediately (dome): Learning visual servoing for one-shot imitation learning,” in 2022 IEEE/RSJ international conference on intelligent robots and systems (IROS), 2022, pp. 8614–8621.
[155]
L. Fu et al., “In-context imitation learning via next-token prediction,” 2025 IEEE international conference on robotics and automation (ICRA), 2025.
[156]
C. Paul, “Morphological computation: A basis for the analysis of morphology and control requirements,” Robotics and Autonomous Systems, vol. 54, no. 8, pp. 619–630, 2006.
[157]
J. Mann et al., “Why do dolphins carry sponges?” PloS one, vol. 3, no. 12, p. e3868, 2008.
[158]
S. H. Drake, “Using compliance in lieu of sensory feedback for automatic assembly.” PhD thesis, Massachusetts Institute of Technology, 1978.
[159]
R. Deimel and O. Brock, “A compliant hand based on a novel pneumatic actuator,” in 2013 IEEE international conference on robotics and automation, 2013, pp. 2047–2053.
[160]
J. Shintake, V. Cacucciolo, D. Floreano, and H. Shea, “Soft robotic grippers,” Advanced materials, vol. 30, no. 29, p. 1707035, 2018.
[161]
J. R. Amend, E. Brown, N. Rodenberg, H. M. Jaeger, and H. Lipson, “A positive pressure universal gripper based on the jamming of granular material,” IEEE transactions on robotics, vol. 28, no. 2, pp. 341–350, 2012.
[162]
S. Song, D.-M. Drotlef, C. Majidi, and M. Sitti, “Controllable load sharing for soft adhesive interfaces on three-dimensional surfaces,” Proceedings of the National Academy of Sciences, vol. 114, no. 22, pp. E4344–E4353, 2017.
[163]
K. Autumn, Copyright © 2006 Kellar Autumn. Accessed: 2026-06-21“Tokay gecko foot image.” https://people.eecs.berkeley.edu/~ronf/Gecko/Interface-slide-adhesion/TokayFoot2-KA.jpg, 2006.
[164]
H. Jiang et al., “A robotic device using gecko-inspired adhesives can grasp and manipulate large objects in microgravity,” Science Robotics, vol. 2, no. 7, p. eaan4545, 2017.
[165]
G. Liang, A. J. Ijspeert, M. Yim, and T. L. Lam, “Modular reconfigurable robots: Toward on-demand multifunctional applications,” Science Robotics, vol. 11, no. 111, p. eadz1999, 2026.
[166]
A. Pagoli, F. Chapelle, J. A. Corrales, Y. Mezouar, and Y. Lapusta, “A soft robotic gripper with an active palm and reconfigurable fingers for fully dexterous in-hand manipulation,” IEEE Robotics and Automation Letters, vol. 6, no. 4, pp. 7706–7713, 2021.
[167]
D. Liconti, Y. Zhou, Y. Toshimitsu, R. Hinchet, and R. K. Katzschmann, “A benchmark of dexterity for anthropomorphic robotic hands,” arXiv preprint arXiv:2604.09294, 2026.
[168]
X. Kang, T. Tian, S.-W. Lee, B. Huang, Y. Li, and Y.-L. Kuo, “Learning force-regulated manipulation with a low-cost tactile-force-controlled gripper,” arXiv preprint arXiv:2602.10013, 2026.
[169]
H. Xue et al., “Reactive diffusion policy: Slow-fast visual-tactile policy learning for contact-rich manipulation,” Robotics: Science and Systems, 2025.
[170]
H. Kress-Gazit et al., “Robot learning as an empirical science: Best practices for policy evaluation,” arXiv preprint arXiv:2409.09491, 2024.
[171]
C. Ferrari and J. Canny, “Planning optimal grasps,” in Proceedings., 1992 IEEE international conference on robotics and automation, 1992., 1992, vol. 3, pp. 2290–2295.
[172]
J. Mahler et al., “Dex-net 2.0: Deep learning to plan robust grasps with synthetic point clouds and analytic grasp metrics,” Robotics: Science and Systems XIII, 2017.
[173]
F. T. Pokorny and D. Kragic, “Data-driven topological motion planning with persistent cohomology,” 2015 Robotics: Science and Systems Conference, vol. 11, 2015.
[174]
S. Bhattacharya, R. Ghrist, and V. Kumar, “Persistent homology for path planning in uncertain environments,” IEEE Transactions on Robotics, vol. 31, no. 3, pp. 578–590, 2015.
[175]
J. Mahler, F. T. Pokorny, S. Niyaz, and K. Goldberg, “Synthesis of energy-bounded planar caging grasps using persistent homology,” IEEE Transactions on Automation Science and Engineering, vol. 15, no. 3, pp. 908–918, 2018.
[176]
C. J. Hasson, T. Shen, and D. Sternad, “Energy margins in dynamic object manipulation,” Journal of Neurophysiology, vol. 108, no. 5, pp. 1349–1365, 2012.
[177]
D. Sternad and C. J. Hasson, “Predictability and robustness in the manipulation of dynamically complex objects,” Progress in Motor Control: Theories and Translations, pp. 55–77, 2016.
[178]
Y. Dong, X. Cheng, and F. T. Pokorny, “Characterizing manipulation robustness through energy margin and caging analysis,” IEEE Robotics and Automation Letters, vol. 9, no. 9, pp. 7525–7532, 2024.
[179]
M. A. Roa and R. Suárez, “Grasp quality measures: Review and performance,” Autonomous robots, vol. 38, no. 1, pp. 65–88, 2015.
[180]
M. Fox, R. Howey, and D. Long, “Exploration of the robustness of plans,” in AAAI, 2006, pp. 834–839.
[181]
O. Maler and D. Nickovic, “Monitoring temporal properties of continuous signals,” in International symposium on formal techniques in real-time and fault-tolerant systems, 2004, pp. 152–166.
[182]
A. Donzé and O. Maler, “Robust satisfaction of temporal logic over real-valued signals,” in International conference on formal modeling and analysis of timed systems, 2010, pp. 92–106.
[183]
P. Kapoor, A. Balakrishnan, and J. V. Deshmukh, “Model-based reinforcement learning from signal temporal logic specifications,” arXiv preprint arXiv:2011.04950, 2020.
[184]
X. Li, C.-I. Vasile, and C. Belta, “Reinforcement learning with temporal logic rewards,” in 2017 IEEE/RSJ international conference on intelligent robots and systems (IROS), 2017, pp. 3834–3839.
[185]
Y. Meng and C. Fan, “Signal temporal logic neural predictive control,” IEEE Robotics and Automation Letters, vol. 8, no. 11, pp. 7719–7726, 2023.
[186]
R. Takano, H. Oyama, and M. Yamakita, “Continuous optimization-based task and motion planning with signal temporal logic specifications for sequential manipulation,” in 2021 IEEE international conference on robotics and automation (ICRA), 2021, pp. 8409–8415.
[187]
P. Varnai and D. V. Dimarogonas, “On robustness metrics for learning STL tasks,” in 2020 american control conference (ACC), 2020, pp. 5394–5399.
[188]
A. Dhonthi, P. Schillinger, L. Rozo, and D. Nardi, “Study of signal temporal logic robustness metrics for robotic tasks optimization,” arXiv preprint arXiv:2110.00339, 2021.
[189]
S. Bazzi and D. Sternad, “Robustness in human manipulation of dynamically complex objects through control contraction metrics,” IEEE robotics and automation letters, vol. 5, no. 2, pp. 2578–2585, 2020.
[190]
K. C. Catania and J. H. Kaas, “The unusual nose and brain of the star-nosed mole,” Bioscience, vol. 46, no. 8, pp. 578–586, 1996.
[191]
L. Brunke et al., “Safe learning in robotics: From learning-based control to safe reinforcement learning,” Annual Review of Control, Robotics, and Autonomous Systems, vol. 5, no. 1, pp. 411–444, 2022.
[192]
E. Aljalbout et al., “The reality gap in robotics: Challenges, solutions, and best practices,” Annual Review of Control, Robotics, and Autonomous Systems, vol. 9, 2025.
[193]
J. Gao, S. Belkhale, S. Dasari, A. Balakrishna, D. Shah, and D. Sadigh, “A taxonomy for evaluating generalist robot manipulation policies,” IEEE Robotics and Automation Letters, 2026.
[194]
A. M. Johnson, S. A. Burden, and D. E. Koditschek, “A hybrid systems model for simple manipulation and self-manipulation systems,” The International Journal of Robotics Research, vol. 35, no. 11, pp. 1354–1392, 2016.
[195]
F. Shi, C. Zhang, T. Miki, J. Lee, M. Hutter, and S. Coros, “Rethinking robustness assessment: Adversarial attacks on learning-based quadrupedal locomotion controllers,” Robotics: Science and System XX, 2024.
[196]
K. To, P.-Y. Lai, and H. Pak, “Jamming of granular flow in a two-dimensional hopper,” Physical review letters, vol. 86, no. 1, p. 71, 2001.
[197]
S. Tao et al., “Maniskill3: Gpu parallelized robotics simulation and rendering for generalizable embodied ai,” arXiv preprint arXiv:2410.00425, 2024.
[198]
S. James, Z. Ma, D. R. Arrojo, and A. J. Davison, “Rlbench: The robot learning benchmark & learning environment,” IEEE Robotics and Automation Letters, vol. 5, no. 2, pp. 3019–3026, 2020.
[199]
Y. Zhu et al., “Robosuite: A modular simulation framework and benchmark for robot learning,” arXiv preprint arXiv:2009.12293, 2020.
[200]
B. Calli et al., “Yale-CMU-berkeley dataset for robotic manipulation research,” The International Journal of Robotics Research, vol. 36, no. 3, pp. 261–268, 2017.
[201]
N. Correll et al., “Analysis and observations from the first amazon picking challenge,” IEEE Transactions on Automation Science and Engineering, vol. 15, no. 1, pp. 172–188, 2016.
[202]
A. Kasper, Z. Xue, and R. Dillmann, “The kit object models database: An object model database for object recognition, localization and manipulation in service robotics,” The International Journal of Robotics Research, vol. 31, no. 8, pp. 927–934, 2012.
[203]
M. Zahid and F. T. Pokorny, “Cloudgripper: An open source cloud robotics testbed for robotic manipulation research, benchmarking and data collection at scale,” in 2024 IEEE international conference on robotics and automation (ICRA), 2024, pp. 12076–12082.
[204]
Y. Chen et al., “ManipulationNet: An infrastructure for benchmarking real-world robot manipulation with physical skill challenges and embodied multimodal reasoning,” arXiv preprint arXiv:2603.04363, 2026.
[205]
P. Atreya et al., “RoboArena: Distributed real-world evaluation of generalist robot policies,” in Conference on robot learning, 2025, pp. 336–364.
[206]
N. M. Amato et al., ‘Data will solve robotics and automation: True or false?’: A debate,” Science Robotics, vol. 10, no. 105, p. eaea7897, 2025.
[207]
G. R. Hunt and R. D. Gray, “The crafting of hook tools by wild new caledonian crows,” Proceedings of the Royal Society of London. Series B: Biological Sciences, vol. 271, no. suppl_3, pp. S88–S90, 2004.