Reinforcement Learning Enables Autonomous Microrobot Navigation and Intervention in Simulated Blood Capillaries

Jannik Drotleff\(^{1,\dagger}\), Samuel Tovey\(^{1,\dagger}\),
Paul Hohenberger\(^{1}\), Christoph Lohrmann\(^{1}\),
Julian Hoßbach\(^{1}\), Konstantin Nikolaou\(^{1}\), Christian Holm\(^{1,*}\)
\(^{1}\)Institute for Computational Physics, University of Stuttgart,
Allmandring 3, 70569 Stuttgart, Baden-Württemberg, Germany
\(^\dagger\)These authors contributed equally to this work.
\(^*\)Correspondence: holm@icp.uni-stuttgart.de


Abstract

Autonomous microrobots navigating biological vasculature could enable targeted drug delivery and thrombolysis, yet training control policies for realistic environments remains an open challenge. Prior reinforcement learning (RL) studies of microrobotic navigation have been limited to idealized geometries that omit complex hydrodynamic flow fields, confined branching structures, and dense cellular obstacles found in vivo. Here, we develop a physically grounded simulation of a blood capillary network, incorporating realistic hydrodynamic flow fields, explicit red blood cell dynamics, and anatomically derived branching geometry, and train deep RL agents to navigate it via chemotaxis. We systematically map the physical limits of navigation across robot size and swimming speed, revealing a forbidden regime where Brownian motion and flow overcome propulsion. Successful agents independently discover multiple universal strategy types, including run-and-rotate and energy-efficient search-and-sit policies, regardless of robot parameters. Without retraining, these agents perform targeted blocking and unblocking of capillary flow, restoring throughput to healthy baseline levels. These results establish RL as a viable framework for developing autonomous microrobotic intervention strategies in complex biological environments.

1 Introduction↩︎

The ability to precisely control robots on the micrometer scale would open a new era of medical applications, from microscale manufacturing [1] and porous media exploration [2], [3] to targeted drug delivery, minimally invasive surgery [4], and thrombolysis [5]. The clinical urgency is substantial: thrombosis-related conditions account for one in four deaths worldwide [6]. Realizing this potential demands solutions to two coupled challenges: building microrobots that can propel themselves at low Reynolds numbers and developing control strategies that function in the complex, noisy environments these robots will encounter.

At the microscale, motion is governed by a low Reynolds number (\(Re\)), where viscous forces dominate inertial effects, and propulsion becomes fundamentally different from macroscopic locomotion. Despite this constraint, hardware progress has been rapid. Externally controlled magnetic microrobots have shown promising results, both as nanoparticle swarms [7][9] and individual colloids [10]. Bio-inspired flagella provide propulsion and orientation mechanisms [11][13], while Janus particles offer asymmetric surface properties that enable control via magnetic fields and laser-driven self-thermophoresis [14], [15]. Most recently, a fully integrated robotic surface swimmer with on-board sensing, computation, memory, communication, and locomotion was realized at the 200 μm scale [16], demonstrating that previously assumed engineering limitations can be overcome. Yet even the most advanced designs offer only rudimentary actuation (typically translation and uncontrolled rotation) and remain confined to surface swimming, far from operating inside the body. Further complicating control is Brownian motion, a stochastic noise arising from thermal interactions with the surrounding fluid, which constantly perturbs the robot’s trajectory.

Reinforcement learning (RL) has become a powerful tool for discovering effective navigation strategies in these challenging environments [17]. Initial applications focused on optimizing navigation within fluidic disturbances: Colabrese et al. [18] utilized Q-learning to enable gravitactic microswimmers to exploit vortex flows, outperforming passive strategies, while Nasiri and Liebchen [19] employed the Advantage Actor-Critic algorithm to guide active particles through complex two-dimensional flow fields. Moving beyond purely synthetic settings, Muiños-Landin et al. [20] demonstrated RL-based microrobotic control under real experimental conditions. Other work has focused on emergent behaviors: Xiong et al. [21] and Tovey et al. [22] showed that RL agents trained via actor-critic architectures can autonomously rediscover chemotactic behaviors, the oriented motion toward higher chemical concentrations, recovering the same run-and-tumble strategies used by bacteria [23][25]. The mechanics underlying these learned policies have been characterized [26], revealing how neural network responses to environmental inputs produce the observed motion patterns. To facilitate such studies, Tovey et al. [27] introduced SwarmRL, an open-source platform that integrates RL with molecular dynamics simulations for active matter research.

However, a significant gap remains between these idealized simulations and the application-specific environments needed for microrobotic deployment. Current studies focus on understanding control algorithms and swimming mechanisms in open domains, simple channels, or uniform flow fields, with little attention given to the realistic environments required to model actual deployment scenarios. One such environment is that of a blood capillary network: a complex geometry of branching paths under the influence of flow fields and filled with passive blood cells. Biological capillaries present severe challenges from both simulation and learning perspectives, as robots must navigate confined geometries where flow fields are further complicated by the dense presence of dynamic obstacles such as red blood cells (RBCs) [28], [29].

Here, we bridge this gap by developing a physically grounded simulation of a blood capillary network to train and evaluate deep RL agents. Rather than aiming for a perfect biological representation, we employ a geometry that qualitatively captures essential properties—such as branching morphology and confined flow. Our environment incorporates realistic branching geometry derived from anatomical illustration, steady-state hydrodynamic flow fields, explicit RBC dynamics, and Brownian fluctuations. Within this framework, we map the physical boundaries of successful chemotactic navigation across robot size and speed, analyze the emergent control strategies, and demonstrate that trained agents can perform targeted blocking and unblocking of capillary flow without retraining—capabilities directly relevant to embolotherapy and thrombolysis.

2 Results↩︎

Capillary environment and navigation challenge↩︎

Figure 1: Simulated capillary environment. a) Capillary geometry and Lattice-Boltzmann flow field, derived from an anatomical capillary illustration [30]. b) Static concentration field for a chemical source at \vec{r}_s = (\SI{119}{\micro \meter}, \SI{175}{\micro \meter})^T, constrained by capillary walls.

To bridge the gap between idealized simulations and the complexity of biological environments, we established a realistic simulation environment modeled after human blood capillary structures. We derived the simulation boundaries from an anatomical capillary illustration from the Encyclopedia Britannica [30], mapping it into a domain of \(\SI{200}{\micro \meter} \times \SI{184.13}{\micro \meter}\) based on physiological capillary diameters of 5 μm to 10 μm [31]. A steady-state flow field was computed using the Lattice-Boltzmann method, with the vessel geometry dictating local flow velocities and creating a training environment where peak velocities reach approximately \(3.5\times\) the mean flow (1a).

Central to our approach is chemotaxis, a directed motion along a chemical concentration gradient toward a source. To represent the chemical signal, we applied a static concentration field that weakens as it spreads. We constrained this field to the inside of the capillaries, ensuring that the gradient accurately traces the anatomical branching structure (1b). To maximize the navigation challenge for the RL agents, the chemical source was placed at \(\vec{r}_s = (\SI{119}{\micro \meter}, \SI{175}{\micro \meter})^T\), in a region of comparatively high flow. This steady-state approximation is justified by a Péclet number \(Pe < 1\) throughout the domain [32], ensuring that diffusion dominates over advective transport of the chemical signal (see 5 for the full Péclet analysis).

Physical limits of navigation↩︎

Figure 2: Navigation performance across robot parameters. a) Probability of successful chemotaxis (opacity scales with success rate); success requires \geq 8 of 10 robots reaching within 20 μm of the source. b) Cumulative reward across all 30 training runs per parameter combination. c) Mean equilibrium distance from the source. d) Mean time to reach the equilibrium distance.

To map the physical boundaries of robotic navigation, we systematically varied the radius of the robot and the swimming speed, training 30 independent agents per combination of parameters with 10 robots each. A successful agent was defined as one that steered at least eight of ten robots to within 20 μm of the source during a two-hour deployment run. The resulting phase diagram (2a) reveals a distinctive forbidden regime for small, slow colloids, where Brownian motion and the opposing capillary flow overwhelm the robots’ limited propulsive force.

Compared to an equivalent phase space analysis in a simpler environment with only Brownian motion [26], this forbidden region shifts to larger and faster robots, as the complex boundary geometry, opposing flow, and RBC collisions form additional hindrances. To confirm that this boundary reflects physics rather than algorithmic limitations, we transferred a single successfully trained model across the entire parameter space (6). The persistent forbidden region under transfer confirms a fundamental physical boundary where environmental forces supersede any possible control strategy, and suggests that the RL agent finds a solution as soon as one is physically possible.

The highest success probability (80 %) was achieved by mid-sized robots (\(r = \SI{1.4}{\micro\meter}\)) at moderate speed (\(\SI{2.5}{\text{body lengths per second}}\)), indicating that maximum propulsion does not guarantee maximum success. This reveals a clear size–speed trade-off. Small robots (\(r < \SI{0.9}{\micro\meter}\)) are less affected by RBC collisions but are severely affected by Brownian fluctuations and local flow drag. In contrast, large robots are resistant to noise and flow, but, given that the RBCs were modeled with \(r_\text{RBC} = \SI{1.5}{\micro\meter}\), they cannot avoid collisions with these obstacles and lack maneuverability to navigate around them.

These physical trade-offs are also reflected in the learning dynamics. The cumulative training reward (2b), which indicates how efficiently the model learned the required policy, shows the same trade-off but with a peak for the fastest robots of the lower-mid-size. This differs from the success rate results: while faster agents accumulate higher rewards by reaching the target quickly, they appear more susceptible to catastrophic failures in crowded flow, whereas larger mid-sized agents provide the physical stability necessary for consistent deployment success.

To characterize the actual performance of the swarm, we implemented successful strategies without further training and measured both the equilibrium distance and arrival time (2c,d). In contrast to findings in idealized bulk environments, equilibrium distances show no clear trend across the parameter space, ranging from 2 μm to 11 μm, comparable to values from our previous work despite substantially increased environmental complexity. Although there is a slight tendency for smaller robots to get closer to the source, the physical constraints of capillary walls, RBC collisions, and opposing flow impose a floor on the equilibrium distance regardless of robot size, suggesting that the learned strategies remain effective at localized targeting even under realistic perturbations. The time to arrival scales inversely with size and speed, but is roughly tenfold longer than in the simple Brownian case (10–80 versus 1–10), reflecting both the reduced net velocity from flow opposition and the tortuous paths imposed by the capillary geometry.

Emergent navigation strategies↩︎

Figure 3: Strategy analysis. a) t-SNE embedding of learned policies colored by robot radius. b) Same embedding colored by k-means cluster assignment (k=11, selected by silhouette analysis [33]). c–h) Mean action probability distributions (bold lines) and standard deviation (shaded) for the six largest clusters as a function of sensed concentration change. Positive values indicate motion toward the source; amplitude reflects proximity. Remaining clusters shown in 8.

The above results demonstrate that RL can solve the challenge of navigation within complex biological environments. To understand whether different robot configurations converge on similar control logic and what those strategies look like, we analyzed the learned actor networks directly rather than relying on trajectories alone, as different strategies do not necessarily produce distinguishable trajectories.

We mapped each agent’s policy, defined as the action probability distribution across a range of chemical concentration changes, into a two-dimensional embedding using the t-distributed stochastic neighborhood embedding (t-SNE) [34] with perplexity 30 and PCA initialization, using the scikit-learn implementation [35] (3a,b). Two prominent clusters emerge on the left and right of the embedding, with several smaller clusters below. Critically, robots of all sizes appear in nearly every cluster (3a), and the same holds when coloring by swim speed (7), demonstrating that the emergent strategies are universal and independent of the robot’s physical parameters.

To quantify the cluster structure, we applied \(k\)-means clustering [36] with the silhouette metric [33] to determine the optimal number of clusters, finding \(k_\text{max} = 11\) (3b) In the mean policy plots (3c–h), the sign of the concentration change denotes the direction of motion relative to the source (positive for approaching, negative for receding), while the amplitude indicates proximity: high-amplitude changes signify nearness to the source due to the steep reciprocal gradient, and low-amplitude changes indicate greater distance.

Analysis of individual strategy distributions reveals distinct navigational phenotypes. The two dominant clusters (1 and 2) implement classic Run-and-Rotate policies, characterized by clockwise or counterclockwise rotation for negative inputs and translation for positive inputs, with a transition around zero concentration change. Cluster 3 employs Brownian Piloting, an energy-efficient strategy where agents remain idle during large negative inputs and translate only when sensing small negative or large positive inputs, effectively conserving energy by relying on Brownian fluctuations for reorientation.

A distinct class of behavior emerges in Cluster 4 and its variants (Clusters 6–8, 10, and 11), which share a Search-and-Sit phenotype specifically adapted to the capillary environment. Here, agents translate with high probability only for small concentration changes, corresponding to the weaker gradients found far from the source. As they approach the source, they transition to idling or rotating. In effect, the agents swim up through the capillary system and, upon reaching the source, sit until the opposing flow drifts them away by a certain distance, at which point they repeat the approach. This cycle emerges regardless of whether agents idle or rotate when being carried back, as these two actions are functionally equivalent in the presence of flow.

Finally, Cluster 5 (and the hybrid Cluster 9) follows a Gradient Gliding strategy, maintaining a nearly continuous translation signal and using rotation only during weak gradient signals for course correction. Although these types of strategy share heritage with policies documented in idealized environments [26], they exhibit specific adaptations to capillary flow conditions, introducing a broader taxonomy of navigation tactics for confined, flow-dominated environments. The remaining cluster strategies (7–11) are shown in 8.

Functional intervention: blocking and unblocking capillary flow↩︎

Figure 4: Targeted capillary intervention. a) Blocking task setup: chemical sources (red crosses) placed at y = \SI{90}{\micro\meter} in each vessel branch. b) Measured leak rate versus number of agents for varying robot radii; larger agents achieve full occlusion with fewer robots. c,d) Vertical unblocking setup and normalized leak rates showing near-complete flow restoration (green) relative to free flow (blue) and blocked baseline (red). e,f) Horizontal unblocking setup and corresponding leak rates.

Although multiple universal strategy types emerged during training, we selected the Run-and-Rotate phenotype from Clusters 1 and 2 for functional task demonstrations, as it was the most prominent emergent strategy. We now show how these behaviors can be leveraged for targeted medical interventions. By strategically placing chemical sources, the microrobotic swarm can be commanded to perform functional tasks, specifically blocking and unblocking microvascular flow, without any retraining. We note that it should be feasible to create the required local chemical concentration fields pharmacologically. Because the learned policy relies solely on local gradient sensing, it can be applied independently of the source location. These tasks mimic the requirements for embolotherapy, where blood flow is restricted to treat hemorrhages or starve tumors, and thrombolysis, where obstructions must be cleared.

For the blocking task, we placed a chemical source in each tube of the capillary network at a height of \(y = \SI{90}{\micro\meter}\), effectively commanding the robots to navigate there (4a). Varying both the radius of the robot and the size of the swarm, we measured the resulting rates of blood leak through the blockage (4b). Blocking efficiency scales with both robot size and number: larger agents achieve total cessation of flow with fewer agents than smaller counterparts, owing to their greater cross-sectional area. The leak rate decreases linearly with the number of agents until a critical density is reached, at which point the capillary becomes fully occluded. These results demonstrate that the learned Run-and-Rotate policy is sufficiently robust to enable the collective formation of physical barriers, and that the system’s flow resistance can be precisely tuned by modulating swarm composition. This provides a versatile framework for autonomous embolotherapy.

For the inverse challenge of unblocking (restoring flow to a pathologically obstructed network), we introduced multiple blockages strong enough to withstand the blood stream and deployed 20 robots per configuration. We measured free flow (no blockages or robots), blocked flow (\(n_\text{blockages}\) present), and restored flow after robot deployment. This was tested in two configurations: blockages added from top to bottom (4c,d) and from left to right (4e,f). In both systems, the unblocked leak rates overlap with the free-flow baseline, as seen in the overlap between the blue and green regions in (4d,f). This near-perfect flow restoration demonstrates complete recovery of capillary throughput by the microrobotic swarm, regardless of the initial number or arrangement of obstructions. The Run-and-Rotate policy proves sufficient to clear even multi-point obstructions, effectively returning the capillary network to its healthy physiological state and demonstrating the capacity of RL to deliver autonomous, physically grounded solutions for in vivo microrobotic intervention.

3 Discussion↩︎

We have demonstrated that deep RL agents can learn robust universal navigation policies in a physically grounded capillary simulation and apply them, without retraining, to perform targeted vascular blocking and unblocking. These results establish a concrete path from simulation-trained control to functional microrobotic intervention, contingent on corresponding advances in microrobot hardware.

The emergence of multiple distinct types of strategy has direct practical implications. The energy-efficient Search-and-Sit and Brownian Piloting policies are particularly relevant given that on-board power at the microscale would be rapidly depleted in a flow environment. Traditional power sources would be exhausted quickly; strategies that selectively idle and allow stochastic fluctuations or local flow patterns to assist in positioning offer a promising route to sustained operation. That even the simplest Run-and-Rotate policy suffices for both blocking and unblocking tasks sets a performance floor and provides a baseline for future studies to investigate whether more specialized policies, such as Search-and-Sit, could offer higher efficiency under specific flow conditions.

For application demonstrations, we utilized only the Run-and-Rotate policy, as it emerged as the most prominent strategy. Its success in solving medically relevant tasks suggests that even the most fundamental emergent behaviors are capable of addressing clinical problems. Notably, because each agent relies solely on local gradient sensing, collective blocking emerges as a natural group behavior without requiring complex inter-robot communication—a significant practical advantage for swarm-based interventions.

Several simplifications in our simulation constrain the current level of realism, and addressing these represents a clear path for future work.

We model blood as a two-component fluid consisting of spherical RBCs and water-like plasma. In nature, the characteristic RBC size is of order 5 μm; they are not spherical but elliptical and deformable, capable of squeezing through the narrowest capillaries [37]. To account for this, we used an effective diameter of 3 μm to phenomenologically represent their high deformability. We also used a rigid-wall approach, assuming that the shape of the vessel is not significantly affected by flow and particles [38].

The Fåhræus-Lindqvist effect describes the migration of RBCs toward vessel centers, which leaves a cell-free layer near the walls [39]. Although not explicitly modeled here, this phenomenon could actually benefit microrobot performance by allowing robots near the walls to avoid the highest concentration of RBCs and thereby reduce collisions.

A critical gap between our simulations and physiological conditions concerns flow velocity. Our simulated flows are of order  μm/s, while physiological capillary blood flow ranges from 0.47 mm/s to 0.84 mm/s [40], with brain capillaries reaching 0.5 mm/s to 1.8 mm/s [41]. Since robots must swim at least as fast as the flow to overcome it, micron-sized robots would require swimming velocities of order \(10^3\) body lengths per second. Our tests in this regime showed that the required angular velocities must increase by a similar factor, imposing a significant engineering challenge. However, since physiological flow speeds are typically measured by tracking RBCs in the vessel center, the actual flow near the walls may be substantially slower. Although most studies assume no-slip boundary conditions, recent in vivo measurements suggest wall-adjacent flow velocities of 30% to 80% of luminal velocity [42], which could relax the propulsion requirements for wall-following navigation strategies.

We also did not account for the pulsatile character of blood flow [43]. Future studies should implement non-static flow fields on a Lattice-Boltzmann foundation, though this would add significant computational cost (particularly when extending to three dimensions) and require addressing the shear-rate-dependent dynamic viscosity of blood [44].

At higher physiological flow speeds, the Péclet number increases and chemical signals will be advected downstream, producing smeared rather than static concentration fields. The local-observable design of our agents makes them inherently well-suited for navigating such dynamic gradients, as the perceived concentration change that determines each action would naturally reflect the advected field. However, it will be crucial to either model the concentration change noise realistically during in silico training or to develop precise sensors that offer accurate reception of the surrounding chemical concentration.

Importantly, the gradient-following logic learned by the RL agents is fundamentally gradient-agnostic: the same control architecture could operate on magnetic or acoustic gradients rather than chemical ones, potentially enabling external steering and control. This flexibility substantially broadens the applicability of trained policies.

Despite the modeling simplifications discussed above, our results demonstrate that, given a sufficiently fast and sensitive agent, chemotaxis inside a blood flow environment is physically viable and functionally useful. The simulation framework we have developed provides a foundation for increasingly realistic studies, while the emergent navigation strategies offer concrete design principles for future microrobotic systems. As microrobot hardware continues to advance, the combination of RL-trained control policies with physically grounded simulation environments will be essential to bridge the gap between laboratory demonstrations and clinical deployment.

4 Methods↩︎

Software↩︎

The reinforcement learning experiments were performed using the open-source Python package SwarmRL [27], with machine learning built on JAX [45] and Flax [46]. All molecular dynamics simulations used ESPResSo [47], while the fluid flow was generated in waLBerla [48].

Simulation environment↩︎

We generate a 2D boundary map by binarizing a capillary structure sketch from the Encyclopedia Britannica [30]. The resulting image was cropped to \(504\times464\) pixels, ensuring compatibility with the periodic boundary conditions that allow flow and particles to re-enter the capillary network. Based on average capillary diameters of 5 μm to 10 μm [31], we set the system width to 200 μm, which yields a height of 184.13 μm, where each pixel represents a square of approximately 0.40 μm side length.

Steady-state fluid dynamics was simulated using the Lattice-Boltzmann (LB) method in waLBerla [48] with parameters listed in 1. To adapt the 3D LB implementation to our 2D model, we extracted the central layer of the flow field and set all \(z\)-components to zero, creating a steady-state flow traversing the network from top to bottom. The velocity fields were normalized to allow arbitrary scaling of the mean flow velocities for different agent configurations. The peak velocities reach approximately \(3.5\times\) the mean flow velocity, creating a challenging dynamic environment where robots must constantly adapt to the varying drag forces.

Table 1: Parameters of the Lattice-Boltzmann simulation.
Parameter Value
Time step 1.0
Kinematic viscosity 0.2
Density 1.0
Agrid 1.0
Tau 1.0
External force density \(5 \times 10^{-6}\)
Inlet velocity 0.001
Height 20.0
Length in \(x\) 500.0
Length in \(y\) 462.738
Flow direction [0, -1, 0]
Number of steps 6960
Resulting box dimensions [504, 464, 22]

Chemical concentration field↩︎

The chemical landscape was modeled as a reciprocal decay field, \(f(\vec{r}) = 1/d_s(\vec{r})\), where \(d_s\) represents the shortest path distance from the source to a given point \(\vec{r}\), respecting wall boundaries. This approach is justified by the assumption of a steady-state equilibrium in which the chemical has had sufficient time to diffuse throughout the capillary structure. To verify this, we compute the Péclet number \(Pe = uL/D\), where \(u\) is the flow velocity, \(L\) the system dimension, and \(D\) the diffusion constant of the chemical attractant [32]. Assuming a glucose-like attractant with \(D = \SI{7.7e-10}{m^2/s}\) [49], we find \(5.181 \times 10^{-3} \leq Pe \leq 0.518\), which confirms that diffusion dominates over advection for the concentration field. To account for the complex geometry, we implemented a custom grid search algorithm that computes the shortest path around boundary obstructions, with the boundary map upscaled using scipy [50] for higher spatial resolution.

Particle dynamics↩︎

Within this framework, blood is treated as a two-component fluid: plasma is modeled as a continuum, and red blood cells (RBCs) are treated as explicit particles. Microrobot and RBC dynamics are governed by the overdamped Langevin equations: \[\begin{align} \tag{1} \dot{\vec{r}}_i &= \gamma_t^{-1}[F(t)\hat{e}_i(\theta_i)-\nabla V(\vec{r}_i, \{\vec{r}_j\})] \notag\\ &\quad + \sqrt{2 k_B T\gamma_t^{-1}}\vec{R}^t_i(t) \\ \tag{2} \dot{\theta}_i &= \gamma_r^{-1} m(t) \notag\\ &\quad + \sqrt{2k_BT\gamma^{-1}_r}R^r_i(t), \end{align}\] where \(\vec{r}_i\) and \(\theta_i\) denote the position and orientation of particle \(i\), \(\gamma_t = 6\pi \mu r\), and \(\gamma_r = 8 \pi \mu r^3\) are the translational and rotational Stokes friction coefficients, and \(\mu\) is the dynamic viscosity. We set \(\mu = \SI{0.89}{mPa \cdot s}\) near the \(\SI{1.2}{mPa \cdot s}\) value of blood plasma at \(T = \SI{37}{\degreeCelsius}\) [44]; the minor discrepancy has negligible impact on the Péclet regimes of active and diffusive motion (see 5). \(F\) and \(m\) are the active force and torque, \(\hat{e} = (\cos\theta, \sin\theta)^T\) is the orientation vector, \(V\) is the interaction potential, \(k_B\) is the Boltzmann constant and \(\vec{R}_i^{(t,r)}\) are stochastic noise terms with \(\langle R_i^{(t,r)}(t) R_j^{(t,r)}(t')\rangle = \delta_{ij}\delta(t-t')\).

Two-body interactions are modeled via the Weeks-Chandler-Andersen (WCA) potential [51]: \[\label{eq:WCA-definition} V(r_{ij}) = \begin{cases} 4V_0\biggl[\Bigl(\tfrac{\sigma}{r_{ij}}\Bigr)^{12} - \Bigl(\tfrac{\sigma}{r_{ij}}\Bigr)^{6}\biggr] + V_0 \\ \hfill \text{if } r_{ij} < r_c, \\[1ex] 0 \hfill \text{else.} \end{cases}\tag{3}\] with \(V_0 = k_B T\), \(\sigma = (a_i + a_j) \cdot 2^{-1/6}\), \(r_c = 2^{1/6}\sigma\), and \(r_{ij} = ||\vec{r}_i - \vec{r}_j ||_2\). For the unblocking task, attractive bonds between blockage particles were implemented using a Lennard-Jones potential with \(\epsilon = \SI{76.94}{J}\), \(\sigma_\text{block} = 2 r_\text{block} \cdot 2^{-1/6}\), cutoff \(r_\text{cutoff} = 2.5\,\sigma_\text{block}\), and \(r_\text{block} = \SI{1.2}{\micro \meter}\), empirically tuned to withstand flow while remaining manipulable by agents. All simulations used a time step of \(\tau = \SI{0.002}{s}\). The capillary geometry and steady-state LB flow field were integrated into the MD environment using ESPResSo’s constraint features, with local flow velocity scaled to 10 % of the agent’s swimming velocity.

Reinforcement learning↩︎

We employ an actor-critic architecture [52], [53] with Proximal Policy Optimization (PPO) [54], where the policy \(\pi\) is represented by a neural network \(\pi_\theta\). The actor outputs an action probability distribution, while the critic estimates the value function \(V(s_t)\). The advantage is estimated using Generalized Advantage Estimation (GAE) [55], and the critic is optimized by minimizing the Huber loss [56] between predicted values and target returns. An entropy term encourages early exploration, yielding the total objective: \[\mathcal{L}_\text{total}= \mathcal{L}_\text{actor} - 0.5 \cdot \mathcal{L}_\text{critic} + \rho \cdot S,\] where \(S\) is the Shannon entropy and \(\rho = 0.02\). The update is maximized via gradient ascent with \(\theta' = \theta + \eta \cdot \nabla_\theta \mathcal{L}\).

We extend this to the multi-agent reinforcement learning (MARL) setting [57], where all agents share a single actor and critic network. The experience collected by all agents is pooled for updates, with the collective objective: \[J_\text{MARL} = \frac{1}{N}\sum_{i=1}^{N_\text{agents}} J_i(\pi_\theta).\] The system follows a decentralized Markov decision process: each agent receives only local observations and individual rewards, sharing knowledge only through pooled training updates [58].

The actor and critic networks each consist of two hidden layers with 128 nodes and ReLU activation, optimized with Adam [59] at the learning rate \(\eta = 0.002\). The discrete action space is \[\mathcal{A}= \left\{ \begin{align} \text{Translate:}\quad & v = n \cdot d,\;\omega = 0, \\ \text{Rotate CCW:}\quad & v = 0,\;\omega = 10.472, \\ \text{Rotate CW:}\quad & v = 0,\;\omega = -10.472, \\ \text{Idle:}\quad & v = 0,\;\omega = 0. \end{align} \right.\] Here, \(v\) is measured in μm/s and \(\omega\) in rad/s; \(n\) sets the swimming speed relative to the diameter of the colloid \(d\), motivated by the empirical scaling of the swimming speed with body size [60]. The rotation speed \(\omega = \SI{600}{°/s} = \SI{10.472}{rad/s}\) was inspired by Escherichia coli [23].

The agent observable is the temporal change in perceived chemical concentration: \[\begin{align} \label{eq:observable} o_i(t) = w \cdot \bigl[f\bigl(\vec{r}_i(t)\bigr) - f\bigl(\vec{r}_i(t-\Delta t)\bigr)\bigr], \end{align}\tag{4}\] with weight factor \(w = 100\), concentration field \(f\), and action interval \(\Delta t = \SI{0.1}{s}\). The reward is the clipped observable: \(r_i(t) = \max(0, o_i(t))\), ensuring that robots are rewarded for approaching the source without explicit punishment for moving away.

For each size–speed combination, 30 independent training runs of 10,000 episodes were conducted with 10 robots each. Each episode spans 4 s of simulated time (40 policy applications at \(\Delta t = \SI{0.1}{s}\)), with the system reset after 1,000 uninterrupted episodes. The robots were randomly initialized in the upper half of the capillary structure. To prevent agents from becoming “blind” after being carried far from the source (where concentration changes approach zero), the system was reset when two agents drifted below \(y_\text{threshold} = -0.4 \times \SI{184}{\micro \meter}\). A custom checkpointer exported the model whenever the current reward exceeded the previous best, protecting against catastrophic forgetting in a noisy training environment. The best-performing checkpoints were selected and evaluated in two-hour deployment simulations without further training.

For functional applications, we implemented a universal Run-and-Rotate policy. It was trained with \(r = \SI{1.8}{\micro\meter}\) and \(\SI{2.5}{\text{body lengths per second}}\). In multi-source scenarios, concentration fields were computed as linear superpositions of individual fields. During unblocking, target sources were deactivated after clearance and agents were turned off once all sources were cleared; final leak rates were measured only after agent deactivation to ensure they reflected the restored physiological state.

Computational resources↩︎

Training and deployment were performed on University of Stuttgart compute resources, with each simulation utilizing one core of an AMD EPYC 9374F CPU node. No GPU was required due to the relatively small network and system sizes. Training required approximately six hours per model, and deployment simulations were completed in 90 minutes.

Data Availability↩︎

The data generated and analyzed in this study will be made publicly available through the DaRUS repository upon publication.

Code Availability↩︎

The analysis scripts used in this study will be made publicly available through the DaRUS repository upon publication. The SwarmRL framework is available as open-source software [27].

Acknowledgments↩︎

This study was funded by the Deutsche Forschungsgemeinschaft (DFG, German Research Foundation) through Compute Cluster grant no..

Author Contributions↩︎

S.T.conceived the project. C.H.supervised the project. S.T., J.D., and P.H.developed the training and deployment scripts. P.H.developed the concentration field algorithm. C.L.designed the capillary structure boundaries and performed the Lattice-Boltzmann simulations. J.D.performed all RL agent training and deployment simulations, developed the application simulation scripts, and analyzed the data. J.D.and S.T.wrote the manuscript and prepared the figures. S.T., K.N., and J.H.provided project guidance and contributed to editing. All authors reviewed the final manuscript.

Competing Interests↩︎

The authors declare no competing interests.

5 Péclet regimes↩︎

Figure 5: Péclet analysis of the navigation phase space. Translational and rotational Péclet boundaries (blue: water viscosity, orange: blood plasma viscosity) overlaid on the training success data. The forbidden regime broadly aligns with the diffusion-dominated Péclet regime, though the capillary flow field shifts the boundary relative to the pure Brownian case. The forbidden regime is reminiscent of the Péclet regime found in previous work [26], where the Péclet number segments motion as either active-driven or diffusion-dominated. However, the capillary flow field prevents a pure diffusion description: the translational Péclet boundary does not precisely match the training success characteristics. The small difference between the Péclet lines computed with water viscosity (\eta_\text{water} = \SI{0.89}{mPa \cdot s}) and blood plasma viscosity (\eta_\text{blood plasma} = \SI{1.2}{mPa \cdot s}) confirms that this choice has negligible impact, producing at most a minor shift in the forbidden region boundary.

6 Model transfer↩︎

Figure 6: Cross-parameter model transfer. A single successfully trained policy was deployed across all robot sizes and speeds (30 simulations per combination). Color indicates the probability of successful chemotaxis (\geq​8 of 10 agents within 20 μm of the source). The persistent forbidden region confirms a fundamental physical boundary rather than an algorithmic limitation.

7 Strategy analysis↩︎

Figure 7: Remaining strategy clusters. Mean action probability distributions (bold lines) and standard deviation (shaded) for clusters 7–11, complementing the six largest clusters shown in the main text.

8 t-SNE analysis for the swim multitudes↩︎

Figure 8: t-SNE embedding colored by swimming speed. The color indicates swim speed in body lengths per second. Robots of all speeds appear across clusters, confirming that emergent strategies are velocity-independent.

References↩︎

[1]
An, Y., He, B., Ma, Z., Guo, Y.&Yang, G.-Z.Microassembly: A review on fundamentals, applications and recent developments. Engineering48, 323–346(2025).
[2]
Liao, C.-T.et al.Propulsion of a three-sphere microrobot in a porous medium. Physical Review E109, 065106(2024).
[3]
Lohrmann, C.&Holm, C.Optimal motility strategies for self-propelled agents to explore porous media. Physical Review E108, 054401(2023).
[4]
Lee, J. G., Raj, R. R., Day, N. B.&Shields, C. W. I.Microrobots for biomedicine: Unsolved challenges and opportunities for translation. ACS Nano17, 14196–14204(2023).
[5]
Lv, S., Tang, L. V.&Hu, Y.Application of nanotechnology and micro/nanorobots in thrombotic diseases. EngMedicine2, 100061(2025).
[6]
Raskob, G.et al.Thrombosis: A major contributor to the global disease burden. Journal of Thrombosis and Haemostasis12, 1580–1590(2014).
[7]
Schuerle, S.et al.Synthetic and living micropropellers for convection-enhanced nanoparticle transport. Science Advances5, eaav4803(2019).
[8]
Yu, J., Yang, L.&Zhang, L.Pattern generation and motion control of a vortex-like paramagnetic nanoparticle swarm. The International Journal of Robotics Research37, 912–930(2018).
[9]
Jiang, J., Yang, L.&Zhang, L.DQN-based on-line path planning method for automatic navigation of miniature robots. In 2023 IEEE International Conference on Robotics and Automation (ICRA), 5407–5413(2023).
[10]
Landers, F. C.et al.Clinically ready magnetic microrobots for targeted therapies. Science390, 710–715(2025).
[11]
Dreyfus, R.et al.Microscopic artificial swimmers. Nature437, 862–865(2005).
[12]
Yamazaki, A.et al.Wireless micro swimming machine with magnetic thin film. Journal of Magnetism and Magnetic Materials272–276, E1741–E1742(2004).
[13]
Zhang, L.et al.Characterizing the swimming properties of artificial bacterial flagella. Nano Letters9, 3663–3667(2009).
[14]
Su, H., Li, S., Yang, G.-Z.&Qian, K.Janus micro/nanorobots in biomedical applications. Advanced Healthcare Materials12, 2202391(2023).
[15]
Jiang, H.-R., Yoshinaga, N.&Sano, M.Active motion of a Janus particle by self-thermophoresis in a defocused laser beam. Physical Review Letters105, 268302(2010).
[16]
Lassiter, M. M.et al.Microscopic robots that sense, think, act, and compute. Science Robotics10, eadu8009(2025).
[17]
Cai, W., Wang, G., Zhang, Y., Qu, X.&Huang, Z.Reinforcement learning for active matter. Biophysics Reviews6, 031302(2025).
[18]
Colabrese, S., Gustavsson, K., Celani, A.&Biferale, L.Flow navigation by smart microswimmers via reinforcement learning. Physical Review Letters118, 158004(2017).
[19]
Nasiri, M.&Liebchen, B.Reinforcement learning of optimal active particle navigation. New Journal of Physics24, 073042(2022).
[20]
Muiños-Landin, S., Fischer, A., Holubec, V.&Cichos, F.Reinforcement learning with artificial microswimmers. Science Robotics6, eabd9285(2021).
[21]
Xiong, T., Liu, Z., Wang, Y., Ong, C. J.&Zhu, L.Chemotactic navigation in robotic swimmers via reset-free hierarchical reinforcement learning. Nature Communications16, 5441(2025).
[22]
Tovey, S.et al.Environmental effects on emergent strategy in micro-scale multi-agent reinforcement learning(2023). .
[23]
Berg, H. C.&Brown, D. A.Chemotaxis in escherichia coli analysed by three-dimensional tracking. Nature239, 500–504(1972).
[24]
Watari, N.&Larson, R. G.The hydrodynamics of a run-and-tumble bacterium propelled by polymorphic helical flagella. Biophysical journal98, 12–17(2010).
[25]
Darnton, N. C., Turner, L., Rojevsky, S.&Berg, H. C.On torque and tumbling in swimming escherichia coli. Journal of Bacteriology189, 1756–1764(2007).
[26]
Tovey, S., Lohrmann, C.&Holm, C.Emergence of chemotactic strategies with multi-agent reinforcement learning. Machine Learning: Science and Technology5, 035054(2024).
[27]
Tovey, S.et al.SwarmRL: Building the future of smart active systems. The European Physical Journal E48, 16(2025).
[28]
Freund, J. B.Numerical simulation of flowing blood cells. Annual Review of Fluid Mechanics46, 67–95(2014).
[29]
Marsden, A. L.&Esmaily-Moghadam, M.Multiscale modeling of cardiovascular flows for clinical decision support. Applied Mechanics Reviews67(2015).
[30]
Britannica, E.Blood vessel definition, anatomy, function, & types britannica. https://www.britannica.com/science/blood-vessel(2026).
[31]
Betts, J. G.et al.Ch. 1 introduction - anatomy and physiology . https://assets.openstax.org/oscms-prodcms/media/documents/anatomy-and-physiology-2e_-_WEB.pdf(2013).
[32]
Bird, R., Stewart, W.&Lightfoot, E.Transport Phenomena(John Wiley and Sons, New York, 2002), 2 edn.
[33]
Rousseeuw, P. J.Silhouettes: A graphical aid to the interpretation and validation of cluster analysis. Journal of Computational and Applied Mathematics20, 53–65(1987).
[34]
van der Maaten, L.&Hinton, G.Visualizing Data using t-SNE. Journal of Machine Learning Research9, 2579–2605(2008).
[35]
Pedregosa, F.et al.Scikit-learn: Machine learning in python. Journal of Machine Learning Research12, 2825–2830(2011).
[36]
Lloyd, S.Least squares quantization in PCM. IEEE Transactions on Information Theory28, 129–137(1982).
[37]
Phillips, R., Kondev, J., Theriot, J.&Garcia, H.What and where: Construction plans for cells and organisms. In Physical Biology of the Cell, 2, 68(Garland Science, New York, 2012), 2 edn.
[38]
Kannojiya, V., Das, A. K.&Das, P. K.Simulation of blood as fluid: A review from rheological aspects. IEEE Reviews in Biomedical Engineering14, 327–341(2021).
[39]
Ascolese, M., Farina, A.&Fasano, A.The Fåhræus-Lindqvist effect in small blood vessels: How does it help the heart?Journal of Biological Physics45, 379–394(2019).
[40]
Bollinger, A., Butti, P., Barras, J. P., Trachsler, H.&Siegenthaler, W.Red blood cell velocity in nailfold capillaries of man measured by a television microscopy technique. Microvascular Research7, 61–72(1974).
[41]
Hudetz, A. G.Blood flow in the cerebral capillary network: A review emphasizing observations with intravital microscopy. Microcirculation4, 233–252(1997).
[42]
Jarolímová, A.et al.In vivo evidence of blood flow slippage: Failure of the no-slip boundary condition assumption(2025). .
[43]
Koutsiaris, A. G.&Pogiatzi, A.Velocity pulse measurements in the mesenteric arterioles of rabbits. Physiological Measurement25, 15(2003).
[44]
Nader, E.et al.Blood rheology: Key parameters, impact on blood flow, role in sickle cell disease and effects of exercise. Frontiers in Physiology10(2019).
[45]
Bradbury, J.et al.JAX: composable transformations of Python+NumPy programs(2018). http://github.com/google/jax.
[46]
Heek, J.et al.Flax: A neural network library and ecosystem for JAX(2023). http://github.com/google/flax.
[47]
Weik, F.et al.ESPResSo 4.0 – an extensible software package for simulating soft matter systems. The European Physical Journal Special Topics227, 1789–1816(2019).
[48]
Bauer, M.et al.waLBerla: A block-structured high-performance framework for multiphysics simulations. Computers & Mathematics with Applications81, 478–501(2021).
[49]
Liu, M., Nicholson, J. K., Parkinson, J. A.&Lindon, J. C.Measurement of biomolecular diffusion coefficients in blood plasma using two-dimensional1 H-1 H diffusion-edited total-correlation NMR spectroscopy. Analytical Chemistry69, 1504–1509(1997).
[50]
Virtanen, P.et al.SciPy 1.0: Fundamental algorithms for scientific computing in Python. Nature Methods17, 261–272(2020).
[51]
Weeks, J. D., Chandler, D.&Andersen, H. C.Role of repulsive forces in determining the equilibrium structure of simple liquids. The Journal of Chemical Physics54, 5237–5247(1971).
[52]
Barto, A. G., Sutton, R. S.&Anderson, C. W.Neuronlike adaptive elements that can solve difficult learning control problems. IEEE Transactions on Systems, Man, and CyberneticsSMC-13, 834–846(1983).
[53]
Grondman, I., Busoniu, L., Lopes, G. A. D.&Babuska, R.A survey of actor-critic reinforcement learning: Standard and natural policy gradients. IEEE Transactions on Systems, Man, and Cybernetics, Part C (Applications and Reviews)42, 1291–1307(2012).
[54]
Schulman, J., Wolski, F., Dhariwal, P., Radford, A.&Klimov, O.Proximal policy optimization algorithms(2017). .
[55]
Schulman, J., Moritz, P., Levine, S., Jordan, M. I.&Abbeel, P.High-dimensional continuous control using generalized advantage estimation(2018). .
[56]
Huber, P. J.Robust estimation of a location parameter. The Annals of Mathematical Statistics35, 73–101(1964).
[57]
Gronauer, S.&Diepold, K.Multi-agent deep reinforcement learning: A survey. Artificial Intelligence Review55, 895–943(2022).
[58]
Oliehoek, F. A.&Amato, C.A Concise Introduction to Decentralized POMDPs. in Intelligent Systems(Springer International Publishing, Cham, 2016).
[59]
Kingma, D. P.&Ba, J.Adam: A Method for Stochastic Optimization. arXiv(2017). .
[60]
Murray, A. G.&Jackson, G. A.Viral dynamics: A model of the effects of size, shape, motion and abundance of single-celled planktonic organisms and other particles. Marine Ecology Progress Series89, 103–116(1992).