Pith. sign in

REVIEW 3 major objections 4 minor 87 references

Kinetic theory of decentralized learning for smart active matter

T0 review · 3 major / 4 minor · reviewed 2026-08-10 · deepseek-v4-flash

Pith's one-line read The paper derives two closed equations for the mean and variance of policies in a swarm learning by local exchange, and shows they imply a universal speed-accuracy tradeoff in two microscopic models.

desk verdict A timely, explicit kinetic theory for decentralized learning in smart active matter, but the P*-expansion makes the predicted optimum and the uncertainty relation depend on an externally supplied guess. read the letter →

arxiv 2501.03948 v2 pith:6AH3UUQM submitted 2025-01-07 cond-mat.stat-mech cond-mat.soft

classification cond-mat.stat-mechcond-mat.soft
keywords decentralizedlearningkinetictheorysmartactivematterpolicydynamicsevolutionaryhydrodynamicequationsuncertaintyrelationswarmrobotics
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper tries to establish that decentralized learning in a population of agents—smart active matter whose members exchange behavioral policies with neighbors—can be reduced to two closed differential equations, one for the average policy and one for the policy diversity. The value of such a reduction is that a whole swarm's learning dynamics, including its convergence speed and the residual spread of policies at steady state, becomes a problem in statistical physics rather than a case-by-case simulation study. The paper demonstrates the reduction on two different microscopic models—phototactic microswimmers whose policy is a rotational diffusion coefficient, and light-sensing robots whose policy is a speed sensitivity—and finds that the kinetic predictions match agent-based simulations. A direct consequence is an uncertainty relation $\tau_L \sigma^2_\infty = \lambda_0^{-1}$ between the learning time and the steady-state policy variance, so that faster learning via stronger mutations always costs larger fluctuations around the target behavior.

What carries the argument

The load-bearing object is the teaching collision integral $I_{\mathrm{teach}}[f]$, an analogue of the Boltzmann collision term in which an agent adopts both the policy and the memory of a neighbor with a reward-dependent teaching probability $p_T$. The closure is achieved by assuming Gaussian dependence of the phase-space density on $M$ and $P$, by exploiting the time-scale hierarchy so that the memory relaxes to a policy-dependent average, and by expanding the average memory $\mu_M(P)$ around a predefined policy $P^*$ (Eq. 7). These steps convert the collision operator into the closed equations (10) and (11), whose coefficients are fixed by physical response functions such as $V'_* = \partial \bar{V}_x/\partial D_\theta$ in the microswimmer model.

What would settle it

In the phototactic microswimmer model, run agent-based simulations with several reference policies $P^*$ placed increasingly far from the true optimal $D_\theta^*$, keeping all other parameters fixed, and measure the long-time mean policy: the theory predicts the converged policy shifts linearly with the error in $P^*$, so its absence, or convergence to the true optimum despite a bad $P^*$, would falsify the closure. A second check is to measure the learning time $\tau_L$ and steady-state variance $\sigma^2_\infty$ over a range of mutation rates $D_{\mathrm{mut}}$, since the product $\tau_L \sigma^2_\infty$ should remain equal to $\lambda_0^{-1}$.

Watch

Extended reading notes

Core claim

The central discovery is that teaching events between agents can be treated like binary collisions in a kinetic theory, leading to a bilinear collision operator $I_{\mathrm{teach}}[f]$ for the single-agent phase-space density. Assuming a hierarchy of time scales $\tau_G \ll \tau_M \ll \tau_T$ and Gaussian distributions in memory and policy, the hierarchy closes into Eqs. (10) and (11) for the mean policy $\mu(t)$ and diversity $\sigma^2(t)$. Under this closure the adaptation rate is proportional to density times diversity, mutations are needed to sustain diversity and adaptation, and the model-specific solutions imply the uncertainty relation $\tau_L \sigma^2_\infty = \lambda_0^{-1}$. The same machinery yields space-dependent hydrodynamic equations, and in both the microswimmer and light-sensing robot models the theory quantitatively reproduces agent-based simulations of the mean policy, the diversity, and the response to spatially varying targets.

Load-bearing premise

The whole derivation leans on expanding the policy-dependent average memory around a preselected reference policy $P^*$, and the predicted optimal policy is exact only when $P^*$ coincides with the true optimum; a poorly chosen $P^*$ shifts the convergence target itself.

Editorial extensions

If this is right

  • The adaptation speed of a swarm is set by the product of population density and policy diversity, so a population with zero policy diversity cannot learn, mirroring Fisher's fundamental theorem of natural selection.
  • Mutations play the role of driving in athermal systems: they keep the policy diversity from decaying to zero and thereby maintain the swarm's ability to adapt, at the price of steady-state fluctuations.
  • The uncertainty relation $\tau_L \sigma^2_\infty = \lambda_0^{-1}$ implies a strict speed-accuracy tradeoff: increasing the mutation rate shortens the learning time but enlarges fluctuations around the target, including fluctuations of the collective velocity.
  • For spatially varying targets the theory predicts a phase lag and amplitude reduction in the mean policy that grow as the ratio of advection time to learning time decreases, so slowly learning swarms cannot track their environment in space.
  • If the reward landscape is asymmetric, as in the light-sensing robot model, mutations shift the long-time mean policy away from the exact optimum to the flatter side of the reward peak, and this shift can be predicted from the fourth-order expansion of the average memory.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A natural extension, not developed in the paper, is to treat the reward function itself as a tunable control field in these equations, which would turn the inverse problem of designing rewards for targeted collective behavior into a systematic optimization task.
  • The same kinetic structure should carry over to discrete policy spaces, where the mutation term becomes a jump process; the paper sketches this formulation, and one can test it against a binary-choice version of the robot model.
  • The fitting procedure the paper uses on simulation data—extracting $\lambda_0$ and $D_{\mathrm{mut}}$ from the decay of $\sigma^2(t)$—suggests an experimental protocol for physical microrobot swarms, where those rates are usually not known in advance.
  • Because the uncertainty relation depends only on $\lambda_0$, an implicit design principle follows: to make a swarm learn faster without losing precision, one should increase the information-exchange rate $\lambda_0$ rather than the mutation rate.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. The paper proposes a kinetic-theory framework for decentralized learning in smart active matter, in which agents exchange policies through stochastic teaching events and mutations. Starting from a single-agent phase-space density, the authors derive closed equations for the average policy and its diversity under a time-scale hierarchy, a Gaussian closure in memory and policy, a first-order expansion of the teaching probability, and an expansion of the policy-dependent average memory around a reference policy P*. The framework is applied to two models: phototactic microswimmers with scalar policy Dθ and light-sensing robots with policy χ. The authors compare the resulting kinetic-theory predictions with agent-based simulations and derive an uncertainty relation τLσ∞² = λ0^{-1} between learning time and steady-state policy variance.

Significance. If the derivation is sound, this is a useful step toward a statistical-physics description of decentralized learning: it provides an explicit microscopic-to-macroscopic route, identifies control parameters (teaching rate, mutation strength), and gives a multi-dimensional generalization. Strengths include the unusually explicit supplemental derivations, the reproducible simulation parameters, and the application to two qualitatively different models. The main reservation is that a load-bearing expansion center P* is left unspecified and the predicted optimum is only a Newton step from that center; this directly affects the central convergence claim and the uncertainty relation. The manuscript is therefore promising but needs substantial revision to justify or remove this dependence.

major comments (3)
  1. [Main text Eq. (7); SM Eqs. (S22), (S28)-(S29)] The closure of Eqs. (10)-(11) is obtained by expanding µM(P) around P* to low order, and for Model 1 the resulting fixed point is Dθ,T = Dθ* + (VT - \bar V_x(Dθ*))/V'_* (SM Eq. S28). Because \bar V_x(Dθ) is nonlinear (SM Eq. S17), Dθ,T coincides with the true reward-maximizing policy only if Dθ* already solves \bar V_x(Dθ*) = VT; otherwise it is a single Newton step and misses the optimum by O((Dθ* - θ_opt)^2). The teaching rate λ0 = 4λT αT ρ V'_*^2, and hence the uncertainty relation τL σ∞² = λ0^{-1}, are also evaluated at Dθ*. The SM concedes after Eq. (S29) that "it is important to choose an appropriate Dθ*" but gives no selection rule or error bound, and the numerical parameter list in SM Sec. I E does not state the Dθ* or χ* used in Figs. 2-4. Since the agent-based model contains no P*, the theory is not closed from first principles: the predicted convergence target and learning rate depend on an external guess, so the central claim that µ(t) converges to the optimal policy is only as good as that guess.
  2. [Fig. 2, Eqs. (12)-(13)] The comparison in Fig. 2 uses Eq. (12) both as the theoretical curve and as a fitting function to extract λ0 and Dmut from agent-based data. Because λ0 is defined through V'_* evaluated at the unspecified Dθ*, the reported values Dfit_mut = 0.094 ± 0.002 and λfit_0 = (2.65 ± 0.01)×10^{-4}, and the claim that they "compare well," cannot be independently checked. The figure also shows no error bars on the simulation data, so the agreement for σ²(t), particularly for Dmut = 0.1, is only visual. A quantitative fitting protocol, including the fit range, the value of Dθ* used for the theory curves, and statistical uncertainties, should be supplied.
  3. [End Matter Eq. (22); SM Sec. I B; Fig. 4] The derivation relies on several uncontrolled approximations that directly affect the quantitative claims: the teaching probability is expanded to first order in αT(R - R') while simulations use αT = 10; the policy distribution is assumed Gaussian even though Dθ is physically non-negative; and the memory variance is assumed policy-independent and drops out of the policy dynamics. None of these approximations is tested against simulations or accompanied by an estimate of its error. In Fig. 4(a) the kinetic theory visibly converges faster than the agent-based model, which the authors attribute to the fourth-order expansion of \bar R(χ); the discrepancy is not quantified, and it is not clear whether the uncontrolled truncation is responsible. A sensitivity analysis or a test of the closures would substantially strengthen the claimed quantitative agreement.
minor comments (4)
  1. [Conclusion] The word "iterativelly" should be spelled "iteratively."
  2. [Main text, notation] The notation Dθ,T for the target policy and VT for the target velocity is easy to confuse; a short table of symbols would improve readability.
  3. [End Matter, Eq. (17)] The transition rate Wf is described as f-dependent, but Eq. (17) uses the marginal f0(M,P); the distinction between the full density f and its marginal f0 should be stated more explicitly at that point.
  4. [SM Eq. (S36)] The piecewise expression for \bar I(χ) would be easier to follow if the authors stated which branch is relevant for the parameters used in Fig. 4 and how χ* was selected in that branch.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the closed policy equations are derived from the microscopic agent-based rules, and the P-star expansion is an explicitly acknowledged approximation rather than a hidden input-output identification.

full rationale

The derivation chain starts from the stated microscopic rules (teaching rate lambda_T, teaching probability p_T, mutation diffusion Dmut, memory dynamics) and constructs the phase-space kinetic equation, Eq. (3), and the teaching collision term, Eqs. (16)-(20). The closed equations for mu and sigma^2, Eqs. (10)-(11), are obtained by moment expansion and a Gaussian closure, with the model-specific functions F1 and F2 fixed by explicit SM calculations (Eqs. S26-S27 for Model 1 and S46-S47 for Model 2). The apparent P-star dependence is an approximation, not a circularity: the SM explicitly states at Eq. (S28) that Dtheta,T = Dtheta* + (VT - Vbar_x(Dtheta*))/V'_* is an approximation of the target policy which maximizes the reward, exact only when Vbar_x(Dtheta*) equals VT, and warns that it is important to choose an appropriate Dtheta*. This is a stated accuracy limitation of the local expansion, not a quantity smuggled in as a prediction. The reported fit of Eq. (12) to agent-based data (Fig. 2) is a parameter-extraction consistency check: the extracted Dmut and lambda_0 are compared with the known simulation inputs (Dmut = 0.1 and lambda_0 = 2.655e-4), and the KT curves themselves are generated from these inputs, so no fitted parameter is relabeled as a prediction. The uncertainty relation tau_L sigma^2_infinity = lambda_0^{-1} follows algebraically from the closed-form solution of the derived ODEs, not from an assumed ansatz. Self-citations to earlier kinetic theory of active matter (Refs. 69, 70, 72, 73) supply standard Fourier-mode techniques with stated assumptions; they are not invoked as an external uniqueness theorem and are not load-bearing in the new derivation. A reproducibility gap remains: the SM numerical-parameters list omits the Dtheta* value used in Model 1, but this is missing support rather than circularity.

Assumptions & free parameters 1 free parameters · 5 assumptions · 0 invented entities

The central claim rests on four approximations: molecular chaos, time scale separation, Gaussian closure, and expansion around P*. These are standard in kinetic theory but are not validated beyond the two examples. No new physical entities are introduced.

free parameters (1)
  • Expansion center P* (Dθ* for model 1, χ* for model 2) = Not specified; chosen by user
    Eq. (7) expands µM(P) around P*; for model 1 the predicted optimal policy Dθ,T = Dθ* + (V_T - \bar Vx(Dθ*))/V'_* is exact only if Dθ* equals the true optimum. A poor choice biases the predicted convergence target.
assumptions (5)
  • domain assumption Two-body phase-space density factorizes (molecular chaos approximation).
    End Matter text: 'we have assumed that the two-body phase-space density can be approximated by multiplying the one-body phase-space densities'.
  • domain assumption Time scale hierarchy τG << τM << τT and local communication (λT δc << v0).
    Main text section 'Kinetic theory'; used to justify memory steady state and local teaching.
  • domain assumption Gaussian closure for f0 in M and P.
    End Matter: 'a Gaussian dependence of f0 on M can be justified with the central limit theorem'; main text assumes Gaussian in P.
  • ad hoc to paper Expansion of µM(P) around P* kept to low order.
    Main text Eq. (7); the choice of P* is not derived, and errors enter the closed equations.
  • domain assumption First-order Taylor expansion of tanh activation (O(α_T^3) neglected).
    End Matter: 'we approximate the activation function to first order, tanh(αT(R-R')) = αT(R-R') + O(α_T^3)'.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Kinetic theory of decentralized learning for smart active matter." pith.science (2026). https://pith.science/paper/6AH3UUQM

@misc{pith2026250103948,
  author       = {Pith},
  title        = {Pith review of: Kinetic theory of decentralized learning for smart active matter},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/6AH3UUQM}},
  note         = {Machine review of arXiv:2501.03948}
}
read the original abstract

Smart active matter has the ability to control its motion guided by individual policies to achieve collective goals. We introduce a theoretical framework to study a decentralized learning process in which agents can locally exchange policies to adapt their behavior and maximize a predefined reward function. We use our formalism to derive explicit hydrodynamic equations for the policy dynamics. We apply the theory to two different microscopic models where policies correspond either to fixed parameters similar to evolutionary dynamics, or to state-dependent controllers known from the field of robotics. We find good agreement between theoretical predictions and agent-based simulations. By deriving fundamental control parameters and uncertainty relations, our work lays the foundations for a statistical physics analysis of decentralized learning.

Figures

Figures reproduced from arXiv: 2501.03948 by the authors.

Figure 1
Figure 1. FIG. 1. Sketch of the decentralized learning procedure de [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. FIG. 2. Adaptation of the microswimmer model to a constant [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 3
Figure 3. FIG. 3. Adaptation of the microswimmer model to a space [PITH_FULL_IMAGE:figures/full_fig_p004_3.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

87 extracted references · 68 canonical work pages

  1. [1]

    Drossel, Biological evolution and statistical physics , Advances in physics 50, 209 (2001)

    B. Drossel, Biological evolution and statistical physics , Advances in physics 50, 209 (2001)

  2. [2]

    To enhance adapt- ability, we also include a small rate of spontaneous policy changes

    only reward differences matter and agents are lacking a con- cept of an absolute reward, which highlights the decen- tralized nature of the optimization. To enhance adapt- ability, we also include a small rate of spontaneous policy changes. By analogy with evolutionary processes in biol- ogy, we call these spontaneous changes ‘mutations’ [ 75] and model th...

  3. [3]

    The vertical arrow indicates τM = 1

    and with mutations ( Dmut = 0 .1). The vertical arrow indicates τM = 1. (a) Average policy (rotation diffusion) µ(t). (b) Diversity σ2(t) (variance of policy). The black dashed- dotted curve is Eq. (

  4. [4]

    Chia and N

    N. Chia and N. Goldenfeld, Statistical mechanics of hor- izontal gene transfer in evolutionary ecology, Journal of Statistical Physics 142, 1287 (2011)

  5. [5]

    ( 12) can be used to fit the long-time behavior of σ2(t) on agent- based simulations data (see Fig

    Additionally, Eq. ( 12) can be used to fit the long-time behavior of σ2(t) on agent- based simulations data (see Fig. 2) to determine λ0 and Dmut, as on experimental data where such parameters are unknown. We extract Dfit mut = 0 .094 ± 0.002 and λfit 0 = (2.65±0.01)·10−4, which compare well to the agent-based model values Dmut = 0.1 and λ0 = 2.655 · 10−4. W...

  6. [6]

    9 1 1 . 1 1 . 2 1 . 3 (b) ¯R(χ ) χ KT AS FIG. 4. Adaptation of the light-sensing robot model to max- imize collected light: Kinetic theory (KT) and agent-based simulations (AS) without mutations ( Dmut = 0) and with mutations (Dmut = 10 −3). (a) Average policy µ(t) = ⟨χ⟩; (b) Average reward ¯R(χ) for robots with a given policy χ. reward. We find for the av...

  7. [7]

    R. A. Watson, S. G. Ficici, and J. B. Pollack, Embodied evolution: Distributing an evolutionary algorithm in a population of robots, Robotics and Autonomous Systems 39, 1 (2002)

  8. [8]

    Sella and A

    G. Sella and A. E. Hirsh, The application of statistical physics to evolutionary biology, Proceedings of the Na- tional Academy of Sciences 102, 9541 (2005)

Show all 87 references
  1. [9]

    Houchmandzadeh and M

    B. Houchmandzadeh and M. Vallade, The fixation prob- ability of a beneficial mutation in a geographically struc- tured population, New Journal of Physics 13, 073020 (2011)

  2. [10]

    and ( 11) (detailed derivations in [ 74]) read, σ2(t) = √ 2Dmut λ0 1 tanh(√2Dmutλ0(τ0 + t)) (12) µ(t) = Dθ,T + η0 √2Dmutλ0 sinh(√2Dmutλ0(τ0 + t)) , (13) 4 0.79 0.82 0.85 0.88 0.91 − 0.5 − 0.25 0 0 .25 0 .5 ⟨Vx(x)⟩ x/L VT (x) KT Dmut=0.05 KT Dmut=1.0 KT Dmut=20.0 AS Dmut=0.05 A...

  3. [11]

    Castellano, S

    C. Castellano, S. Fortunato, and V. Loreto, Statistical physics of social dynamics, Rev. Mod. Phys. 81, 591 (2009)

  4. [12]

    tain a sufficient level of diversity and thus enable efficient optimization

    fitted to the AS data with Dmut = 0.1. tain a sufficient level of diversity and thus enable efficient optimization. Qualitatively, mutations therefore play a role similar to driving in athermal physical systems, as they break detailed balance and maintain a steady-state dynamics at...

  5. [13]

    Lorenz, Continuous opinion dynamics under bounded confidence: A survey, International Journal of Modern Physics C 18, 1819 (2007)

    J. Lorenz, Continuous opinion dynamics under bounded confidence: A survey, International Journal of Modern Physics C 18, 1819 (2007)

  6. [14]

    Bredeche, E

    N. Bredeche, E. Haasdijk, and A. Prieto, Embodied evolution in collective robotics: A review, Frontiers in Robotics and AI 5, 10.3389/frobt.2018.00012 (2018)

  7. [15]

    P. Long, T. Fan, X. Liao, W. Liu, H. Zhang, and J. Pan, Towards optimally decentralized multi-robot col- lision avoidance via deep reinforcement learning, in 2018 IEEE international conference on robotics and automa- tion (ICRA) (IEEE, 2018) pp. 6252–6259

  8. [16]

    Zheng, C

    Y. Zheng, C. Huepe, and Z. Han, Experimental capabili- ties and limitations of a position-based control algorithm for swarm robotics, Adaptive Behavior 30, 19 (2022) , https://doi.org/10.1177/1059712320930418

  9. [17]

    M. Y. Ben Zion, J. Fersula, N. Bredeche, and O. Dauchot, Morphological computation and decentralized learning in a swarm of sterically interacting robots, Science Robotics 8, eabo6140 (2023)

  10. [18]

    Bredeche and N

    N. Bredeche and N. Fontbonne, Social learning in swarm robotics, Philosophical Transactions of the Royal Society B: Biological Sciences 377, 20200309 (2022)

  11. [19]

    Verdoucq, G

    M. Verdoucq, G. Theraulaz, R. Escobedo, C. Sire, and G. Hattenberger, Bio-inspired control for collective mo- tion in swarms of drones, in 2022 International Confer- ence on Unmanned Aircraft Systems (ICUAS) (2022) pp. 1626–1631

  12. [20]

    J. Wang, G. Wang, H. Chen, Y. Liu, P. Wang, D. Yuan, X. Ma, X. Xu, Z. Cheng, B. Ji, M. Yang, J. Shuai, F. Ye, J. Wang, Y. Jiao, and L. Liu, Robo-matter towards recon- figurable multifunctional smart materials, Nat. Commun. 15, 8853 (2024)

  13. [21]

    D. Ma, J. Chen, S. Cutler, and K. Petersen, Smarticle 2.0: Design of scalable, entangled smart matter, in Dis- tributed Autonomous Robotic Systems , edited by J. Bour- geois, J. Paik, B. Piranda, J. Werfel, S. Hauert, A. Pier- son, H. Hamann, T. L. Lam, F. Matsuno, N. Mehr, an...

  14. [22]

    Ozkan-Aydin, D

    Y. Ozkan-Aydin, D. I. Goldman, and M. S. Bhamla, Collective dynamics in entangled worm and robot blobs, Proc. Natl. Acad. Sci. USA 118, e2010542118 (2021) . 6

  15. [23]

    S. Li, B. Dutta, S. Cannon, J. J. Daymude, R. Avinery, E. Aydin, A. W. Richa, D. I. Goldman, and D. Randall, Programming active cohesive granular matter with me- chanically induced phase changes, Science Advances 7, eabe8494 (2021)

  16. [24]

    Y. Zhou, J. Zu, and J. Liu, Programmable intelligent liquid matter: material, science and technology, Jour- nal of Micromechanics and Microengineering 32, 103001 (2022)

  17. [25]

    Kotikian, C

    A. Kotikian, C. McMahan, E. C. Davidson, J. M. Muhammad, R. D. Weeks, C. Daraio, and J. A. Lewis, Untethered soft robotic matter with passive control of shape morphing and propulsion, Science Robotics 4, eaax7044 (2019)

  18. [26]

    Zeravcic, V

    Z. Zeravcic, V. N. Manoharan, and M. P. Brenner, Col- loquium: Toward living matter with colloidal particles, Rev. Mod. Phys. 89, 031001 (2017)

  19. [27]

    L. G. Nava, R. Großmann, and F. Peruani, Markovian robots: Minimal navigation strategies for active particles, Phys. Rev. E 97, 042604 (2018)

  20. [28]

    Pishvar and R

    M. Pishvar and R. L. Harne, Foundations for soft, smart matter by active mechanical metamaterials, Advanced Science 7, 2001384 (2020)

  21. [29]

    Cichos, K

    F. Cichos, K. Gustavsson, B. Mehlig, and G. Volpe, Ma- chine learning for active matter, Nat. Mach. Intell. 2, 94–103 (2020)

  22. [30]

    Kaspar, B

    C. Kaspar, B. J. Ravoo, W. G. van der Wiel, S. V. Weg- ner, and W. H. P. Pernice, The rise of intelligent matter, Nature 594, 345 (2021)

  23. [31]

    Levine and D

    H. Levine and D. I. Goldman, Physics of smart active matter: integrating active matter and control to gain insights into living systems, Soft Matter 19, 4204 (2023)

  24. [32]

    D. I. Goldman and D. Z. Rocklin, Robot swarms meet soft matter physics, Science Robotics 9, eadn6035 (2024)

  25. [33]

    Nasiri, E

    M. Nasiri, E. Loran, and B. Liebchen, Smart active par- ticles learn and transcend bacterial foraging strategies, Proceedings of the National Academy of Sciences 121, e2317618121 (2024)

  26. [34]

    G. M. Whitesides, Soft robotics, Angewandte Chemie In- ternational Edition 57, 4258 (2018)

  27. [35]

    Majidi, Soft-matter engineering for soft robotics, Ad- vanced Materials Technologies 4, 1800477 (2019)

    C. Majidi, Soft-matter engineering for soft robotics, Ad- vanced Materials Technologies 4, 1800477 (2019)

  28. [36]

    A. C. H. Tsang, E. Demir, Y. Ding, and O. S. Pak, Roads to smart artificial microswimmers, Advanced Intelligent Systems 2, 1900137 (2020)

  29. [37]

    L. Xu, R. J. Wagner, S. Liu, Q. He, T. Li, W. Pan, Y. Feng, H. Feng, Q. Meng, X. Zou, Y. Fu, X. Shi, D. Zhao, J. Ding, and F. J. Vernerey, Locomotion of an untethered, worm-inspired soft robot driven by a shape- memory alloy skeleton, Sci. Rep. 12, 12392 (2022)

  30. [38]

    Saintyves, M

    B. Saintyves, M. Spenko, and H. M. Jaeger, A self- organizing robotic aggregate using solid and liquid-like collective states, Science Robotics 9, eadh4130 (2024)

  31. [39]

    Savoie, T

    W. Savoie, T. A. Berrueta, Z. Jackson, A. Pervan, R. Warkentin, S. Li, T. D. Murphey, K. Wiesenfeld, and D. I. Goldman, A robot made of robots: Emergent transport and control of a smarticle ensemble, Science Robotics 4, eaax4316 (2019)

  32. [40]

    A. T. Liu, M. Hempel, J. F. Yang, A. M. Brooks, A. Per- van, V. B. Koman, G. Zhang, D. Kozawa, S. Yang, D. I. Goldman, M. Z. Miskin, A. W. Richa, D. Randall, T. D. Murphey, T. Palacios, and M. S. Strano, Colloidal robotics, Nat. Mater. 22, 1453–1462 (2023)

  33. [41]

    Stern, C

    M. Stern, C. Arinze, L. Perez, and A. Murugan, Super- vised learning through physical changes in a mechanical system, Proc. Natl. Acad. Sci. USA 117, 14843 (2020)

  34. [42]

    Stern, M

    M. Stern, M. B. Pinson, and A. Murugan, Continual learning of multiple memories in mechanical networks, Phys. Rev. X 10, 031044 (2020)

  35. [43]

    Dillavou, M

    S. Dillavou, M. Stern, A. J. Liu, and D. J. Durian, Demonstration of decentralized physics-driven learning, Phys. Rev. Appl. 18, 014040 (2022)

  36. [44]

    Mandal, R

    R. Mandal, R. Huang, M. Fruchart, P. G. Moerman, S. Vaikuntanathan, A. Murugan, and V. Vitelli, Learn- ing dynamical behaviors in physical systems (2024), arXiv:2406.07856 [cond-mat.soft]

  37. [45]

    Garrad, G

    M. Garrad, G. Soter, A. T. Conn, H. Hauser, and J. Rossiter, A soft matter computer for soft robots, Sci- ence Robotics 4, eaaw6060 (2019)

  38. [46]

    Muinos-Landin, A

    S. Muinos-Landin, A. Fischer, V. Holubec, and F. Ci- chos, Reinforcement learning with artificial microswim- mers, Science Robotics 6, eabd9285 (2021)

  39. [47]

    Cichos, S

    F. Cichos, S. Mui˜ nos Landin, and R. Pradip, Chapter 5 - artificial intelligence (ai) enhanced nanomotors and active matter, in Intelligent Nanotechnology , Materials Today, edited by Y. Zheng and Z. Wu (Elsevier, 2023) pp. 113–144

  40. [48]

    Colabrese, K

    S. Colabrese, K. Gustavsson, A. Celani, and L. Biferale, Flow navigation by smart microswimmers via reinforce- ment learning, Phys. Rev. Lett. 118, 158004 (2017)

  41. [49]

    Gustavsson, L

    K. Gustavsson, L. Biferale, A. Celani, and S. Co- labrese, Finding efficient swimming strategies in a three- dimensional chaotic flow by reinforcement learning, The European Physical Journal E 40, 1 (2017)

  42. [50]

    Durve, F

    M. Durve, F. Peruani, and A. Celani, Learning to flock through reinforcement, Phys. Rev. E 102, 012601 (2020)

  43. [51]

    M. J. Falk, V. Alizadehyazdi, H. Jaeger, and A. Murugan, Learning to control active matter, Phys. Rev. Res. 3, 033291 (2021)

  44. [52]

    Sankaewtong, J

    K. Sankaewtong, J. J. Molina, M. S. Turner, and R. Ya- mamoto, Learning to swim efficiently in a nonuniform flow field, Phys. Rev. E 107, 065102 (2023)

  45. [53]

    Z. Zou, Y. Liu, A. C. H. Tsang, Y.-N. Young, and O. S. Pak, Adaptive micro-locomotion in a dynamically changing environment via context detection, Communi- cations in Nonlinear Science and Numerical Simulation 128, 107666 (2024)

  46. [54]

    Xiong, Z

    T. Xiong, Z. Liu, C. J. Ong, and L. Zhu, Enabling micro- robotic chemotaxis via reset-free hierarchical reinforce- ment learning (2024), arXiv:2408.07346 [cond-mat.soft]

  47. [55]

    Grauer, F

    J. Grauer, F. J. Schwarzendahl, H. L¨ owen, and B. Liebchen, Optimizing collective behavior of commu- nicating active particles with machine learning, Machine Learning: Science and Technology 5, 015014 (2024)

  48. [56]

    L. G. Mo C. and B. X., Challenges and attempts to make intelligent microswimmers, Front. Phys. 11, 1279883 (2023)

  49. [57]

    Cazenille, M

    L. Cazenille, M. Toquebiau, N. Lobato-Dauzier, A. Loi, L. Macabre, N. Aubert-Kato, A. Genot, and N. Bredeche, Signaling and social learning in swarms of robots (2024), arXiv:2411.11616 [cs.RO]

  50. [58]

    Tennenbaum, Z

    M. Tennenbaum, Z. Liu, D. Hu, and A. Fernandez- Nieves, Mechanics of fire ant aggregations, Nat. Mater. 15, 54–59 (2016)

  51. [59]

    R. J. Wagner and F. J. Vernerey, Computational explo- ration of treadmilling and protrusion growth observed in fire ant rafts, PLOS Computational Biology 18, 1 (2022)

  52. [60]

    Liebchen and H

    B. Liebchen and H. L¨ owen, Optimal navigation strate- 7 gies for active particles, Europhysics Letters 127, 34003 (2019)

  53. [61]

    L. Piro, E. Tang, and R. Golestanian, Optimal navigation strategies for microswimmers on curved manifolds, Phys. Rev. Res. 3, 023125 (2021)

  54. [62]

    L. Piro, B. Mahault, and R. Golestanian, Optimal navi- gation of microswimmers in complex and noisy environ- ments, New Journal of Physics 24, 093037 (2022)

  55. [63]

    L. Piro, R. Golestanian, and B. Mahault, Efficiency of navigation strategies for active particles in rugged land- scapes, Front. Phys. 10, 1034267 (2022)

  56. [64]

    Nasiri, H

    M. Nasiri, H. L¨ owen, and B. Liebchen, Optimal active particle navigation meets machine learning, Europhysics Letters 142, 17001 (2023)

  57. [65]

    Borra, M

    F. Borra, M. Cencini, and A. Celani, Optimal collision avoidance in swarms of active brownian particles, Journal of Statistical Mechanics: Theory and Experiment 2021, 083401 (2021)

  58. [66]

    L. Yang, J. Jiang, X. Gao, Q. Wang, Q. Dou, and L. Zhang, Autonomous environment-adaptive micro- robot swarm navigation enabled by deep learning-based real-time distribution planning, Nature Machine Intelli- gence 4, 480 (2022)

  59. [67]

    Echeverr ´ ıa-Huarte and A

    I. Echeverr ´ ıa-Huarte and A. Nicolas, Body and mind: Decoding the dynamics of pedestrians and the effect of smartphone distraction by coupling mechanical and decisional processes, Transportation Research Part C: Emerging Technologies 157, 104365 (2023)

  60. [68]

    Bonnemain, M

    T. Bonnemain, M. Butano, T. Bonnet, I. n. Echeverr ´ ıa- Huarte, A. Seguin, A. Nicolas, C. Appert-Rolland, and D. Ullmo, Pedestrians in static crowds are not grains, but game players, Phys. Rev. E 107, 024612 (2023)

  61. [69]

    VanSaders and V

    B. VanSaders and V. Vitelli, Informational active matter, arXiv preprint arXiv:2302.07402 (2023)

  62. [70]

    Cocconi, B

    L. Cocconi, B. Mahault, and L. Piro, Dissipation- accuracy tradeoffs in autonomous control of smart active matter (2024), arXiv:2409.12595 [cond-mat.stat-mech]

  63. [71]

    Ziepke, I

    A. Ziepke, I. Maryshev, I. S. Aranson, and E. Frey, Multi- scale organization in communicating active matter, Na- ture communications 13, 6727 (2022)

  64. [72]

    Ramaswamy, The mechanics and statistics of active matter, Annual Review of Condensed Matter Physics 1, 323 (2010)

    S. Ramaswamy, The mechanics and statistics of active matter, Annual Review of Condensed Matter Physics 1, 323 (2010)

  65. [73]

    M. C. Marchetti, J. F. Joanny, S. Ramaswamy, T. B. Liverpool, J. Prost, M. Rao, and R. A. Simha, Hydrody- namics of soft active matter, Rev. Mod. Phys. 85, 1143 (2013)

  66. [74]

    Ramaswamy, Active fluids, Nat

    S. Ramaswamy, Active fluids, Nat. Rev. Phys. 1, 640–642 (2019)

  67. [75]

    Bertin, M

    E. Bertin, M. Droz, and G. Gr´ egoire, Boltzmann and hy- drodynamic description for self-propelled particles, Phys. Rev. E 74, 022101 (2006)

  68. [76]

    Bertin, M

    E. Bertin, M. Droz, and G. Gr´ egoire, Hydrodynamic equations for self-propelled particles: microscopic deriva- tion and stability analysis, Journal of Physics A: Mathe- matical and Theoretical 42, 445001 (2009)

  69. [77]

    Ihle, Kinetic theory of flocking: Derivation of hydro- dynamic equations, Physical Review E—Statistical, Non- linear, and Soft Matter Physics 83, 030901 (2011)

    T. Ihle, Kinetic theory of flocking: Derivation of hydro- dynamic equations, Physical Review E—Statistical, Non- linear, and Soft Matter Physics 83, 030901 (2011)

  70. [78]

    Bertin, H

    E. Bertin, H. Chat´ e, F. Ginelli, S. Mishra, A. Peshkov, and S. Ramaswamy, Mesoscopic theory for fluctuating ac- tive nematics, New journal of physics 15, 085032 (2013)

  71. [79]

    Ihle, Towards a quantitative kinetic theory of polar active matter, The European Physical Journal Special Topics 223, 1293 (2014)

    T. Ihle, Towards a quantitative kinetic theory of polar active matter, The European Physical Journal Special Topics 223, 1293 (2014)

  72. [80]

    See Supplemental Material at [URL will be inserted by publisher] for analytical details on the derivation of the two specific examples and the multi-dimensional theory

  73. [81]

    Chard` es, A

    V. Chard` es, A. Mazzolini, T. Mora, and A. M. Walczak, Evolutionary stability of antigenically escaping viruses, Proceedings of the National Academy of Sciences 120, e2307712120 (2023) , https://www.pnas.org/doi/pdf/10.1073/pnas.2307712120

  74. [82]

    Ewens, An interpretation and proof of the fundamen- tal theorem of natural selection, Theoretical Population Biology 36, 167 (1989)

    W. Ewens, An interpretation and proof of the fundamen- tal theorem of natural selection, Theoretical Population Biology 36, 167 (1989)

  75. [83]

    Romanczuk, M

    P. Romanczuk, M. B¨ ar, W. Ebeling, B. Lindner, and L. Schimansky-Geier, Active brownian particles: From individual to collective stochastic dynamics, The Euro- pean Physical Journal Special Topics 202, 1 (2012)

  76. [84]

    Martin, A

    M. Martin, A. Barzyk, E. Bertin, P. Peyla, and S. Rafai, Photofocusing: Light and flow of phototactic microswim- mer suspension, Phys. Rev. E 93, 051101 (2016)

  77. [85]

    Risken and H

    H. Risken and H. Risken, Fokker-planck equation (Springer, 1996). End Matter – We describe here the general frame- work of kinetic theory of decentralized learning intro- duced in this work. Calculation details for specific models are reported in the SM [ 74]. The term Iteach[f...

  78. [86]

    into an evolution equation for the probability distribution of M [79] leads to the explicit relation for the memory dynamics, Imem[f ] = τ −1 M ( f (S, M, P) + (M −G(S)) ∂f (S, M, P) ∂M ) , (21) we find a closed equation for the time-dependence of the marginal distribution, ∂f0...

  79. [87]

    Kinetic theory of decentralized learning for smart active matter

    becomes, ∂fP,0(M) ∂t = −∇J0 + τ −1 M ( f0 + (M − ¯GP ) ∂fP,0 ∂M ) +2λT f0 ∫ dM′ tanh(αT (R − R′)) ∑ P ′∈P0 fP ′,0(M′) + ∑ P ′∈P0 λmut,P ′→P fP ′,0. (25) The only differences to Eq. ( 22) are the replacement of the integral by a sum, and introducing a discrete master equation de...

Pith tools

Reviewed August 10, 2026 · model on record in the stance chip above.