Pith. sign in

Paper Citation Record · LEDGER

Pretraining in Actor-Critic Reinforcement Learning for Locomotion

As of 18 August 2026, this Paper Citation Record lists 62 of 62 outbound references and 0 inbound Pith citation observations for arXiv:2510.12363.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2510.12363 v4

Coverage vector

measured 62 of 62 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T10:03:34.319088Z

measured 62 of 62 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

62 of 62 outbound references displayed

  • verified exact1
  • verified fuzzy0
  • unresolved60
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f15addc7-28da-460f-b71d-971be214492b · outbound

This paper cites GPT-4 Technical Report.

Pretraining in Actor-Critic Reinforcement Learning for Locomotion GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T10:03:27.741507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:03:27.741507Z digest=sha256:cb6c63c397efbcc4dfa30c3d6dccd77fb731faa9031e0c129a7bb46acadc8fc3

Observation 2affba33-b178-456a-94eb-b934c1f5b208 · outbound

This paper cites Learning markov state abstractions for deep reinforcement learning.

Pretraining in Actor-Critic Reinforcement Learning for Locomotion Learning markov state abstractions for deep reinforcement learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-04T10:03:27.934564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:03:27.934564Z digest=sha256:c2123be0948894c412470c1f48929573497f8c9de20e6bfe9f6777ee0ea79409

Observation c9f7fe64-278c-4da5-8650-2a6a0b25df32 · outbound

This paper cites Pedipulate: Enabling Manipulation Skills using a Quadruped Robot 's Leg , 2024.

Pretraining in Actor-Critic Reinforcement Learning for Locomotion Pedipulate: Enabling Manipulation Skills using a Quadruped Robot 's Leg , 2024

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T10:03:28.112596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:03:28.112596Z digest=sha256:783ae9e53b363c7053f21d39479d7a50d24f9bdab0600d4c3be01aef71022fff

Observation ef68a4bb-1729-4e5d-85f4-cd39121c6e7c · outbound

This paper cites Scaling mlps: A tale of inductive bias.

Pretraining in Actor-Critic Reinforcement Learning for Locomotion Scaling mlps: A tale of inductive bias

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T10:03:28.249358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:03:28.249358Z digest=sha256:b02314ac2c1e7b3538281e7915abfbdc03719db81bf1aeabe4a99ad9b1f5182a

Observation 80327dcc-2f64-48fe-b0db-a5206da5b4be · outbound

This paper cites A Careful Examination of Large Behavior Models for Multitask Dexterous Manipulation.

Pretraining in Actor-Critic Reinforcement Learning for Locomotion A Careful Examination of Large Behavior Models for Multitask Dexterous Manipulation

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T10:03:28.316077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:03:28.316077Z digest=sha256:0798934d4b02406132e64cdf5c782a87b3c1a0a3975812036428a3efa635182e

Observation b07a8e93-b328-4795-96e7-077d6dc56b0b · outbound

This paper cites Dario Bellicoso, Koen Krämer, Markus Stäuble, Dhionis Sako, Fabian Jenelten, Marko Bjelonic, and Marco Hutter.

Pretraining in Actor-Critic Reinforcement Learning for Locomotion Dario Bellicoso, Koen Krämer, Markus Stäuble, Dhionis Sako, Fabian Jenelten, Marko Bjelonic, and Marco Hutter

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T10:03:28.430158Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:03:28.430158Z digest=sha256:3477c0d67a716419908e53f39a6fdc767e1164964dab252026500ff12bd10267

Observation e5795325-1ce7-43c4-8afd-fea62184fc21 · outbound

This paper cites $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control.

Pretraining in Actor-Critic Reinforcement Learning for Locomotion $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T10:03:28.554612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:03:28.554612Z digest=sha256:3779e35072e4ee8739468416b530288888b8f3f02411fd202596635418dd4aa5

Observation 3514963b-92f6-4a69-99f2-bd323b52994a · outbound

This paper cites RT -2: Vision - Language - Action Models Transfer Web Knowledge to Robotic Control.

Pretraining in Actor-Critic Reinforcement Learning for Locomotion RT -2: Vision - Language - Action Models Transfer Web Knowledge to Robotic Control

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T10:03:28.674648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:03:28.674648Z digest=sha256:95a52c991fd1e0ffac2f97be4ef0ba41085d26dbaa702e141a73677fee923347

Observation c592f4f9-47c4-4a93-a14a-e06d16c38e58 · outbound

This paper cites Symmetric Reinforcement Learning Loss for Robust Learning on Diverse Tasks and Model Scales.

Pretraining in Actor-Critic Reinforcement Learning for Locomotion Symmetric Reinforcement Learning Loss for Robust Learning on Diverse Tasks and Model Scales

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T10:03:28.803985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:03:28.803985Z digest=sha256:6817fb715f32e1a878aa23c79460bc3789d0b464ebc735983b3998bee1e4566b

Observation 7f716a29-44bd-4cbd-b1ff-2b30ce74574c · outbound

This paper cites Learning quadrupedal locomotion on deformable terrain.

Pretraining in Actor-Critic Reinforcement Learning for Locomotion Learning quadrupedal locomotion on deformable terrain

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T10:03:28.919419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:03:28.919419Z digest=sha256:6cbfa3e786e8c3881b531f53c99030fc11b91b25dc5a26cfddbcfed070e0af33

Observation 0e3d4d30-45c6-4df2-a7f1-e33638c2408f · outbound

This paper cites Transfer from Simulation to Real World through Learning Deep Inverse Dynamics Model , 2016.

Pretraining in Actor-Critic Reinforcement Learning for Locomotion Transfer from Simulation to Real World through Learning Deep Inverse Dynamics Model , 2016

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T10:03:29.008153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:03:29.008153Z digest=sha256:f2636fdf0d05fc8b8eefd67f015930a3577a4598b71aadf80211372bc674f761

Observation e7b80cc8-74d3-4886-baea-b64c506064ab · outbound

This paper cites Deep reinforcement learning in a handful of trials using probabilistic dynamics models.

Pretraining in Actor-Critic Reinforcement Learning for Locomotion Deep reinforcement learning in a handful of trials using probabilistic dynamics models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T10:03:29.150862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:03:29.150862Z digest=sha256:144bb4adde974aaa420616b152fa393dde3dfe3553b2b2f6277f15e1d1ff2fcb

Observation 16ba2e58-1971-48ef-8948-a52e7b83341a · outbound

This paper cites Efficient model-based reinforcement learning through optimistic policy search and planning.

Pretraining in Actor-Critic Reinforcement Learning for Locomotion Efficient model-based reinforcement learning through optimistic policy search and planning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T10:03:29.312751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:03:29.312751Z digest=sha256:de628fe3e9ceef8211549ad5b0512078bc36a2e6eb57363716d53988ca649722

Observation edaf0117-a842-4c2e-86fd-223adf5a3d99 · outbound

This paper cites BERT : Pre-training of deep bidirectional transformers for language understanding.

Pretraining in Actor-Critic Reinforcement Learning for Locomotion BERT : Pre-training of deep bidirectional transformers for language understanding

Reference 14

Resolution
malformed identifier
no resolver link, observed 2026-08-04T10:03:29.476493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:03:29.476493Z digest=sha256:c9cf644dddcfa2b99e615b6faa6a7f4bfbde856e448cb2969ec52588dbe472dd

Observation ac754f07-fbaa-4e61-a25c-35c30c927275 · outbound

This paper cites Roloma: Robust loco-manipulation for quadruped robots with arms.

Pretraining in Actor-Critic Reinforcement Learning for Locomotion Roloma: Robust loco-manipulation for quadruped robots with arms

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T10:03:29.586681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:03:29.586681Z digest=sha256:2eee67aead7f7da858e0b3d7db391e1d8f542676d40ce17acd1bf5d9594e1a31

Observation b5b79247-7cdd-428f-a2c6-7c0e617fe6c7 · outbound

This paper cites Delving deep into rectifiers: Surpassing human-level performance on imagenet classification.

Pretraining in Actor-Critic Reinforcement Learning for Locomotion Delving deep into rectifiers: Surpassing human-level performance on imagenet classification

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T10:03:29.667876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:03:29.667876Z digest=sha256:cc8acfdd0a852edf0734016c7821cb28242f2370bf223eea167bcdcb3c58efcc

Observation 41482c13-611f-48bf-a375-14c43c80edf8 · outbound

This paper cites Masked autoencoders are scalable vision learners.

Pretraining in Actor-Critic Reinforcement Learning for Locomotion Masked autoencoders are scalable vision learners

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T10:03:29.744272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:03:29.744272Z digest=sha256:e6f35063ccb619defd578d48e3f2d2b357b5be1dbcad77fb246a8bb4aa187d4c

Observation 5594a5f6-a13f-4886-a309-84a9f27e46e1 · outbound

This paper cites ANYmal Parkour : Learning Agile Navigation for Quadrupedal Robots , 2023.

Pretraining in Actor-Critic Reinforcement Learning for Locomotion ANYmal Parkour : Learning Agile Navigation for Quadrupedal Robots , 2023

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T10:03:29.849265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:03:29.849265Z digest=sha256:6917f2998d05955e35472f4ce75472f5ee7fe86ad19e6b6b420a90ff878dda53

Observation 6340f53c-ba44-446c-9a66-8a8d01786479 · outbound

This paper cites Dario Bellicoso, Vassilios Tsounis, Jemin Hwangbo, Karen Bodie, Peter Fankhauser, Michael Bloesch, Remo Diethelm, Samuel Bachmann, Amir Melzer, and Mark Hoepflinger.

Pretraining in Actor-Critic Reinforcement Learning for Locomotion Dario Bellicoso, Vassilios Tsounis, Jemin Hwangbo, Karen Bodie, Peter Fankhauser, Michael Bloesch, Remo Diethelm, Samuel Bachmann, Amir Melzer, and Mark Hoepflinger

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T10:03:29.953190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:03:29.953190Z digest=sha256:6c5c32d7950e96e5f59d921b9f732d9fe959295fec6eb0f76c370aca5d1a13e0

Observation 669f2f07-67dd-4779-bbc0-846e3ae63956 · outbound

This paper cites Learning agile and dynamic motor skills for legged robots.

Pretraining in Actor-Critic Reinforcement Learning for Locomotion Learning agile and dynamic motor skills for legged robots

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T10:03:30.089031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:03:30.089031Z digest=sha256:cfcdf57bcae93237560f8a497f6f85c866cf56e02355f218b132cdf3328a9c88

Observation 10c998ed-53b6-4904-8ad3-5c7c84b0ce88 · outbound

This paper cites Bellman Eluder Dimension: New Rich Classes of RL Problems, and Sample-Efficient Algorithms.

Pretraining in Actor-Critic Reinforcement Learning for Locomotion Bellman Eluder Dimension: New Rich Classes of RL Problems, and Sample-Efficient Algorithms

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-04T10:03:30.180905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:03:30.180905Z digest=sha256:76e749df6575d86511edbae665d40f1be7f19f4c27a8f562ee6e2ae277feef87

Observation 08439231-3013-4d01-9fd5-9dd36269bf6b · outbound

This paper cites The Role of Domain Randomization in Training Diffusion Policies for Whole-Body Humanoid Control.

Pretraining in Actor-Critic Reinforcement Learning for Locomotion The Role of Domain Randomization in Training Diffusion Policies for Whole-Body Humanoid Control

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T10:03:30.308293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:03:30.308293Z digest=sha256:5ab8d49f9cf5d1460df1ebc0f3add8108f21051b26751f04b0ddfef492c3df62

Observation 74fc8085-0de7-424d-8198-0aedeeb842ca · outbound

This paper cites Adam: A Method for Stochastic Optimization.

Pretraining in Actor-Critic Reinforcement Learning for Locomotion Adam: A Method for Stochastic Optimization

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T10:03:30.394974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:03:30.394974Z digest=sha256:173c0edaccf8080e2d2113b99f24b47e82f1919dbdf04ad867677bd6ca84120a

Observation ecfe473e-fcf0-4964-8b31-caf6109d252e · outbound

This paper cites Actor-critic algorithms.

Pretraining in Actor-Critic Reinforcement Learning for Locomotion Actor-critic algorithms

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T10:03:30.428840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:03:30.428840Z digest=sha256:d626bb3452ae421d0be03c993c6d90ef11145b36cb20e4445a3976438a391308

Observation 8f01d23c-e5f1-4e4e-b97b-61dd58461855 · outbound

This paper cites Learning quadrupedal locomotion over challenging terrain.

Pretraining in Actor-Critic Reinforcement Learning for Locomotion Learning quadrupedal locomotion over challenging terrain

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-04T10:03:30.508818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:03:30.508818Z digest=sha256:a47fcdb8e3699c21bef086febc16a96e7c6f54124e091742306c27b422128d54

Observation afc4d0f2-b871-4947-b693-7132bc5c1050 · outbound

This paper cites Learning to Walk from Three Minutes of Real - World Data with Semi -structured Dynamics Models , 2024.

Pretraining in Actor-Critic Reinforcement Learning for Locomotion Learning to Walk from Three Minutes of Real - World Data with Semi -structured Dynamics Models , 2024

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T10:03:30.562093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:03:30.562093Z digest=sha256:8ea8a8cbb9c640fd30636a56efc6592a35ccb072e304b46bfe687619455f64b8

Observation 4a4d8fb1-e292-4dc3-925e-709f08923fa9 · outbound

This paper cites A Survey : Learning Embodied Intelligence from Physical Simulators and World Models , 2025.

Pretraining in Actor-Critic Reinforcement Learning for Locomotion A Survey : Learning Embodied Intelligence from Physical Simulators and World Models , 2025

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-04T10:03:30.635467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:03:30.635467Z digest=sha256:a9fae22e5bc6811cea8569f8fc33e50c34f74d954311d6995fdc4460e1faec1d

Observation af29eafa-30a8-482a-a7a8-4fa71a91ed1b · outbound

This paper cites Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks.

Pretraining in Actor-Critic Reinforcement Learning for Locomotion Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-04T10:03:30.714843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:03:30.714843Z digest=sha256:4f27f6b2091c9eefeaf041310d2fc90ab98eefb5d12aae50b769cafc5e1726a6

Observation 4fdc2f0b-89a6-46b4-bf0c-d451a8c7dd29 · outbound

This paper cites Combining physics and deep learning to learn continuous-time dynamics models.

Pretraining in Actor-Critic Reinforcement Learning for Locomotion Combining physics and deep learning to learn continuous-time dynamics models

Reference 29

Resolution
verified exact
doi, observed 2026-08-04T10:08:39.069158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-04T10:03:30.778393Z digest=sha256:5f53c30474e323ea36dbd22f158aa58885574d62c8fb94240267384de6fdf9fa

Observation 3b3c0a4f-6c29-4485-a882-771da65cbd07 · outbound

This paper cites UMAP: Uniform Manifold Approximation and Projection for Dimension Reduction.

Pretraining in Actor-Critic Reinforcement Learning for Locomotion UMAP: Uniform Manifold Approximation and Projection for Dimension Reduction

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-04T10:03:30.827491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:03:30.827491Z digest=sha256:ab627065b892f076ea28b95d3280b0ee3015bbb839ebcb095ae63a302fcc7812

Observation 51032244-5508-49df-a6f8-c77b79d71d82 · outbound

This paper cites Learning robust perceptive locomotion for quadrupedal robots in the wild.

Pretraining in Actor-Critic Reinforcement Learning for Locomotion Learning robust perceptive locomotion for quadrupedal robots in the wild

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-04T10:03:30.921104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:03:30.921104Z digest=sha256:162a0c7d0a06e34a0a07f2ddba273b54f06b618775e20edee5c5071e331575df

Observation 5f1577d2-c062-4549-b5ca-1eb502819c71 · outbound

This paper cites Orbit: A unified simulation framework for interactive robot learning environments.

Pretraining in Actor-Critic Reinforcement Learning for Locomotion Orbit: A unified simulation framework for interactive robot learning environments

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-04T10:03:30.967138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:03:30.967138Z digest=sha256:ecac583d0f2f22cec3fae45d1b508a5b05765a4225ee87c13d316c3fb95275cf

Observation 9d1ddd80-807e-42f9-8996-cee712795bd9 · outbound

This paper cites Symmetry considerations for learning task symmetric robot policies.

Pretraining in Actor-Critic Reinforcement Learning for Locomotion Symmetry considerations for learning task symmetric robot policies

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-04T10:03:31.029123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:03:31.029123Z digest=sha256:e5698bd3c806ecd806946d2321de49e09ec4558f307da1f73235e5cacda320b7

Observation 3b8908e4-05e2-4906-b8bb-6f0a35fb9460 · outbound

This paper cites Murphy, Benjamin J.

Pretraining in Actor-Critic Reinforcement Learning for Locomotion Murphy, Benjamin J

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-04T10:03:31.051659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:03:31.051659Z digest=sha256:5baea21e292486852df47c774c7690ea3d2737a5167f23a6a97ee1defec83040

Observation e64f4e7d-64d3-46be-9ee8-a7e3687320a1 · outbound

This paper cites Information-Directed Exploration for Deep Reinforcement Learning.

Pretraining in Actor-Critic Reinforcement Learning for Locomotion Information-Directed Exploration for Deep Reinforcement Learning

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-04T10:03:31.128989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:03:31.128989Z digest=sha256:4a13469931d6e374250b9aa33fa5e8f282daf3642699e7fac9e0b19a8e9c06af

Observation ef271075-0b1c-45af-9c5d-b0198420bb88 · outbound

This paper cites AMP: Adversarial Motion Priors for Stylized Physics-Based Character Control.

Pretraining in Actor-Critic Reinforcement Learning for Locomotion AMP: Adversarial Motion Priors for Stylized Physics-Based Character Control

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-04T10:03:31.288541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:03:31.288541Z digest=sha256:d5c67f199991dc33e172f2478c6ffae057f7813bab9294efee5b5c00237b828d

Observation 064f7195-f83e-4fd7-baea-b83735b3e5ca · outbound

This paper cites ASE : large-scale reusable adversarial skill embeddings for physically simulated characters.

Pretraining in Actor-Critic Reinforcement Learning for Locomotion ASE : large-scale reusable adversarial skill embeddings for physically simulated characters

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-04T10:03:31.431057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:03:31.431057Z digest=sha256:e67ee42ad496ccaf9294eb714a40e8179429099e6d66a66a835c9d6729c5327e

Observation 7746778e-0817-4026-9edc-9cf8c71397ca · outbound

This paper cites Whole-body end-effector pose tracking, 2024.

Pretraining in Actor-Critic Reinforcement Learning for Locomotion Whole-body end-effector pose tracking, 2024

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-04T10:03:31.553721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:03:31.553721Z digest=sha256:c2ae6647eb5f281fac648624ce475b66b35cc65686ead8f297b64ff278a848d4

Observation ccd69885-dcd6-4e78-bd2f-9fc122b88fd0 · outbound

This paper cites Language models are unsupervised multitask learners.

Pretraining in Actor-Critic Reinforcement Learning for Locomotion Language models are unsupervised multitask learners

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-04T10:03:31.707213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:03:31.707213Z digest=sha256:bf9b738291b7476681768e191b767345e2312cba3d57e36da6f0becdb25d031e

Observation 6062ef33-4f90-4afc-9db4-82b0d4a3a133 · outbound

This paper cites Learning to walk in minutes using massively parallel deep reinforcement learning.

Pretraining in Actor-Critic Reinforcement Learning for Locomotion Learning to walk in minutes using massively parallel deep reinforcement learning

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-04T10:03:31.866986Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:03:31.866986Z digest=sha256:e8c9d7b6084f80aa3ac733cea19a955c4c713eecb64d0b36e12c0c4bba3b96c4

Observation 808ac16b-76dd-43dc-9d59-e05db3c19195 · outbound

This paper cites Parkour in the Wild : Learning a General and Extensible Agile Locomotion Policy Using Multi -expert Distillation and RL Fine -tuning, 2025.

Pretraining in Actor-Critic Reinforcement Learning for Locomotion Parkour in the Wild : Learning a General and Extensible Agile Locomotion Policy Using Multi -expert Distillation and RL Fine -tuning, 2025

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-04T10:03:31.980649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:03:31.980649Z digest=sha256:78036890051f49a3e2c48d4119ea0b7c3c9f16f17beb81feacee401fe9bd4164

Observation df215aec-7d86-4320-884b-0e5eb467c7ef · outbound

This paper cites Proximal Policy Optimization Algorithms.

Pretraining in Actor-Critic Reinforcement Learning for Locomotion Proximal Policy Optimization Algorithms

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-04T10:03:32.155903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:03:32.155903Z digest=sha256:c4eba40ef18a179b4accfb54fd56f7ecc839c5a0e9eaaa8ee6b33126c423721d

Observation 22d664b8-f072-4b50-acea-3046f2fa3812 · outbound

This paper cites RSL-RL: A Learning Library for Robotics Research.

Pretraining in Actor-Critic Reinforcement Learning for Locomotion RSL-RL: A Learning Library for Robotics Research

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-04T10:03:32.299059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:03:32.299059Z digest=sha256:d91d5031c6121724ac205da094b2f54950158fa293ee15950db2e39c04c83946

Observation d75942fd-dba1-42da-890a-104e6d089818 · outbound

This paper cites Devon Hjelm, Philip Bachman, and Aaron C.

Pretraining in Actor-Critic Reinforcement Learning for Locomotion Devon Hjelm, Philip Bachman, and Aaron C

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-04T10:03:32.469853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:03:32.469853Z digest=sha256:844ff059c77e3364ecad9b2f97dd781a5666d6d2e8d5f4c0d609226846aaedeb

Observation 635783d2-0bd9-4137-9b8e-e962ae853a9d · outbound

This paper cites Planning to explore via self-supervised world models.

Pretraining in Actor-Critic Reinforcement Learning for Locomotion Planning to explore via self-supervised world models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-04T10:03:32.560939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:03:32.560939Z digest=sha256:1f5ed1d09139c2fbc672ec64432486e848f23ce8439401b3bbcebf24e929245a

Observation dd9f9ead-8bea-47ce-bb5c-4210fca988c7 · outbound

This paper cites Blind Bipedal Stair Traversal via Sim-to-Real Reinforcement Learning.

Pretraining in Actor-Critic Reinforcement Learning for Locomotion Blind Bipedal Stair Traversal via Sim-to-Real Reinforcement Learning

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-04T10:03:32.644933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:03:32.644933Z digest=sha256:d8fa7ac65860bb5f9236e6fadc6cba8a12ce3f19bd8ff0aede6366ac1db66da7

Observation 0c2add18-d72e-4490-8f31-36d43eb586a7 · outbound

This paper cites A Unified MPC Framework for Whole-Body Dynamic Locomotion and Manipulation.

Pretraining in Actor-Critic Reinforcement Learning for Locomotion A Unified MPC Framework for Whole-Body Dynamic Locomotion and Manipulation

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-04T10:03:32.760620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:03:32.760620Z digest=sha256:9886651e90bbc5dbc6522e4881f69eb52a72f72db0a69200b2f97efccff08b77

Observation 556918e7-efe7-47be-b5b0-a76e796d2ce2 · outbound

This paper cites Guided Reinforcement Learning for Robust Multi - Contact Loco - Manipulation , 2024.

Pretraining in Actor-Critic Reinforcement Learning for Locomotion Guided Reinforcement Learning for Robust Multi - Contact Loco - Manipulation , 2024

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-04T10:03:32.828377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:03:32.828377Z digest=sha256:858077fdb159699d936a7a9bf56544d8ebfddc81cfb564bc3395efe638851c0e

Observation b55a044d-5d02-447e-9d50-e8e07a1bb76e · outbound

This paper cites Perceptive Pedipulation with Local Obstacle Avoidance , 2024.

Pretraining in Actor-Critic Reinforcement Learning for Locomotion Perceptive Pedipulation with Local Obstacle Avoidance , 2024

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-04T10:03:32.859664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:03:32.859664Z digest=sha256:31416da72dd489804b118fce4e007fc09fa5334158d8b0c45a1fa3e3287b54bd

Observation 16b6cdf5-1ab6-4da4-9f48-d6a3ca6eddea · outbound

This paper cites Gemini Robotics: Bringing AI into the Physical World.

Pretraining in Actor-Critic Reinforcement Learning for Locomotion Gemini Robotics: Bringing AI into the Physical World

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-04T10:03:32.943196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:03:32.943196Z digest=sha256:0c6356d402f6af0fa9f76c0bcd8ea3f6db4d0da46100fef6bcc8cea0f60ec8ae

Observation 729cd458-865d-47fd-8b9e-db05840368d5 · outbound

This paper cites Octo: An Open-Source Generalist Robot Policy.

Pretraining in Actor-Critic Reinforcement Learning for Locomotion Octo: An Open-Source Generalist Robot Policy

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-04T10:03:32.996524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:03:32.996524Z digest=sha256:3eba952301d8268015f8879e7c26bc5d7889001e67d7881c8dc8fb8fd96e08b6

Observation 7c9ac8a9-0987-4798-ba2b-c9eea7696d2b · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Pretraining in Actor-Critic Reinforcement Learning for Locomotion Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-04T10:03:33.116502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:03:33.116502Z digest=sha256:3f43741af003a05f312b793be0b1e5502de631211645210324bbcc2490b97aef

Observation 67680e13-674a-45b2-8344-3c1051849caf · outbound

This paper cites Advanced Skills through Multiple Adversarial Motion Priors in Reinforcement Learning.

Pretraining in Actor-Critic Reinforcement Learning for Locomotion Advanced Skills through Multiple Adversarial Motion Priors in Reinforcement Learning

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-04T10:03:33.186830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:03:33.186830Z digest=sha256:95d7a9bd5480080d3170bad5af24092b3a91afe0beb62af56c47c436d9479a5d

Observation 09a85d7d-c0f3-4016-8ec8-9ef83cfdd9a4 · outbound

This paper cites Pretraining in Deep Reinforcement Learning : A Survey , 2022.

Pretraining in Actor-Critic Reinforcement Learning for Locomotion Pretraining in Deep Reinforcement Learning : A Survey , 2022

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-04T10:03:33.336742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:03:33.336742Z digest=sha256:e0a90d27055180006c1d0cc4d0f9afb23402653196ebf659aea2f1183381f099

Observation 2cadf785-2002-4825-9e71-e5fb6d831ca3 · outbound

This paper cites Neural Robot Dynamics.

Pretraining in Actor-Critic Reinforcement Learning for Locomotion Neural Robot Dynamics

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-04T10:03:33.453741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:03:33.453741Z digest=sha256:c989f6223b086f67175171cba20c49fc2a22dab90a70cce8eca58c32c87abd4a

Observation 0720e17b-93fc-47e2-baf5-2e48e902188f · outbound

This paper cites Neural Volumetric Memory for Visual Locomotion Control.

Pretraining in Actor-Critic Reinforcement Learning for Locomotion Neural Volumetric Memory for Visual Locomotion Control

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-04T10:03:33.537036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:03:33.537036Z digest=sha256:a66d9100d3cb207d445db88862794cabbc648bdb0b3ebc71402b9d0a3920f505

Observation 38f54183-f9cb-48f4-a584-dbf9ffd60c9a · outbound

This paper cites Distillation-PPO: A Novel Two-Stage Reinforcement Learning Framework for Humanoid Robot Perceptive Locomotion.

Pretraining in Actor-Critic Reinforcement Learning for Locomotion Distillation-PPO: A Novel Two-Stage Reinforcement Learning Framework for Humanoid Robot Perceptive Locomotion

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-04T10:03:33.681668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:03:33.681668Z digest=sha256:97138bbfafc142c5792b6f35e192bf33e3cbf85061c914f685a89bbf8d48f993

Observation 9732485e-cbf4-4a40-8b4f-fd73d5922e81 · outbound

This paper cites Intention- Conditioned Flow Occupancy Models , 2025.

Pretraining in Actor-Critic Reinforcement Learning for Locomotion Intention- Conditioned Flow Occupancy Models , 2025

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-04T10:03:33.824504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:03:33.824504Z digest=sha256:32287b10979b3b84fdd5959b90b1f0014c455e763b61bffc948cb5adffef850e

Observation 671711d3-e693-4cf3-ab56-581b6eff25de · outbound

This paper cites write newline.

Pretraining in Actor-Critic Reinforcement Learning for Locomotion write newline

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-04T10:03:33.977983Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:03:33.977983Z digest=sha256:056489dd573f05b59ac72f783b19a1c0a1e85a54d31f94ed2bb176d4d8d864df

Observation f6fdb2e7-a1c5-4abf-af22-05565b03b945 · outbound

This paper cites @esa (Ref.

Pretraining in Actor-Critic Reinforcement Learning for Locomotion @esa (Ref

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-04T10:03:34.095585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:03:34.095585Z digest=sha256:20b97d24bbc8bad426be3a9319340d50c4d54a21f06a07734a6f106bbcf4035e

Observation c2278c11-d25d-4377-bf2c-d4b6aae46917 · outbound

This paper cites an unresolved cited work.

Pretraining in Actor-Critic Reinforcement Learning for Locomotion Unresolved cited work

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-04T10:03:34.173302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:03:34.173302Z digest=sha256:07e54a3d785ca338d6abf26aad521520a78d4c02494e5f71a8653c22b0acbc47

Observation 316195b7-0025-4aaf-ad1a-4d3106a29923 · outbound

This paper cites an unresolved cited work.

Pretraining in Actor-Critic Reinforcement Learning for Locomotion Unresolved cited work

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-04T10:03:34.319088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:03:34.319088Z digest=sha256:0e5ee58dc727587129ef780cad8f9f9188e7d015e80fe8352891e57a2a952faf

Pith citing papers

No inbound Pith citation observations are available.