Pith. sign in

Paper Citation Record · LEDGER

Pretraining in Actor-Critic Reinforcement Learning for Locomotion

As of 11 August 2026, this Paper Citation Record lists 62 of 62 outbound references and 0 inbound Pith citation observations for arXiv:2510.12363.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2510.12363 v4

Coverage vector

measured 62 of 62 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T10:03:34.319088Z

measured 62 of 62 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

62 of 62 outbound references displayed

  • verified exact1
  • verified fuzzy0
  • unresolved60
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f15addc7-28da-460f-b71d-971be214492b · outbound

This paper cites GPT-4 Technical Report.

Pretraining in Actor-Critic Reinforcement Learning for Locomotion GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T10:03:27.741507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:03:27.741507Z digest=sha256:de0c5549b713553df845f133e4ab9d2888d0ba3db7da88611d8451e5d9722e38

Observation 2affba33-b178-456a-94eb-b934c1f5b208 · outbound

This paper cites Learning markov state abstractions for deep reinforcement learning.

Pretraining in Actor-Critic Reinforcement Learning for Locomotion Learning markov state abstractions for deep reinforcement learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-04T10:03:27.934564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:03:27.934564Z digest=sha256:d68851ecf7872662a1ce8c0839b8baf65af94a7d4b2f673e82cc9007f1d1e39f

Observation c9f7fe64-278c-4da5-8650-2a6a0b25df32 · outbound

This paper cites Pedipulate: Enabling Manipulation Skills using a Quadruped Robot 's Leg , 2024.

Pretraining in Actor-Critic Reinforcement Learning for Locomotion Pedipulate: Enabling Manipulation Skills using a Quadruped Robot 's Leg , 2024

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T10:03:28.112596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:03:28.112596Z digest=sha256:834438b8f04a38de77e85ac7762398c65d9b138dafbdcbff40d2a4da2bfcf44d

Observation ef68a4bb-1729-4e5d-85f4-cd39121c6e7c · outbound

This paper cites Scaling mlps: A tale of inductive bias.

Pretraining in Actor-Critic Reinforcement Learning for Locomotion Scaling mlps: A tale of inductive bias

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T10:03:28.249358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:03:28.249358Z digest=sha256:b4e9f2d142a0c1d314c9b69d8138600de461ac7cbedf14adfcdfa8c3f21411ed

Observation 80327dcc-2f64-48fe-b0db-a5206da5b4be · outbound

This paper cites A Careful Examination of Large Behavior Models for Multitask Dexterous Manipulation.

Pretraining in Actor-Critic Reinforcement Learning for Locomotion A Careful Examination of Large Behavior Models for Multitask Dexterous Manipulation

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T10:03:28.316077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:03:28.316077Z digest=sha256:e2054fd50fe171e42067efde338d6d9386eb2ce2924577e673fbbb0c24bd0d85

Observation b07a8e93-b328-4795-96e7-077d6dc56b0b · outbound

This paper cites Dario Bellicoso, Koen Krämer, Markus Stäuble, Dhionis Sako, Fabian Jenelten, Marko Bjelonic, and Marco Hutter.

Pretraining in Actor-Critic Reinforcement Learning for Locomotion Dario Bellicoso, Koen Krämer, Markus Stäuble, Dhionis Sako, Fabian Jenelten, Marko Bjelonic, and Marco Hutter

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T10:03:28.430158Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:03:28.430158Z digest=sha256:768dc73750f68ebb3f0dce31d91ece5e00878570cc04da5b1d1bdcbc1177446c

Observation e5795325-1ce7-43c4-8afd-fea62184fc21 · outbound

This paper cites $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control.

Pretraining in Actor-Critic Reinforcement Learning for Locomotion $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T10:03:28.554612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:03:28.554612Z digest=sha256:8ebf9718ecb60a7675f0c838991e187bd3e71ae5e51b330844cba7cbb8c7ce3e

Observation 3514963b-92f6-4a69-99f2-bd323b52994a · outbound

This paper cites RT -2: Vision - Language - Action Models Transfer Web Knowledge to Robotic Control.

Pretraining in Actor-Critic Reinforcement Learning for Locomotion RT -2: Vision - Language - Action Models Transfer Web Knowledge to Robotic Control

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T10:03:28.674648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:03:28.674648Z digest=sha256:65c0a7150cfa4be18cf07f4d5dfa45e6d943ff5ea85044e1526de0c6d51af619

Observation c592f4f9-47c4-4a93-a14a-e06d16c38e58 · outbound

This paper cites Symmetric Reinforcement Learning Loss for Robust Learning on Diverse Tasks and Model Scales.

Pretraining in Actor-Critic Reinforcement Learning for Locomotion Symmetric Reinforcement Learning Loss for Robust Learning on Diverse Tasks and Model Scales

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T10:03:28.803985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:03:28.803985Z digest=sha256:eff68935b7e1204c9a0c333c2ab81b61841542446df5d34907acb5b648b4377e

Observation 7f716a29-44bd-4cbd-b1ff-2b30ce74574c · outbound

This paper cites Learning quadrupedal locomotion on deformable terrain.

Pretraining in Actor-Critic Reinforcement Learning for Locomotion Learning quadrupedal locomotion on deformable terrain

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T10:03:28.919419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:03:28.919419Z digest=sha256:1c58eaec395f85ad365a5d0e54195b38ea75bae4b5bd2d048431dab886cf9e0b

Observation 0e3d4d30-45c6-4df2-a7f1-e33638c2408f · outbound

This paper cites Transfer from Simulation to Real World through Learning Deep Inverse Dynamics Model , 2016.

Pretraining in Actor-Critic Reinforcement Learning for Locomotion Transfer from Simulation to Real World through Learning Deep Inverse Dynamics Model , 2016

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T10:03:29.008153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:03:29.008153Z digest=sha256:e14ad4d9037c3032c3cf3cd63565f0d13f5a6221478a6690292aa1b6c900cf7b

Observation e7b80cc8-74d3-4886-baea-b64c506064ab · outbound

This paper cites Deep reinforcement learning in a handful of trials using probabilistic dynamics models.

Pretraining in Actor-Critic Reinforcement Learning for Locomotion Deep reinforcement learning in a handful of trials using probabilistic dynamics models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T10:03:29.150862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:03:29.150862Z digest=sha256:996917c3b984c01b5d83086af42bddf6739d2539fcd47e72b6adc426cab1a5fc

Observation 16ba2e58-1971-48ef-8948-a52e7b83341a · outbound

This paper cites Efficient model-based reinforcement learning through optimistic policy search and planning.

Pretraining in Actor-Critic Reinforcement Learning for Locomotion Efficient model-based reinforcement learning through optimistic policy search and planning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T10:03:29.312751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:03:29.312751Z digest=sha256:d08eb4802e6f3450ccff2edecc1b0de91bba04d25aaaed5b009d867b4df4450a

Observation edaf0117-a842-4c2e-86fd-223adf5a3d99 · outbound

This paper cites BERT : Pre-training of deep bidirectional transformers for language understanding.

Pretraining in Actor-Critic Reinforcement Learning for Locomotion BERT : Pre-training of deep bidirectional transformers for language understanding

Reference 14

Resolution
malformed identifier
no resolver link, observed 2026-08-04T10:03:29.476493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:03:29.476493Z digest=sha256:acd49edc61c4108b7b40fd940b530c2a93d9e700d2e0394ce2b9c9acf7740a47

Observation ac754f07-fbaa-4e61-a25c-35c30c927275 · outbound

This paper cites Roloma: Robust loco-manipulation for quadruped robots with arms.

Pretraining in Actor-Critic Reinforcement Learning for Locomotion Roloma: Robust loco-manipulation for quadruped robots with arms

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T10:03:29.586681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:03:29.586681Z digest=sha256:62bf5dac5110878f48310023494f8b18a753527c7ae574dbc9026f76bab8ab26

Observation b5b79247-7cdd-428f-a2c6-7c0e617fe6c7 · outbound

This paper cites Delving deep into rectifiers: Surpassing human-level performance on imagenet classification.

Pretraining in Actor-Critic Reinforcement Learning for Locomotion Delving deep into rectifiers: Surpassing human-level performance on imagenet classification

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T10:03:29.667876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:03:29.667876Z digest=sha256:792b2cab770ddde66c508486e22588cc1cdb5e38f713f621ed3e2c76426c9563

Observation 41482c13-611f-48bf-a375-14c43c80edf8 · outbound

This paper cites Masked autoencoders are scalable vision learners.

Pretraining in Actor-Critic Reinforcement Learning for Locomotion Masked autoencoders are scalable vision learners

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T10:03:29.744272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:03:29.744272Z digest=sha256:b139398672e8459eef0a9f55fc23b3cfcf6c229d261226ce8de2e04d24b10b5c

Observation 5594a5f6-a13f-4886-a309-84a9f27e46e1 · outbound

This paper cites ANYmal Parkour : Learning Agile Navigation for Quadrupedal Robots , 2023.

Pretraining in Actor-Critic Reinforcement Learning for Locomotion ANYmal Parkour : Learning Agile Navigation for Quadrupedal Robots , 2023

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T10:03:29.849265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:03:29.849265Z digest=sha256:7d2c5ed1548a1cd8caeaa059f8ca89baad965cef6d4e68cf6875d91ca22f53b0

Observation 6340f53c-ba44-446c-9a66-8a8d01786479 · outbound

This paper cites Dario Bellicoso, Vassilios Tsounis, Jemin Hwangbo, Karen Bodie, Peter Fankhauser, Michael Bloesch, Remo Diethelm, Samuel Bachmann, Amir Melzer, and Mark Hoepflinger.

Pretraining in Actor-Critic Reinforcement Learning for Locomotion Dario Bellicoso, Vassilios Tsounis, Jemin Hwangbo, Karen Bodie, Peter Fankhauser, Michael Bloesch, Remo Diethelm, Samuel Bachmann, Amir Melzer, and Mark Hoepflinger

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T10:03:29.953190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:03:29.953190Z digest=sha256:7e9253979b3b11d38adc06faca7d3956bfad0add1f8d8a4f02a23bd3743bf265

Observation 669f2f07-67dd-4779-bbc0-846e3ae63956 · outbound

This paper cites Learning agile and dynamic motor skills for legged robots.

Pretraining in Actor-Critic Reinforcement Learning for Locomotion Learning agile and dynamic motor skills for legged robots

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T10:03:30.089031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:03:30.089031Z digest=sha256:fb99616d28ded3e4a56d61eec6e28b5df20427117cfcca46bebbd3e4d025c94a

Observation 10c998ed-53b6-4904-8ad3-5c7c84b0ce88 · outbound

This paper cites Bellman Eluder Dimension: New Rich Classes of RL Problems, and Sample-Efficient Algorithms.

Pretraining in Actor-Critic Reinforcement Learning for Locomotion Bellman Eluder Dimension: New Rich Classes of RL Problems, and Sample-Efficient Algorithms

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-04T10:03:30.180905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:03:30.180905Z digest=sha256:7c031ebafb65e571af0cf0f36d10b53547992090a977edd5a9b65000e4f1fce6

Observation 08439231-3013-4d01-9fd5-9dd36269bf6b · outbound

This paper cites The Role of Domain Randomization in Training Diffusion Policies for Whole-Body Humanoid Control.

Pretraining in Actor-Critic Reinforcement Learning for Locomotion The Role of Domain Randomization in Training Diffusion Policies for Whole-Body Humanoid Control

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T10:03:30.308293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:03:30.308293Z digest=sha256:21a2cadf09d92e38ed07293dbaa83dd6818b015eb5b527232dc41920fbd4a440

Observation 74fc8085-0de7-424d-8198-0aedeeb842ca · outbound

This paper cites Adam: A Method for Stochastic Optimization.

Pretraining in Actor-Critic Reinforcement Learning for Locomotion Adam: A Method for Stochastic Optimization

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T10:03:30.394974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:03:30.394974Z digest=sha256:ce4c137699645025f31c7fc073c40d075d8e5a632bb4abd3a7674ca565ae8aa1

Observation ecfe473e-fcf0-4964-8b31-caf6109d252e · outbound

This paper cites Actor-critic algorithms.

Pretraining in Actor-Critic Reinforcement Learning for Locomotion Actor-critic algorithms

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T10:03:30.428840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:03:30.428840Z digest=sha256:f469b74dae7791924c176e683daa9fc40afa6f144cb62f29dcd2ddc3cbd9a09c

Observation 8f01d23c-e5f1-4e4e-b97b-61dd58461855 · outbound

This paper cites Learning quadrupedal locomotion over challenging terrain.

Pretraining in Actor-Critic Reinforcement Learning for Locomotion Learning quadrupedal locomotion over challenging terrain

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-04T10:03:30.508818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:03:30.508818Z digest=sha256:0c9ff8d683486cab129ebf4b13ef2fe7bd53e667a6b3ca5e8f66767756285885

Observation afc4d0f2-b871-4947-b693-7132bc5c1050 · outbound

This paper cites Learning to Walk from Three Minutes of Real - World Data with Semi -structured Dynamics Models , 2024.

Pretraining in Actor-Critic Reinforcement Learning for Locomotion Learning to Walk from Three Minutes of Real - World Data with Semi -structured Dynamics Models , 2024

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T10:03:30.562093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:03:30.562093Z digest=sha256:a8c534d19055512068f40c153318f2dd4e53f25964ff36ef7358ca2d3de275c7

Observation 4a4d8fb1-e292-4dc3-925e-709f08923fa9 · outbound

This paper cites A Survey : Learning Embodied Intelligence from Physical Simulators and World Models , 2025.

Pretraining in Actor-Critic Reinforcement Learning for Locomotion A Survey : Learning Embodied Intelligence from Physical Simulators and World Models , 2025

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-04T10:03:30.635467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:03:30.635467Z digest=sha256:f457632029b9ea5d02743d0a5cc7830e73048ba35acbc5914e5de6a1391eb5fc

Observation af29eafa-30a8-482a-a7a8-4fa71a91ed1b · outbound

This paper cites Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks.

Pretraining in Actor-Critic Reinforcement Learning for Locomotion Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-04T10:03:30.714843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:03:30.714843Z digest=sha256:cc84e27c45c0ba42f6a80486fcff6411bf78ce806ce63e0d29cfa9016a329279

Observation 4fdc2f0b-89a6-46b4-bf0c-d451a8c7dd29 · outbound

This paper cites Combining physics and deep learning to learn continuous-time dynamics models.

Pretraining in Actor-Critic Reinforcement Learning for Locomotion Combining physics and deep learning to learn continuous-time dynamics models

Reference 29

Resolution
verified exact
doi, observed 2026-08-04T10:08:39.069158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-04T10:03:30.778393Z digest=sha256:4efba20f0b03ee1fbce3c58515645885ce007aa995c73d16e8fa85876bcfe968

Observation 3b3c0a4f-6c29-4485-a882-771da65cbd07 · outbound

This paper cites UMAP: Uniform Manifold Approximation and Projection for Dimension Reduction.

Pretraining in Actor-Critic Reinforcement Learning for Locomotion UMAP: Uniform Manifold Approximation and Projection for Dimension Reduction

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-04T10:03:30.827491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:03:30.827491Z digest=sha256:cfaa8203c03cb3f0f6c3c042a8e5c6788e3b57765d72f72f10994157d15f3017

Observation 51032244-5508-49df-a6f8-c77b79d71d82 · outbound

This paper cites Learning robust perceptive locomotion for quadrupedal robots in the wild.

Pretraining in Actor-Critic Reinforcement Learning for Locomotion Learning robust perceptive locomotion for quadrupedal robots in the wild

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-04T10:03:30.921104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:03:30.921104Z digest=sha256:d30cb376c782a4f1e31a0146be7e0334d6541dfcfb23861446d89e19a7414dce

Observation 5f1577d2-c062-4549-b5ca-1eb502819c71 · outbound

This paper cites Orbit: A unified simulation framework for interactive robot learning environments.

Pretraining in Actor-Critic Reinforcement Learning for Locomotion Orbit: A unified simulation framework for interactive robot learning environments

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-04T10:03:30.967138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:03:30.967138Z digest=sha256:ca8859257d91b5e8cbe4b9fd0d1543aae630ca3b2ab486b49630aa7aa9ba2c77

Observation 9d1ddd80-807e-42f9-8996-cee712795bd9 · outbound

This paper cites Symmetry considerations for learning task symmetric robot policies.

Pretraining in Actor-Critic Reinforcement Learning for Locomotion Symmetry considerations for learning task symmetric robot policies

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-04T10:03:31.029123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:03:31.029123Z digest=sha256:e28879d21510a58f5ad9f05f4fb5e778849db43ca4b39f4f8c53f5f0e0b74a08

Observation 3b8908e4-05e2-4906-b8bb-6f0a35fb9460 · outbound

This paper cites Murphy, Benjamin J.

Pretraining in Actor-Critic Reinforcement Learning for Locomotion Murphy, Benjamin J

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-04T10:03:31.051659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:03:31.051659Z digest=sha256:6a5db0ee8e81c062bd57cfb2e5759dd6a8ed266d5f453a730242c3a1f047304a

Observation e64f4e7d-64d3-46be-9ee8-a7e3687320a1 · outbound

This paper cites Information-Directed Exploration for Deep Reinforcement Learning.

Pretraining in Actor-Critic Reinforcement Learning for Locomotion Information-Directed Exploration for Deep Reinforcement Learning

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-04T10:03:31.128989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:03:31.128989Z digest=sha256:b5ca174bd47de0278ca2098a1125159d069e19d420a9267f4e37ef06e81ddba8

Observation ef271075-0b1c-45af-9c5d-b0198420bb88 · outbound

This paper cites AMP: Adversarial Motion Priors for Stylized Physics-Based Character Control.

Pretraining in Actor-Critic Reinforcement Learning for Locomotion AMP: Adversarial Motion Priors for Stylized Physics-Based Character Control

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-04T10:03:31.288541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:03:31.288541Z digest=sha256:eea5149ad33507239c4b3eeadebd56f060bb52161e7d4da1f78a14811a995e24

Observation 064f7195-f83e-4fd7-baea-b83735b3e5ca · outbound

This paper cites ASE : large-scale reusable adversarial skill embeddings for physically simulated characters.

Pretraining in Actor-Critic Reinforcement Learning for Locomotion ASE : large-scale reusable adversarial skill embeddings for physically simulated characters

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-04T10:03:31.431057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:03:31.431057Z digest=sha256:466f42ef39827b041a8c862af176178e6dbf59767f8dfafaa462598f64eecf4b

Observation 7746778e-0817-4026-9edc-9cf8c71397ca · outbound

This paper cites Whole-body end-effector pose tracking, 2024.

Pretraining in Actor-Critic Reinforcement Learning for Locomotion Whole-body end-effector pose tracking, 2024

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-04T10:03:31.553721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:03:31.553721Z digest=sha256:4132fa57d152ffa5cac6708f171d18a15f9ba02dfe569432f7a4be670261ec90

Observation ccd69885-dcd6-4e78-bd2f-9fc122b88fd0 · outbound

This paper cites Language models are unsupervised multitask learners.

Pretraining in Actor-Critic Reinforcement Learning for Locomotion Language models are unsupervised multitask learners

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-04T10:03:31.707213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:03:31.707213Z digest=sha256:1db32beecdd9b3cd799c366b9f8c15afb538e7c2bb70edb8a57ee935857ca484

Observation 6062ef33-4f90-4afc-9db4-82b0d4a3a133 · outbound

This paper cites Learning to walk in minutes using massively parallel deep reinforcement learning.

Pretraining in Actor-Critic Reinforcement Learning for Locomotion Learning to walk in minutes using massively parallel deep reinforcement learning

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-04T10:03:31.866986Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:03:31.866986Z digest=sha256:df191790da5155f9202cccc0356a47a31c631edfe105762644c6a11bcda3f899

Observation 808ac16b-76dd-43dc-9d59-e05db3c19195 · outbound

This paper cites Parkour in the Wild : Learning a General and Extensible Agile Locomotion Policy Using Multi -expert Distillation and RL Fine -tuning, 2025.

Pretraining in Actor-Critic Reinforcement Learning for Locomotion Parkour in the Wild : Learning a General and Extensible Agile Locomotion Policy Using Multi -expert Distillation and RL Fine -tuning, 2025

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-04T10:03:31.980649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:03:31.980649Z digest=sha256:51936f98c94eea9cd5467d9f5871b13a5d9fbd7161711df7a8a53107491c7add

Observation df215aec-7d86-4320-884b-0e5eb467c7ef · outbound

This paper cites Proximal Policy Optimization Algorithms.

Pretraining in Actor-Critic Reinforcement Learning for Locomotion Proximal Policy Optimization Algorithms

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-04T10:03:32.155903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:03:32.155903Z digest=sha256:79e1463574b38a7d9c37cdb226f5098a07f2cecdc05be45ab087ff403b8b539b

Observation 22d664b8-f072-4b50-acea-3046f2fa3812 · outbound

This paper cites RSL-RL: A Learning Library for Robotics Research.

Pretraining in Actor-Critic Reinforcement Learning for Locomotion RSL-RL: A Learning Library for Robotics Research

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-04T10:03:32.299059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:03:32.299059Z digest=sha256:9649a78c7bbe2e498e54d2392c5cdd2f0758e2a92f4c97a2de7a77bc297f8b13

Observation d75942fd-dba1-42da-890a-104e6d089818 · outbound

This paper cites Devon Hjelm, Philip Bachman, and Aaron C.

Pretraining in Actor-Critic Reinforcement Learning for Locomotion Devon Hjelm, Philip Bachman, and Aaron C

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-04T10:03:32.469853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:03:32.469853Z digest=sha256:bbd620183f5863298a0edbf2bacadf0cdbdf4ea69b9e8ba8328e5d7f425fe223

Observation 635783d2-0bd9-4137-9b8e-e962ae853a9d · outbound

This paper cites Planning to explore via self-supervised world models.

Pretraining in Actor-Critic Reinforcement Learning for Locomotion Planning to explore via self-supervised world models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-04T10:03:32.560939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:03:32.560939Z digest=sha256:c37163128f998d538110ec6d5febb3e69371e37c185985ba9ac12eccc73518c7

Observation dd9f9ead-8bea-47ce-bb5c-4210fca988c7 · outbound

This paper cites Blind Bipedal Stair Traversal via Sim-to-Real Reinforcement Learning.

Pretraining in Actor-Critic Reinforcement Learning for Locomotion Blind Bipedal Stair Traversal via Sim-to-Real Reinforcement Learning

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-04T10:03:32.644933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:03:32.644933Z digest=sha256:e94c858d271fce2b9d583c4f49631ce55c95843ff63c0666a94df8a33e85e73e

Observation 0c2add18-d72e-4490-8f31-36d43eb586a7 · outbound

This paper cites A Unified MPC Framework for Whole-Body Dynamic Locomotion and Manipulation.

Pretraining in Actor-Critic Reinforcement Learning for Locomotion A Unified MPC Framework for Whole-Body Dynamic Locomotion and Manipulation

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-04T10:03:32.760620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:03:32.760620Z digest=sha256:410ac4e98f94a73dc821f9dd483da022d5273f5932b2c0a45801506cfa10f533

Observation 556918e7-efe7-47be-b5b0-a76e796d2ce2 · outbound

This paper cites Guided Reinforcement Learning for Robust Multi - Contact Loco - Manipulation , 2024.

Pretraining in Actor-Critic Reinforcement Learning for Locomotion Guided Reinforcement Learning for Robust Multi - Contact Loco - Manipulation , 2024

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-04T10:03:32.828377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:03:32.828377Z digest=sha256:4adecf5bdd45d7d17397649a34420509b14f4a675511d76dbd155d5b12125757

Observation b55a044d-5d02-447e-9d50-e8e07a1bb76e · outbound

This paper cites Perceptive Pedipulation with Local Obstacle Avoidance , 2024.

Pretraining in Actor-Critic Reinforcement Learning for Locomotion Perceptive Pedipulation with Local Obstacle Avoidance , 2024

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-04T10:03:32.859664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:03:32.859664Z digest=sha256:3bfc1a8b4250dcb0e348a6dd29fb31db5c054615355b728cfeab36c8171014e1

Observation 16b6cdf5-1ab6-4da4-9f48-d6a3ca6eddea · outbound

This paper cites Gemini Robotics: Bringing AI into the Physical World.

Pretraining in Actor-Critic Reinforcement Learning for Locomotion Gemini Robotics: Bringing AI into the Physical World

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-04T10:03:32.943196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:03:32.943196Z digest=sha256:bad2d444fc781e5f24b7c9ef5f5abd66b4e24671804fdf01fa7b5bc3fd1a4046

Observation 729cd458-865d-47fd-8b9e-db05840368d5 · outbound

This paper cites Octo: An Open-Source Generalist Robot Policy.

Pretraining in Actor-Critic Reinforcement Learning for Locomotion Octo: An Open-Source Generalist Robot Policy

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-04T10:03:32.996524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:03:32.996524Z digest=sha256:c8dc34e6cc946099279fb0aa09836503c896b52e3cbb316fc6be3a4d589cb5d8

Observation 7c9ac8a9-0987-4798-ba2b-c9eea7696d2b · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Pretraining in Actor-Critic Reinforcement Learning for Locomotion Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-04T10:03:33.116502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:03:33.116502Z digest=sha256:600109372fac63829c9ea2c192b8d420fc0f2bc7107c40e2da9cb1684e858b4c

Observation 67680e13-674a-45b2-8344-3c1051849caf · outbound

This paper cites Advanced Skills through Multiple Adversarial Motion Priors in Reinforcement Learning.

Pretraining in Actor-Critic Reinforcement Learning for Locomotion Advanced Skills through Multiple Adversarial Motion Priors in Reinforcement Learning

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-04T10:03:33.186830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:03:33.186830Z digest=sha256:cc677913ee7f6059098d9bd5296ce67dc56e41782fe1fbb7e1b379a0901a053f

Observation 09a85d7d-c0f3-4016-8ec8-9ef83cfdd9a4 · outbound

This paper cites Pretraining in Deep Reinforcement Learning : A Survey , 2022.

Pretraining in Actor-Critic Reinforcement Learning for Locomotion Pretraining in Deep Reinforcement Learning : A Survey , 2022

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-04T10:03:33.336742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:03:33.336742Z digest=sha256:0cdc865f98f98053d1f69a0e5eb14434569dac5922d53a604f16043b59d76a00

Observation 2cadf785-2002-4825-9e71-e5fb6d831ca3 · outbound

This paper cites Neural Robot Dynamics.

Pretraining in Actor-Critic Reinforcement Learning for Locomotion Neural Robot Dynamics

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-04T10:03:33.453741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:03:33.453741Z digest=sha256:d967b4d99dba497c34655e3d9a31ed2c90cbd9f2f38ac0c3c3c533b74276285a

Observation 0720e17b-93fc-47e2-baf5-2e48e902188f · outbound

This paper cites Neural Volumetric Memory for Visual Locomotion Control.

Pretraining in Actor-Critic Reinforcement Learning for Locomotion Neural Volumetric Memory for Visual Locomotion Control

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-04T10:03:33.537036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:03:33.537036Z digest=sha256:c5c359c267e14031e0a08c77caef83faadd325a9e0f90acd7cba595900a3e163

Observation 38f54183-f9cb-48f4-a584-dbf9ffd60c9a · outbound

This paper cites Distillation-PPO: A Novel Two-Stage Reinforcement Learning Framework for Humanoid Robot Perceptive Locomotion.

Pretraining in Actor-Critic Reinforcement Learning for Locomotion Distillation-PPO: A Novel Two-Stage Reinforcement Learning Framework for Humanoid Robot Perceptive Locomotion

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-04T10:03:33.681668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:03:33.681668Z digest=sha256:120b4181784fb77069d5a41179ad7c8d446f2e76edf8bb3082a17b0dd1c73da5

Observation 9732485e-cbf4-4a40-8b4f-fd73d5922e81 · outbound

This paper cites Intention- Conditioned Flow Occupancy Models , 2025.

Pretraining in Actor-Critic Reinforcement Learning for Locomotion Intention- Conditioned Flow Occupancy Models , 2025

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-04T10:03:33.824504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:03:33.824504Z digest=sha256:1bb0dfaa9f09f61af37a80c91f94b363d7807e9dd0645f7187446635b44f99bf

Observation 671711d3-e693-4cf3-ab56-581b6eff25de · outbound

This paper cites write newline.

Pretraining in Actor-Critic Reinforcement Learning for Locomotion write newline

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-04T10:03:33.977983Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:03:33.977983Z digest=sha256:abad35d1ae3845e623137545dc3d528ceed66ee71293fd1eb3de2e2021faf260

Observation f6fdb2e7-a1c5-4abf-af22-05565b03b945 · outbound

This paper cites @esa (Ref.

Pretraining in Actor-Critic Reinforcement Learning for Locomotion @esa (Ref

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-04T10:03:34.095585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:03:34.095585Z digest=sha256:cf67695aae01d6bdb6da97a8d3a1c9e4e3e1df710754be44d33a3bc2cd6164d6

Observation c2278c11-d25d-4377-bf2c-d4b6aae46917 · outbound

This paper cites an unresolved cited work.

Pretraining in Actor-Critic Reinforcement Learning for Locomotion Unresolved cited work

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-04T10:03:34.173302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:03:34.173302Z digest=sha256:cfb94834d69828aa535491d9750539e235f34eb7be8dfdc90925a9b0abba87cf

Observation 316195b7-0025-4aaf-ad1a-4d3106a29923 · outbound

This paper cites an unresolved cited work.

Pretraining in Actor-Critic Reinforcement Learning for Locomotion Unresolved cited work

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-04T10:03:34.319088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:03:34.319088Z digest=sha256:b8c8c4bd463833cb796f021c187a5ab96a954b5230a4d5f68401f0b26576475e

Pith citing papers

No inbound Pith citation observations are available.