Pith. sign in

Paper Citation Record · LEDGER

Efficient Real-World Online Reinforcement Learning for Robot Manipulation via Centralized Training and Critic Decomposition

As of 15 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 0 inbound Pith citation observations for arXiv:2608.09762.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.09762 v1

Coverage vector

measured 31 of 31 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T11:20:20.073159Z

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

31 of 31 outbound references displayed

  • verified exact2
  • verified fuzzy9
  • unresolved20
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 7953fa82-b810-499b-bb1c-efcf13a82c01 · outbound

This paper cites A review of learning-based dynamics models for robotic manipulation |Science Robotics.

Efficient Real-World Online Reinforcement Learning for Robot Manipulation via Centralized Training and Critic Decomposition A review of learning-based dynamics models for robotic manipulation |Science Robotics

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T11:20:19.874750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:20:19.874750Z digest=sha256:790718bb380ce65d695b3ea4150274e041b3bc1b40624f1f35eda55fbedbf695

Observation 4f175a95-4231-4d2c-8106-761b11c7705a · outbound

This paper cites Learning Con- tinuous Control Actions for Robotic Grasping with Reinforcement Learning,.

Efficient Real-World Online Reinforcement Learning for Robot Manipulation via Centralized Training and Critic Decomposition Learning Con- tinuous Control Actions for Robotic Grasping with Reinforcement Learning,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:20:21.194190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T11:20:19.881716Z digest=sha256:5f4cca05bbdf76e37ef01f0e0682a36ebc3693360cdfb0ad40ad0a33e24a68a8

Observation d0cd25a3-7530-46ae-af01-50fc90390d03 · outbound

This paper cites General- Purpose Sim2Real Protocol for Learning Contact-Rich Manipulation With Marker-Based Visuotactile Sensors,.

Efficient Real-World Online Reinforcement Learning for Robot Manipulation via Centralized Training and Critic Decomposition General- Purpose Sim2Real Protocol for Learning Contact-Rich Manipulation With Marker-Based Visuotactile Sensors,

Reference 3

Resolution
verified exact
raw_fallback, observed 2026-08-11T11:20:20.852697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T11:20:19.889277Z digest=sha256:6d1209420f3b94a247548f667443e2cf736e6aa379f816666ebe58149d64cd29

Observation 3eac9bce-f94b-4229-a64e-e3e40b92467d · outbound

This paper cites Towards bridging the gap: Systematic sim-to-real transfer for diverse legged robots.

Efficient Real-World Online Reinforcement Learning for Robot Manipulation via Centralized Training and Critic Decomposition Towards bridging the gap: Systematic sim-to-real transfer for diverse legged robots

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T11:20:19.895124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:20:19.895124Z digest=sha256:fd31d88108e073057590e2fffcbfdb7861b23dbc9f54a01feb28347401e0b3d5

Observation 13405272-286b-4dc9-b277-7b1d07376018 · outbound

This paper cites Srl-vic: A variable stiffness-based safe reinforcement learning for contact-rich robotic tasks,.

Efficient Real-World Online Reinforcement Learning for Robot Manipulation via Centralized Training and Critic Decomposition Srl-vic: A variable stiffness-based safe reinforcement learning for contact-rich robotic tasks,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:20:21.160147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T11:20:19.901989Z digest=sha256:c5c99a64aa6fbd11e438932e75de7ab4121fc3c01c42f0422d9fad0d16aee8fa

Observation 95e6ccef-cb1e-475c-8382-cf36fc8ff216 · outbound

This paper cites Deep reinforcement learning for robotics: A survey of real-world successes,.

Efficient Real-World Online Reinforcement Learning for Robot Manipulation via Centralized Training and Critic Decomposition Deep reinforcement learning for robotics: A survey of real-world successes,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T11:20:19.907467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:20:19.907467Z digest=sha256:5b5b47ade76ecb74c48d2a9cfea70428394c10447903b9dd6dbad65617501909

Observation ee437a5f-e70d-4f2a-be11-d2916a6e6eee · outbound

This paper cites Improving vision-language-action model with online reinforcement learning,.

Efficient Real-World Online Reinforcement Learning for Robot Manipulation via Centralized Training and Critic Decomposition Improving vision-language-action model with online reinforcement learning,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:20:21.121932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T11:20:19.913331Z digest=sha256:4c3ee595a18952b8ee6c634dfcd97b31d9b2672bf50990a9145ea22c9f67509a

Observation ae2eab6f-486a-4215-bef2-fd81a99c225d · outbound

This paper cites Hi-WM: Human-in-the-World-Model for Scalable Robot Post-Training.

Efficient Real-World Online Reinforcement Learning for Robot Manipulation via Centralized Training and Critic Decomposition Hi-WM: Human-in-the-World-Model for Scalable Robot Post-Training

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T11:20:19.919454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:20:19.919454Z digest=sha256:c77e2578b84054340c17825fd8aab50ab11329036b6cb230c177e74d1fa62ac5

Observation 2bab35a4-89e1-4c86-b096-dcf22b84ad41 · outbound

This paper cites Rl-100: Performant robotic manipulation with real-world reinforcement learning,.

Efficient Real-World Online Reinforcement Learning for Robot Manipulation via Centralized Training and Critic Decomposition Rl-100: Performant robotic manipulation with real-world reinforcement learning,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T11:20:19.924795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:20:19.924795Z digest=sha256:b8a15c51fcf4bbe9f66491f7488736cc5208151c719cfa762f9ec5b7a2cfd16d

Observation 96e95e0a-14ed-41b8-afe7-ff5c1cc4a317 · outbound

This paper cites Rlinf-user: A unified and extensible system for real-world online policy learning in embodied ai,.

Efficient Real-World Online Reinforcement Learning for Robot Manipulation via Centralized Training and Critic Decomposition Rlinf-user: A unified and extensible system for real-world online policy learning in embodied ai,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T11:20:19.930023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:20:19.930023Z digest=sha256:b1da0e616e6b7ecba7120f59d30c14d0bcb4fae2f5126a04bfd2345a6b30e962

Observation dca26e27-0d33-4020-bab7-77c33600fa49 · outbound

This paper cites Towards Autonomous Reinforcement Learning for Real-World Robotic Manipulation With Large Language Models,.

Efficient Real-World Online Reinforcement Learning for Robot Manipulation via Centralized Training and Critic Decomposition Towards Autonomous Reinforcement Learning for Real-World Robotic Manipulation With Large Language Models,

Reference 11

Resolution
verified exact
raw_fallback, observed 2026-08-11T11:20:20.456064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T11:20:19.935620Z digest=sha256:7ea4738bb43737ae819a6505b83c9941ff0ba24d839264e7cda44ef90a9c9fd1

Observation d314379b-deee-43a2-aa96-5c385f267da2 · outbound

This paper cites Serl: A software suite for sample- efficient robotic reinforcement learning,.

Efficient Real-World Online Reinforcement Learning for Robot Manipulation via Centralized Training and Critic Decomposition Serl: A software suite for sample- efficient robotic reinforcement learning,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T11:20:19.942026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:20:19.942026Z digest=sha256:f303653c9a575cf86b88248b9f19861047ffb9371e0008fd402f7de63b47ff43

Observation b6fd6e15-207f-490b-b2ad-28e8b8d3db58 · outbound

This paper cites Precise and dexterous robotic manipulation via human-in-the-loop reinforcement learning,.

Efficient Real-World Online Reinforcement Learning for Robot Manipulation via Centralized Training and Critic Decomposition Precise and dexterous robotic manipulation via human-in-the-loop reinforcement learning,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T11:20:19.947525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:20:19.947525Z digest=sha256:c3f204548f3eccb771438d5b72ca46effe9cc207f4f763a27879d764c75433c0

Observation 8fe47e55-fee9-49fc-b126-586f27ee244d · outbound

This paper cites VLA-RL: Towards Masterful and General Robotic Manipulation with Scalable Reinforcement Learning.

Efficient Real-World Online Reinforcement Learning for Robot Manipulation via Centralized Training and Critic Decomposition VLA-RL: Towards Masterful and General Robotic Manipulation with Scalable Reinforcement Learning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T11:20:19.953071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:20:19.953071Z digest=sha256:3b9d8fa90029fbc3ff56546b5a1ed4ee3acae9b7e20ae15a2f16115f1a90b38f

Observation 22e82f08-1792-4285-9c83-703042f22bcf · outbound

This paper cites Reinflow: Fine-tuning flow matching policy with online reinforcement learning,.

Efficient Real-World Online Reinforcement Learning for Robot Manipulation via Centralized Training and Critic Decomposition Reinflow: Fine-tuning flow matching policy with online reinforcement learning,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:20:21.073542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T11:20:19.958730Z digest=sha256:cf2b0cbc20454063b2f16223301fc8cad6558e7261b86ea71802efe2736a20be

Observation 61f3d213-e333-428e-8051-c675de64a402 · outbound

This paper cites Hybrid reward architecture for reinforcement learning,.

Efficient Real-World Online Reinforcement Learning for Robot Manipulation via Centralized Training and Critic Decomposition Hybrid reward architecture for reinforcement learning,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:20:21.053010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T11:20:19.968287Z digest=sha256:cdd0dc53daf0cf6bee7ce74854d4b2cdf6203baebdb46a650572b3707a37ded7

Observation 0142220e-b7b5-49fc-8fe4-3080dc7f2d20 · outbound

This paper cites Efficient online reinforcement learning with offline data,.

Efficient Real-World Online Reinforcement Learning for Robot Manipulation via Centralized Training and Critic Decomposition Efficient online reinforcement learning with offline data,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T11:20:19.974107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:20:19.974107Z digest=sha256:098ed6ae0db436786cc1c24a4450cdc06e9256e21abf2b586a3a059f90508cc5

Observation df969530-88a0-45bf-8ae8-eb2b1f682f6d · outbound

This paper cites Deep Reinforcement Learning in Parameterized Action Space.

Efficient Real-World Online Reinforcement Learning for Robot Manipulation via Centralized Training and Critic Decomposition Deep Reinforcement Learning in Parameterized Action Space

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T11:20:19.980036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:20:19.980036Z digest=sha256:618dc4a02ae798e4b7bbdbcd0c2274e6f600455afd084edfcacc320bbb8cac24

Observation f8be1923-6991-4df6-864a-890db2122199 · outbound

This paper cites Parametrized Deep Q-Networks Learning: Reinforcement Learning with Discrete-Continuous Hybrid Action Space.

Efficient Real-World Online Reinforcement Learning for Robot Manipulation via Centralized Training and Critic Decomposition Parametrized Deep Q-Networks Learning: Reinforcement Learning with Discrete-Continuous Hybrid Action Space

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T11:20:19.985496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:20:19.985496Z digest=sha256:9a2aeab2b7099881f7681c8da03f24115f41feaeed734634d11dc95da8e7a31b

Observation fe4d146d-8b36-413d-9a1c-539d15bd0047 · outbound

This paper cites Multi-Pass Q-Networks for Deep Reinforcement Learning with Parameterised Action Spaces.

Efficient Real-World Online Reinforcement Learning for Robot Manipulation via Centralized Training and Critic Decomposition Multi-Pass Q-Networks for Deep Reinforcement Learning with Parameterised Action Spaces

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T11:20:19.998476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:20:19.998476Z digest=sha256:bb718abb719392dd9c035d69e1c1daccf7587c2fdddc6b25a9ff521e77113df3

Observation 8a078d8e-7ac0-4a00-8fc4-eb2c7f5db353 · outbound

This paper cites Multi-agent actor-critic for mixed cooperative-competitive environments,.

Efficient Real-World Online Reinforcement Learning for Robot Manipulation via Centralized Training and Critic Decomposition Multi-agent actor-critic for mixed cooperative-competitive environments,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T11:20:20.006679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:20:20.006679Z digest=sha256:eb2fc0f8402565b14cc45805b5d1b836434d5618fb36eeed15a8b85a8d384aea

Observation 36ffc7a6-71f5-4857-84c9-17b733237f35 · outbound

This paper cites The surprising effectiveness of ppo in cooperative multi-agent games,.

Efficient Real-World Online Reinforcement Learning for Robot Manipulation via Centralized Training and Critic Decomposition The surprising effectiveness of ppo in cooperative multi-agent games,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T11:20:20.012254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:20:20.012254Z digest=sha256:44099a49491e25c981a3fae7100f565cdcbe96a93990989eeb081193c63e4c8a

Observation 55b51ff4-6707-4e05-9489-6b5836aefa71 · outbound

This paper cites Asynchronous actor-critic for multi-agent reinforcement learning,.

Efficient Real-World Online Reinforcement Learning for Robot Manipulation via Centralized Training and Critic Decomposition Asynchronous actor-critic for multi-agent reinforcement learning,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:20:20.985986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T11:20:20.021539Z digest=sha256:02fe7bed1c0201ee4bb6fd093c8bc04d7a0d8aaaa4e6cf78d8c6b17743ba7e8f

Observation 5d5e5942-bb6d-414a-812b-bd687254c0a5 · outbound

This paper cites Action decoupled sac reinforcement learning with discrete-continuous hybrid action spaces,.

Efficient Real-World Online Reinforcement Learning for Robot Manipulation via Centralized Training and Critic Decomposition Action decoupled sac reinforcement learning with discrete-continuous hybrid action spaces,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:20:20.956233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T11:20:20.027981Z digest=sha256:b95a5f4f63242f213afb5887bbe59bced68713e79ea8f32826b0ba511ec26c60

Observation 83659ad8-fe57-4c94-b34d-b03c64bccc8a · outbound

This paper cites Effective multi-agent deep reinforcement learning control with relative entropy regularization,.

Efficient Real-World Online Reinforcement Learning for Robot Manipulation via Centralized Training and Critic Decomposition Effective multi-agent deep reinforcement learning control with relative entropy regularization,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:20:20.934523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T11:20:20.036886Z digest=sha256:205d5c0a82b8e288e352eb9e7517ae0c42e8cda1b82179ddb0eef2bba342f24a

Observation e5eeae7a-6621-4cf4-a573-156865f8c01e · outbound

This paper cites Soft Actor-Critic Algorithms and Applications.

Efficient Real-World Online Reinforcement Learning for Robot Manipulation via Centralized Training and Critic Decomposition Soft Actor-Critic Algorithms and Applications

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T11:20:20.042024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:20:20.042024Z digest=sha256:aec16389ca4c05fa24f760a25e6f8ee7c1e2979adc3dbf50afe3c450a5aef86e

Observation f2cabc70-02ae-44eb-b8ae-be067ec5d036 · outbound

This paper cites Deep residual learning for image recognition,.

Efficient Real-World Online Reinforcement Learning for Robot Manipulation via Centralized Training and Critic Decomposition Deep residual learning for image recognition,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T11:20:20.048651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:20:20.048651Z digest=sha256:25fd90ebf9c42f13fa3f718239e01f925b066492d7bb56d642ce85d35aba7857

Observation 5074cb71-2bf8-4341-9c92-132d71c9fe2e · outbound

This paper cites An Introduction to Centralized Training for Decentralized Execution in Cooperative Multi-Agent Reinforcement Learning.

Efficient Real-World Online Reinforcement Learning for Robot Manipulation via Centralized Training and Critic Decomposition An Introduction to Centralized Training for Decentralized Execution in Cooperative Multi-Agent Reinforcement Learning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T11:20:20.055396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:20:20.055396Z digest=sha256:01e82e98dd3f8fe27f7b21e09a91a46e974d59e4fc5fe0a6cc31fe7a12c190bf

Observation 641735f2-9b57-48e6-98be-efa0f519e538 · outbound

This paper cites Soft Actor-Critic for Discrete Action Settings.

Efficient Real-World Online Reinforcement Learning for Robot Manipulation via Centralized Training and Critic Decomposition Soft Actor-Critic for Discrete Action Settings

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T11:20:20.061953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:20:20.061953Z digest=sha256:7d877c8fdc8fe2d7d1189105784313e62f1de09d1de479b88ad9303e78a9facd

Observation c43875eb-ce6b-4761-a192-06de815cc0a2 · outbound

This paper cites A high-force gripper with embedded multimodal sensing for powerful and perception driven grasping,.

Efficient Real-World Online Reinforcement Learning for Robot Manipulation via Centralized Training and Critic Decomposition A high-force gripper with embedded multimodal sensing for powerful and perception driven grasping,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:20:20.894818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T11:20:20.067472Z digest=sha256:26ef3f6ae9a8dbec465c148147bd1c990298f81dbd9dfe99ab3ceb5a9fb16f37

Observation 80032bd8-1fab-403b-ad4b-09aa3173a7e7 · outbound

This paper cites Root mean square layer normalization,.

Efficient Real-World Online Reinforcement Learning for Robot Manipulation via Centralized Training and Critic Decomposition Root mean square layer normalization,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T11:20:20.073159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:20:20.073159Z digest=sha256:ebfa4482fe62d65ecbca752651232da02f24769fb42b3fc577b2a08db7e4faea

Pith citing papers

No inbound Pith citation observations are available.