Pith. sign in

Paper Citation Record · LEDGER

Towards Human-like Physical Intelligence: Lifelong Vision-Language-Action Learning for Robotic Manipulation

As of 9 August 2026, this Paper Citation Record lists 59 of 59 outbound references and 0 inbound Pith citation observations for arXiv:2607.14852.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.14852 v2

Coverage vector

measured 59 of 59 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-02T00:57:45.185859Z

measured 59 of 59 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

59 of 59 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved59
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6d4e5abc-5875-460f-9143-60c6687f9e05 · outbound

This paper cites Do As I Can, Not As I Say: Grounding Language in Robotic Affordances.

Towards Human-like Physical Intelligence: Lifelong Vision-Language-Action Learning for Robotic Manipulation Do As I Can, Not As I Say: Grounding Language in Robotic Affordances

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-02T00:57:38.361430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:57:38.361430Z digest=sha256:323431eba265f86c23a453dfc2a5f0b61c957984cc0d1b07a4974930a8f9240b

Observation 8da652d4-1b83-4fb1-b00e-29b2ac5e0a2e · outbound

This paper cites $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control.

Towards Human-like Physical Intelligence: Lifelong Vision-Language-Action Learning for Robotic Manipulation $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-02T00:57:38.446664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:57:38.446664Z digest=sha256:d876e8d26585ff273e0d5d002e16d0a5253fde4fd3659b5b3040ebe639761fcf

Observation 5ecd9849-3188-4f1e-872d-c057bfe1feaf · outbound

This paper cites RoboCat: A Self-Improving Generalist Agent for Robotic Manipulation.

Towards Human-like Physical Intelligence: Lifelong Vision-Language-Action Learning for Robotic Manipulation RoboCat: A Self-Improving Generalist Agent for Robotic Manipulation

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-02T00:57:38.564461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:57:38.564461Z digest=sha256:8009ef7912f8c42290dc9badb8284e701ab2221f4e8db40166b3a9bfca14bf55

Observation d23d7b75-13be-4f71-b49f-811a87159e0e · outbound

This paper cites RT-1: Robotics Transformer for Real-World Control at Scale.

Towards Human-like Physical Intelligence: Lifelong Vision-Language-Action Learning for Robotic Manipulation RT-1: Robotics Transformer for Real-World Control at Scale

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-02T00:57:38.675970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:57:38.675970Z digest=sha256:1dd88461516870c745f11d95e8e61a55a9a490d6d7d7a0e21618540c41917c42

Observation bfd3f953-bb97-48a6-8ad8-0565fa985ce2 · outbound

This paper cites RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control.

Towards Human-like Physical Intelligence: Lifelong Vision-Language-Action Learning for Robotic Manipulation RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-02T00:57:38.798291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:57:38.798291Z digest=sha256:8a2e48e5e26c86447e73a8942d63206e9e5d8da80ff5b521ef30f492f25d4195

Observation 8fa60eea-fa48-4433-b372-6b24d4ce6af2 · outbound

This paper cites Riemannian walk for incremental learning: Understanding forgetting and intransigence.

Towards Human-like Physical Intelligence: Lifelong Vision-Language-Action Learning for Robotic Manipulation Riemannian walk for incremental learning: Understanding forgetting and intransigence

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-02T00:57:38.868336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:57:38.868336Z digest=sha256:2e8b894029109cdcc6b945ef642955cbc52838b329e97996a33ea8e0b4b0156c

Observation ac24ec8c-df59-44f7-9361-66b1e17324df · outbound

This paper cites On Tiny Episodic Memories in Continual Learning.

Towards Human-like Physical Intelligence: Lifelong Vision-Language-Action Learning for Robotic Manipulation On Tiny Episodic Memories in Continual Learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-02T00:57:38.974751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:57:38.974751Z digest=sha256:4c0fadcfee921c3d98e7635ab802452722e0e652c37eb7b16ededd0d6c7ce679

Observation 97243f6c-be38-4b5a-a6fe-ca2c7dec90b3 · outbound

This paper cites Diffusion Policy: Visuomotor Policy Learning via Action Diffusion.

Towards Human-like Physical Intelligence: Lifelong Vision-Language-Action Learning for Robotic Manipulation Diffusion Policy: Visuomotor Policy Learning via Action Diffusion

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-02T00:57:39.140456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:57:39.140456Z digest=sha256:c9fc8e0e4da2be5ce67375d462c396ef95a28f8080060a63016676836e3aaf1a

Observation cf28c9fc-2b4b-4556-8395-f2bd1a2639dc · outbound

This paper cites OpenHelix: A Short Survey, Empirical Analysis, and Open-Source Dual-System VLA Model for Robotic Manipulation.

Towards Human-like Physical Intelligence: Lifelong Vision-Language-Action Learning for Robotic Manipulation OpenHelix: A Short Survey, Empirical Analysis, and Open-Source Dual-System VLA Model for Robotic Manipulation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-02T00:57:39.308136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:57:39.308136Z digest=sha256:81397f587549ebff9ce66c2d174856b002bf16b9509b463d97d56e1d0da07def

Observation 1a53177d-b354-436d-bfe3-879228996277 · outbound

This paper cites Loss of plasticity in deep continual learning.Nature, 632(8026):768–774, 2024.

Towards Human-like Physical Intelligence: Lifelong Vision-Language-Action Learning for Robotic Manipulation Loss of plasticity in deep continual learning.Nature, 632(8026):768–774, 2024

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-02T00:57:39.473600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:57:39.473600Z digest=sha256:d89bcc262645d3b92443bedeea33f79f1b54df816c6e8f75251a8488d8aba6f2

Observation 6c03f0c3-b7d1-4ecc-85d8-62fe8f30faf9 · outbound

This paper cites Palm-e: An embodied multimodal language model.

Towards Human-like Physical Intelligence: Lifelong Vision-Language-Action Learning for Robotic Manipulation Palm-e: An embodied multimodal language model

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-02T00:57:39.612805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:57:39.612805Z digest=sha256:ffa927cd945ac5aa7e1683180ab61187baeb6620d71908c34947b32b1bd21ebe

Observation 81402915-0aa4-4d73-9b3f-e82aec51676b · outbound

This paper cites Benchmarking Vision, Language, & Action Models in Procedurally Generated, Open Ended Action Environments.

Towards Human-like Physical Intelligence: Lifelong Vision-Language-Action Learning for Robotic Manipulation Benchmarking Vision, Language, & Action Models in Procedurally Generated, Open Ended Action Environments

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-02T00:57:39.775665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:57:39.775665Z digest=sha256:fccd7cbac6979d20b980f2ef2a3e2e3a2a9f037a40f2677b01f289164202a577

Observation 4ba87ad8-c25a-46fe-ae7d-0666ed9f4d4b · outbound

This paper cites Lora: Low-rank adaptation of large language models.

Towards Human-like Physical Intelligence: Lifelong Vision-Language-Action Learning for Robotic Manipulation Lora: Low-rank adaptation of large language models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T00:57:39.937349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:57:39.937349Z digest=sha256:8bdf61f72b3cb3e46330daa99471352a7cc968d32d3d29265d262a1b55fb0170

Observation e5a8d7e6-4f24-4cf7-b19c-973217c16924 · outbound

This paper cites VoxPoser: Composable 3D Value Maps for Robotic Manipulation with Language Models.

Towards Human-like Physical Intelligence: Lifelong Vision-Language-Action Learning for Robotic Manipulation VoxPoser: Composable 3D Value Maps for Robotic Manipulation with Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-02T00:57:40.106650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:57:40.106650Z digest=sha256:4434193bf273aebfd78f030ebabd18d0de3c32660c60767e0847b4976e15f736

Observation cae95d81-6b39-474d-a382-5de0ee2bbb6e · outbound

This paper cites $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization.

Towards Human-like Physical Intelligence: Lifelong Vision-Language-Action Learning for Robotic Manipulation $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-02T00:57:40.272666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:57:40.272666Z digest=sha256:259d4acb2f88499174a5e55c9851d815d186fc90ba25e57b4bb37a02b66b838a

Observation e7c4049a-4dde-4140-8e42-6f4f21c82016 · outbound

This paper cites VIMA: General Robot Manipulation with Multimodal Prompts.

Towards Human-like Physical Intelligence: Lifelong Vision-Language-Action Learning for Robotic Manipulation VIMA: General Robot Manipulation with Multimodal Prompts

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-02T00:57:40.434766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:57:40.434766Z digest=sha256:88a1d846c7c4ddb7c33de54b881f45e11e550fb7172a90b6a534933e8efc60c6

Observation 671d6ba0-d7df-499c-9274-8b1d699d88fd · outbound

This paper cites Srt-h: A hierarchical framework for autonomous surgery via language-conditioned imitation learning.Science robotics, 10(104):eadt5254, 2025.

Towards Human-like Physical Intelligence: Lifelong Vision-Language-Action Learning for Robotic Manipulation Srt-h: A hierarchical framework for autonomous surgery via language-conditioned imitation learning.Science robotics, 10(104):eadt5254, 2025

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-02T00:57:40.557073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:57:40.557073Z digest=sha256:27b7cc4ca945b114f79a001b64251e900e54d5eb947c527fddc14c0f550f38b7

Observation 17855fb2-1f2f-45cd-96f9-767e61ea49ca · outbound

This paper cites OpenVLA: An Open-Source Vision-Language-Action Model.

Towards Human-like Physical Intelligence: Lifelong Vision-Language-Action Learning for Robotic Manipulation OpenVLA: An Open-Source Vision-Language-Action Model

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-02T00:57:40.564504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:57:40.564504Z digest=sha256:ca8302a19034058a06c25d04181115924bf240f0b9c3bf1fc5b74be4a49ef37f

Observation 7f1747ac-9dad-42dc-b167-da8070caa109 · outbound

This paper cites Overcoming catastrophic forgetting in neural networks.Proceedings of the national academy of sciences, 114(13): 3521–3526, 2017.

Towards Human-like Physical Intelligence: Lifelong Vision-Language-Action Learning for Robotic Manipulation Overcoming catastrophic forgetting in neural networks.Proceedings of the national academy of sciences, 114(13): 3521–3526, 2017

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-02T00:57:40.626531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:57:40.626531Z digest=sha256:5355c379bb8e1dcfd9e9c5235ad1fc1fbd7f0298c26f7acf481c258319022a51

Observation 973ce0bb-7239-40ee-8772-a2cee4eca53c · outbound

This paper cites Reinforcement Fine-Tuning Naturally Mitigates Forgetting in Continual Post-Training.

Towards Human-like Physical Intelligence: Lifelong Vision-Language-Action Learning for Robotic Manipulation Reinforcement Fine-Tuning Naturally Mitigates Forgetting in Continual Post-Training

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-02T00:57:40.746075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:57:40.746075Z digest=sha256:dd73d3c15e1a896e76ddad18a6317cd9e62b83ea761965c6e7cbd23fe1ec57a5

Observation 920df1b8-d243-4890-b5e7-f1a1efd5aa4b · outbound

This paper cites The power of scale for parameter-efficient prompt tuning.

Towards Human-like Physical Intelligence: Lifelong Vision-Language-Action Learning for Robotic Manipulation The power of scale for parameter-efficient prompt tuning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-02T00:57:40.923216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:57:40.923216Z digest=sha256:85cc24b0a7640572e09cead7ed7f0f44763df92c3781b3821d82b1f26cfcc987

Observation 1c832505-4417-4d0b-ab31-5a3909d5fd35 · outbound

This paper cites Remem-vla: Empowering vision-language-action model with memory via dual-level recurrent queries.arXiv preprint arXiv:2603.12942, 2026.

Towards Human-like Physical Intelligence: Lifelong Vision-Language-Action Learning for Robotic Manipulation Remem-vla: Empowering vision-language-action model with memory via dual-level recurrent queries.arXiv preprint arXiv:2603.12942, 2026

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-02T00:57:41.063842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:57:41.063842Z digest=sha256:6dbce8f9b014aa4084f960d9ff28c72899afe1dacf3906b230472368c17c350f

Observation 1db9480a-222b-4401-a7ef-7963e83fa241 · outbound

This paper cites Prefix-tuning: Optimizing continuous prompts for generation.

Towards Human-like Physical Intelligence: Lifelong Vision-Language-Action Learning for Robotic Manipulation Prefix-tuning: Optimizing continuous prompts for generation

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-02T00:57:41.166042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:57:41.166042Z digest=sha256:f2f9d8c1ac81e5cb263b177ce7c0cfb2d42edad6bd08da7e2e813330e44024bb

Observation df862715-64fe-47c8-a6f2-2a3e0c02350e · outbound

This paper cites Learning without forgetting.IEEE transactions on pattern analysis and machine intelligence, 40(12):2935–2947, 2017.

Towards Human-like Physical Intelligence: Lifelong Vision-Language-Action Learning for Robotic Manipulation Learning without forgetting.IEEE transactions on pattern analysis and machine intelligence, 40(12):2935–2947, 2017

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-02T00:57:41.271781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:57:41.271781Z digest=sha256:a11e5d56beef7adef0644f5d18465b5fb3e60bdbbc23c773db1fd45dc59dd216

Observation 6d394b78-c309-43d3-aaa3-a0f41b8b07f1 · outbound

This paper cites Learning without forgetting.

Towards Human-like Physical Intelligence: Lifelong Vision-Language-Action Learning for Robotic Manipulation Learning without forgetting

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-02T00:57:41.376843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:57:41.376843Z digest=sha256:44851665d399f34474e42ed3ccbf2f09cc3e4425f8c60af1ed2b3d19b98e2f1e

Observation d083236b-4763-4028-b504-68e83c95dba4 · outbound

This paper cites Code as policies: Language model programs for embodied control.

Towards Human-like Physical Intelligence: Lifelong Vision-Language-Action Learning for Robotic Manipulation Code as policies: Language model programs for embodied control

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-02T00:57:41.519701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:57:41.519701Z digest=sha256:e8fa45effae2135748b904d8e14dc9e714cd54a4f47cbc200a07a13c7559abba

Observation 8a6204cf-c9a6-406a-81ed-fce26d6c0c00 · outbound

This paper cites Never-ending behavior-cloning agent for robotic manipulation.arXiv preprint arXiv:2403.00336, 2024.

Towards Human-like Physical Intelligence: Lifelong Vision-Language-Action Learning for Robotic Manipulation Never-ending behavior-cloning agent for robotic manipulation.arXiv preprint arXiv:2403.00336, 2024

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-02T00:57:41.679900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:57:41.679900Z digest=sha256:077f85b80718c2e09d37564709300fd7c93accc0bbe5b41aab520a087eebd8d0

Observation ae038217-681c-40b4-934c-8fbd01cfe4a2 · outbound

This paper cites Pixelvla: Advancing pixel-level understanding in vision-language-action model.arXiv preprint arXiv:2511.01571, 2025.

Towards Human-like Physical Intelligence: Lifelong Vision-Language-Action Learning for Robotic Manipulation Pixelvla: Advancing pixel-level understanding in vision-language-action model.arXiv preprint arXiv:2511.01571, 2025

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-02T00:57:41.790039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:57:41.790039Z digest=sha256:b2307567c519f20e0fee64309a60cdfc5b120699d89a6ffce13f5ab180f9af8e

Observation f4e12658-d377-4652-acaf-556b08b76456 · outbound

This paper cites Discrete Diffusion VLA: Bringing Discrete Diffusion to Action Decoding in Vision-Language-Action Policies.

Towards Human-like Physical Intelligence: Lifelong Vision-Language-Action Learning for Robotic Manipulation Discrete Diffusion VLA: Bringing Discrete Diffusion to Action Decoding in Vision-Language-Action Policies

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-02T00:57:41.909520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:57:41.909520Z digest=sha256:4f2210b503c8813cdef18743b157511e1d9183f3c1b1a05cbcdbe47d892b630a

Observation 75d14dfb-e1ee-489d-903a-193620f4e0f4 · outbound

This paper cites Showui: One vision-language-action model for gui visual agent.

Towards Human-like Physical Intelligence: Lifelong Vision-Language-Action Learning for Robotic Manipulation Showui: One vision-language-action model for gui visual agent

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-02T00:57:42.023492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:57:42.023492Z digest=sha256:dc81231345608528c35f823aefff3fff7eb573f99bc9394cb456f991a019abb6

Observation c9dda9e5-a346-4835-9596-a5462827eb2a · outbound

This paper cites Libero: Benchmarking knowledge transfer for lifelong robot learning.Advances in Neural Information Processing Systems, 36:44776–44791, 2023.

Towards Human-like Physical Intelligence: Lifelong Vision-Language-Action Learning for Robotic Manipulation Libero: Benchmarking knowledge transfer for lifelong robot learning.Advances in Neural Information Processing Systems, 36:44776–44791, 2023

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-02T00:57:42.163804Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:57:42.163804Z digest=sha256:c46cf68b27068175d4fdb709a31c858fb454a954ae1e4bfde2ef4c37a6ee4056

Observation 2861effb-1148-4105-b9e3-6352c2a770b6 · outbound

This paper cites Pretrained vision-language-action models are surprisingly resistant to forgetting in continual learning.arXiv preprint arXiv:2603.03818, 2026.

Towards Human-like Physical Intelligence: Lifelong Vision-Language-Action Learning for Robotic Manipulation Pretrained vision-language-action models are surprisingly resistant to forgetting in continual learning.arXiv preprint arXiv:2603.03818, 2026

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-02T00:57:42.309881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:57:42.309881Z digest=sha256:f0a9dd4dfcc3e16df315a8d1efb045846cb44430089618043a8f5009b40fba33

Observation ad0e677d-ee68-4ae8-b596-9ab348fca764 · outbound

This paper cites Spatial-Temporal Aware Visuomotor Diffusion Policy Learning.

Towards Human-like Physical Intelligence: Lifelong Vision-Language-Action Learning for Robotic Manipulation Spatial-Temporal Aware Visuomotor Diffusion Policy Learning

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-02T00:57:42.384758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:57:42.384758Z digest=sha256:556c1af7d82700d64ae96401711f935b25f64f3712e4747c09e634850f03dd7b

Observation 76a1a350-ba0f-4403-b458-12b73c023994 · outbound

This paper cites Packnet: Adding multiple tasks to a single network by iterative pruning.

Towards Human-like Physical Intelligence: Lifelong Vision-Language-Action Learning for Robotic Manipulation Packnet: Adding multiple tasks to a single network by iterative pruning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-02T00:57:42.499558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:57:42.499558Z digest=sha256:63d1b80078029a77c078e2d3c04584f5c38a083c71d7871e878071d76108a31d

Observation f859b398-a502-45a9-b0de-7c514b6607a0 · outbound

This paper cites Preserving and combining knowledge in robotic lifelong reinforcement learning.Nature Machine Intelligence, 7(2):256–269, 2025.

Towards Human-like Physical Intelligence: Lifelong Vision-Language-Action Learning for Robotic Manipulation Preserving and combining knowledge in robotic lifelong reinforcement learning.Nature Machine Intelligence, 7(2):256–269, 2025

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-02T00:57:42.643648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:57:42.643648Z digest=sha256:93ed34566d0c89606d65e61462b243999aa564f5b2574002096935a8e683b0d4

Observation e21013a9-57c6-4ceb-8be4-35d30bdc961b · outbound

This paper cites LLARVA: Vision-Action Instruction Tuning Enhances Robot Learning.

Towards Human-like Physical Intelligence: Lifelong Vision-Language-Action Learning for Robotic Manipulation LLARVA: Vision-Action Instruction Tuning Enhances Robot Learning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-02T00:57:42.768376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:57:42.768376Z digest=sha256:723ff5c150b03cca9dad23d090bff63208b1ebea3c5eebfd3d85b20aa9583e39

Observation b7e19fac-c9c1-4f25-801d-4be669680dc4 · outbound

This paper cites Open X-Embodiment: Robotic Learning Datasets and RT-X Models.

Towards Human-like Physical Intelligence: Lifelong Vision-Language-Action Learning for Robotic Manipulation Open X-Embodiment: Robotic Learning Datasets and RT-X Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-02T00:57:42.876938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:57:42.876938Z digest=sha256:36646a4981d8365231f62cd30d2fdb2ad78a547337cc70531775b72b5e8e370e

Observation 8e5de4bf-5370-457c-bf11-394968808e29 · outbound

This paper cites FAST: Efficient Action Tokenization for Vision-Language-Action Models.

Towards Human-like Physical Intelligence: Lifelong Vision-Language-Action Learning for Robotic Manipulation FAST: Efficient Action Tokenization for Vision-Language-Action Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-02T00:57:42.983225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:57:42.983225Z digest=sha256:299c8ccc044ab54dd8eb4cafad2044e16cf11bee9339bf0b655e28f213e427bb

Observation f093673d-4b63-4894-9d1d-2b3ffc21433c · outbound

This paper cites ${\pi}_{0.7}$: a Steerable Generalist Robotic Foundation Model with Emergent Capabilities.

Towards Human-like Physical Intelligence: Lifelong Vision-Language-Action Learning for Robotic Manipulation ${\pi}_{0.7}$: a Steerable Generalist Robotic Foundation Model with Emergent Capabilities

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-02T00:57:43.044566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:57:43.044566Z digest=sha256:bc0b6b47ecc3caf743bf1d07cf279947e79c969320ca0db167c42ab91f6e2cbf

Observation deffecab-8279-43d4-a60f-555ffe976a3b · outbound

This paper cites Incremental classifier and representation learning.

Towards Human-like Physical Intelligence: Lifelong Vision-Language-Action Learning for Robotic Manipulation Incremental classifier and representation learning

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-02T00:57:43.119826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:57:43.119826Z digest=sha256:c9bc4f52a419f8e7bb9b7f6962c96e1cbdd828173d03aec617bf812d72b03480

Observation 86645130-7d20-4dd4-838d-caa7f7cfebf8 · outbound

This paper cites A Generalist Agent.

Towards Human-like Physical Intelligence: Lifelong Vision-Language-Action Learning for Robotic Manipulation A Generalist Agent

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-02T00:57:43.200599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:57:43.200599Z digest=sha256:80de79b1a9144250d1a5fbcf4139244db0028a535035ce84be395b22cc7fa14a

Observation e9c286bb-d113-4d4a-bf70-b76e4fcdae19 · outbound

This paper cites Progressive Neural Networks.

Towards Human-like Physical Intelligence: Lifelong Vision-Language-Action Learning for Robotic Manipulation Progressive Neural Networks

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-02T00:57:43.302391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:57:43.302391Z digest=sha256:8f364e77572f8bc7ca7a6513594a9dacaed0ec39d281fcf1fefb3b78c5e5be39

Observation 1d130e34-7e7e-4643-96c8-fc4b7eb9824f · outbound

This paper cites RL's Razor: Why Online Reinforcement Learning Forgets Less.

Towards Human-like Physical Intelligence: Lifelong Vision-Language-Action Learning for Robotic Manipulation RL's Razor: Why Online Reinforcement Learning Forgets Less

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-02T00:57:43.368877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:57:43.368877Z digest=sha256:49623fa66681cadd6ea8419b57e882e4a7016a0a74b371b1816cbf274b4ae7b0

Observation 4091b805-2fee-4570-bc78-e314dad59af3 · outbound

This paper cites Cliport: What and where pathways for robotic manipulation.

Towards Human-like Physical Intelligence: Lifelong Vision-Language-Action Learning for Robotic Manipulation Cliport: What and where pathways for robotic manipulation

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-02T00:57:43.430502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:57:43.430502Z digest=sha256:8476f92807206c37a041cd3b171278396ae51f7e0fcb1b827c322a175ca728a7

Observation 30b3ec67-9175-4544-8a41-6fa4ed8db12b · outbound

This paper cites Perceiver-actor: A multi-task transformer for robotic manipulation.

Towards Human-like Physical Intelligence: Lifelong Vision-Language-Action Learning for Robotic Manipulation Perceiver-actor: A multi-task transformer for robotic manipulation

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-02T00:57:43.556355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:57:43.556355Z digest=sha256:dbbfb05c849e87671c18382c6cd372154d5e00641a86e690e45ac8b6c7b7fe6f

Observation 2b2a4d7c-00e1-47e2-a65c-e5d731c23ac4 · outbound

This paper cites SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics.

Towards Human-like Physical Intelligence: Lifelong Vision-Language-Action Learning for Robotic Manipulation SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-02T00:57:43.651170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:57:43.651170Z digest=sha256:8d3950f71950e01e21bf6a8c4832c8ccc2021e6e2fd0411502a76942865f4d6d

Observation 68528695-96c2-4fa0-a50c-627e6815c1ab · outbound

This paper cites Coda-prompt: Continual decomposed attention- based prompting for rehearsal-free continual learning.

Towards Human-like Physical Intelligence: Lifelong Vision-Language-Action Learning for Robotic Manipulation Coda-prompt: Continual decomposed attention- based prompting for rehearsal-free continual learning

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-02T00:57:43.723645Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:57:43.723645Z digest=sha256:862088a99349f330a160bd7869c59c84c179b09fc25584c172e78db2c9d3ff02

Observation 58294f10-753f-46e2-895f-e3ea0d455fe2 · outbound

This paper cites Continual Learning in Vision-Language Models via Aligned Model Merging.

Towards Human-like Physical Intelligence: Lifelong Vision-Language-Action Learning for Robotic Manipulation Continual Learning in Vision-Language Models via Aligned Model Merging

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-02T00:57:43.881234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:57:43.881234Z digest=sha256:2ecde6fc3cc4738f35a3c93ad47fd3de68a43a22c3361710046d888f8a07d04b

Observation 3449da03-ae32-4162-bdf3-0cbecd78b540 · outbound

This paper cites Octo: An Open-Source Generalist Robot Policy.

Towards Human-like Physical Intelligence: Lifelong Vision-Language-Action Learning for Robotic Manipulation Octo: An Open-Source Generalist Robot Policy

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-02T00:57:44.038690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:57:44.038690Z digest=sha256:409b313cf549f745b4e574344ad9ff0566e6581e3eb92385535d8e39f5c06af6

Observation a5dcc6ca-0af3-4495-8190-012c45d24af2 · outbound

This paper cites A comprehensive survey of continual learning: Theory, method and application.IEEE transactions on pattern analysis and machine intelligence, 46(8): 5362–5383, 2024.

Towards Human-like Physical Intelligence: Lifelong Vision-Language-Action Learning for Robotic Manipulation A comprehensive survey of continual learning: Theory, method and application.IEEE transactions on pattern analysis and machine intelligence, 46(8): 5362–5383, 2024

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-02T00:57:44.188337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:57:44.188337Z digest=sha256:377ed2e861ab8498bb0a7fa11c4a402641daab3ac7fe038595b7ad14a3a50d50

Observation 697e9aba-0960-4c83-8869-10f31a36f1bd · outbound

This paper cites DualPrompt: Complementary Prompting for Rehearsal-free Continual Learning.

Towards Human-like Physical Intelligence: Lifelong Vision-Language-Action Learning for Robotic Manipulation DualPrompt: Complementary Prompting for Rehearsal-free Continual Learning

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-02T00:57:44.354469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:57:44.354469Z digest=sha256:2fbf64b85b71ff2828485f0c8c39aad34c0fb7c3c5d89c88610de80f1e1d6d2b

Observation c29e4a62-2969-42e5-b3e9-618575a75da8 · outbound

This paper cites Tinyvla: Towards fast, data-efficient vision-language-action models for robotic manipulation.IEEE Robotics and Automation Letters, 2025.

Towards Human-like Physical Intelligence: Lifelong Vision-Language-Action Learning for Robotic Manipulation Tinyvla: Towards fast, data-efficient vision-language-action models for robotic manipulation.IEEE Robotics and Automation Letters, 2025

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-02T00:57:44.474480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:57:44.474480Z digest=sha256:926269ddf7e896b0171d5b608ae112e724e5fe2e4de85edddeeadcc2a8cd24b9

Observation 69001641-1349-411b-8022-4794ddb7017b · outbound

This paper cites Continual world: A robotic benchmark for continual reinforcement learning.

Towards Human-like Physical Intelligence: Lifelong Vision-Language-Action Learning for Robotic Manipulation Continual world: A robotic benchmark for continual reinforcement learning

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-02T00:57:44.733577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:57:44.733577Z digest=sha256:9f23c29bee0cb4f917d1106c7850bc6a233493efb617c006b1569d42a747ca4c

Observation 543c0b8f-1bf6-4ffa-b97d-a34062faa093 · outbound

This paper cites Long-horizon language-conditioned imitation learning for robotic manipulation.

Towards Human-like Physical Intelligence: Lifelong Vision-Language-Action Learning for Robotic Manipulation Long-horizon language-conditioned imitation learning for robotic manipulation

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-02T00:57:44.809271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:57:44.809271Z digest=sha256:a59d3802f1b3d07fe4c087770438d5a40eca151b3cf950250f0e9957648f07d5

Observation 9c3e7bb1-682a-4721-b50b-ea52e0df196e · outbound

This paper cites Boosting continual learning of vision-language models via mixture-of-experts adapters.

Towards Human-like Physical Intelligence: Lifelong Vision-Language-Action Learning for Robotic Manipulation Boosting continual learning of vision-language models via mixture-of-experts adapters

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-02T00:57:44.957923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:57:44.957923Z digest=sha256:428f8e61db374c49c622126586f0ea83eea0a096c5a9306371571d5b15bf7443

Observation 6a83ff64-d458-4670-9690-e4b773621fde · outbound

This paper cites Atomicvla: Unlocking the potential of atomic skill learning in robots.arXiv preprint arXiv:2603.07648, 2026.

Towards Human-like Physical Intelligence: Lifelong Vision-Language-Action Learning for Robotic Manipulation Atomicvla: Unlocking the potential of atomic skill learning in robots.arXiv preprint arXiv:2603.07648, 2026

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-02T00:57:45.006248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:57:45.006248Z digest=sha256:abf75c36e775b96a8a1ac7ef3775f96c6185dbbc5571475c77d7504d874f3a47

Observation 20ba1cbf-ecb0-48d1-9ff7-2422168cd8fa · outbound

This paper cites Mllm-cl: Continual learning for multimodal large language models.arXiv preprint arXiv:2506.05453, 2025.

Towards Human-like Physical Intelligence: Lifelong Vision-Language-Action Learning for Robotic Manipulation Mllm-cl: Continual learning for multimodal large language models.arXiv preprint arXiv:2506.05453, 2025

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-02T00:57:45.048809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:57:45.048809Z digest=sha256:c21e75c2a122c62e250b72374160ce6d2240a55a4bca72e807ab501101685db0

Observation 3bf46d08-2adf-45ed-b526-6ccba4889e8b · outbound

This paper cites Information-theoretic constraints for continual vision-language-action alignment.arXiv preprint arXiv:2603.13335, 2026.

Towards Human-like Physical Intelligence: Lifelong Vision-Language-Action Learning for Robotic Manipulation Information-theoretic constraints for continual vision-language-action alignment.arXiv preprint arXiv:2603.13335, 2026

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-02T00:57:45.102602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:57:45.102602Z digest=sha256:3b4a89576436c0cf9728adedce4166416fbd51230dd20ea330b0e8e048a8ef77

Observation 965b4692-6aed-46b2-801c-c004dcfbc883 · outbound

This paper cites imanip: Skill-incremental learning for robotic manipulation.

Towards Human-like Physical Intelligence: Lifelong Vision-Language-Action Learning for Robotic Manipulation imanip: Skill-incremental learning for robotic manipulation

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-02T00:57:45.185859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:57:45.185859Z digest=sha256:5fb3559b6e6f53c118fb5af28fbb07d26067474d9476c076ea4248ff71a3cfea

Pith citing papers

No inbound Pith citation observations are available.