Pith. sign in

Paper Citation Record · LEDGER

Towards Human-like Physical Intelligence: Lifelong Vision-Language-Action Learning for Robotic Manipulation

As of 7 August 2026, this Paper Citation Record lists 59 of 59 outbound references and 0 inbound Pith citation observations for arXiv:2607.14852.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.14852 v2

Coverage vector

measured 59 of 59 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-02T00:57:45.185859Z

measured 59 of 59 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

59 of 59 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved59
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6d4e5abc-5875-460f-9143-60c6687f9e05 · outbound

This paper cites Do As I Can, Not As I Say: Grounding Language in Robotic Affordances.

Towards Human-like Physical Intelligence: Lifelong Vision-Language-Action Learning for Robotic Manipulation Do As I Can, Not As I Say: Grounding Language in Robotic Affordances

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-02T00:57:38.361430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:57:38.361430Z digest=sha256:76aa84d0a5a17cbc04732f628780208033b6a8790cc6a8588f8721b000001266

Observation 8da652d4-1b83-4fb1-b00e-29b2ac5e0a2e · outbound

This paper cites $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control.

Towards Human-like Physical Intelligence: Lifelong Vision-Language-Action Learning for Robotic Manipulation $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-02T00:57:38.446664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:57:38.446664Z digest=sha256:00df0cc8f9851d27575cde14d2a117fe7518a0674459b486b122b009b5be1d1d

Observation 5ecd9849-3188-4f1e-872d-c057bfe1feaf · outbound

This paper cites RoboCat: A Self-Improving Generalist Agent for Robotic Manipulation.

Towards Human-like Physical Intelligence: Lifelong Vision-Language-Action Learning for Robotic Manipulation RoboCat: A Self-Improving Generalist Agent for Robotic Manipulation

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-02T00:57:38.564461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:57:38.564461Z digest=sha256:3acbe7be1d38af0b093e5b542e28b5d71d8d2f2b1b6047b814306d43e0d6319c

Observation d23d7b75-13be-4f71-b49f-811a87159e0e · outbound

This paper cites RT-1: Robotics Transformer for Real-World Control at Scale.

Towards Human-like Physical Intelligence: Lifelong Vision-Language-Action Learning for Robotic Manipulation RT-1: Robotics Transformer for Real-World Control at Scale

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-02T00:57:38.675970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:57:38.675970Z digest=sha256:83ac78fc1e7f5714cf76e7ef23b605d40ef4d1968d0f1f65c5116439060fe2a3

Observation bfd3f953-bb97-48a6-8ad8-0565fa985ce2 · outbound

This paper cites RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control.

Towards Human-like Physical Intelligence: Lifelong Vision-Language-Action Learning for Robotic Manipulation RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-02T00:57:38.798291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:57:38.798291Z digest=sha256:790f5a34dc46373c5e1cdb72073756b23958c5c27dd51d5dcd5108e5f7c2c51b

Observation 8fa60eea-fa48-4433-b372-6b24d4ce6af2 · outbound

This paper cites Riemannian walk for incremental learning: Understanding forgetting and intransigence.

Towards Human-like Physical Intelligence: Lifelong Vision-Language-Action Learning for Robotic Manipulation Riemannian walk for incremental learning: Understanding forgetting and intransigence

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-02T00:57:38.868336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:57:38.868336Z digest=sha256:ab602a7d68af8e138961497111c787683b72385cd65bb7ed9cc5cbcd43f90194

Observation ac24ec8c-df59-44f7-9361-66b1e17324df · outbound

This paper cites On Tiny Episodic Memories in Continual Learning.

Towards Human-like Physical Intelligence: Lifelong Vision-Language-Action Learning for Robotic Manipulation On Tiny Episodic Memories in Continual Learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-02T00:57:38.974751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:57:38.974751Z digest=sha256:e37964a0310d773c91d2e0d7788b9feffdd6dd4518c1986296ef9fe874880c7b

Observation 97243f6c-be38-4b5a-a6fe-ca2c7dec90b3 · outbound

This paper cites Diffusion Policy: Visuomotor Policy Learning via Action Diffusion.

Towards Human-like Physical Intelligence: Lifelong Vision-Language-Action Learning for Robotic Manipulation Diffusion Policy: Visuomotor Policy Learning via Action Diffusion

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-02T00:57:39.140456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:57:39.140456Z digest=sha256:664a45f6446cc7191e8d9b5dc671bbed829188d459855e9ac8cafe99f98cdbac

Observation cf28c9fc-2b4b-4556-8395-f2bd1a2639dc · outbound

This paper cites OpenHelix: A Short Survey, Empirical Analysis, and Open-Source Dual-System VLA Model for Robotic Manipulation.

Towards Human-like Physical Intelligence: Lifelong Vision-Language-Action Learning for Robotic Manipulation OpenHelix: A Short Survey, Empirical Analysis, and Open-Source Dual-System VLA Model for Robotic Manipulation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-02T00:57:39.308136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:57:39.308136Z digest=sha256:7dcc8d20df20ff0aef6bbe2bc3f9f3150391e3144d8c006a88e33717526b7b72

Observation 1a53177d-b354-436d-bfe3-879228996277 · outbound

This paper cites Loss of plasticity in deep continual learning.Nature, 632(8026):768–774, 2024.

Towards Human-like Physical Intelligence: Lifelong Vision-Language-Action Learning for Robotic Manipulation Loss of plasticity in deep continual learning.Nature, 632(8026):768–774, 2024

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-02T00:57:39.473600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:57:39.473600Z digest=sha256:180dbe6e7af328326a71c76257809a13fd403062435f26dd3d8f79b97859430c

Observation 6c03f0c3-b7d1-4ecc-85d8-62fe8f30faf9 · outbound

This paper cites Palm-e: An embodied multimodal language model.

Towards Human-like Physical Intelligence: Lifelong Vision-Language-Action Learning for Robotic Manipulation Palm-e: An embodied multimodal language model

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-02T00:57:39.612805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:57:39.612805Z digest=sha256:0265a1b7082e8d6780bf9e3404d7e96b4c0682a4c0d77f66bffd471c3139ae5c

Observation 81402915-0aa4-4d73-9b3f-e82aec51676b · outbound

This paper cites Benchmarking Vision, Language, & Action Models in Procedurally Generated, Open Ended Action Environments.

Towards Human-like Physical Intelligence: Lifelong Vision-Language-Action Learning for Robotic Manipulation Benchmarking Vision, Language, & Action Models in Procedurally Generated, Open Ended Action Environments

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-02T00:57:39.775665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:57:39.775665Z digest=sha256:cee1b8827c4c1d92c4679713208ff8fef9243233aaba9073896abfb154715514

Observation 4ba87ad8-c25a-46fe-ae7d-0666ed9f4d4b · outbound

This paper cites Lora: Low-rank adaptation of large language models.

Towards Human-like Physical Intelligence: Lifelong Vision-Language-Action Learning for Robotic Manipulation Lora: Low-rank adaptation of large language models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T00:57:39.937349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:57:39.937349Z digest=sha256:b252484193c3b602520a289e7ce67df5ab5246b48bb672b566e1521594b5dcc2

Observation e5a8d7e6-4f24-4cf7-b19c-973217c16924 · outbound

This paper cites VoxPoser: Composable 3D Value Maps for Robotic Manipulation with Language Models.

Towards Human-like Physical Intelligence: Lifelong Vision-Language-Action Learning for Robotic Manipulation VoxPoser: Composable 3D Value Maps for Robotic Manipulation with Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-02T00:57:40.106650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:57:40.106650Z digest=sha256:ede741a7c95cbc46e343a2be8c7d5f3444b627d14b85431ce650c28d04f8efda

Observation cae95d81-6b39-474d-a382-5de0ee2bbb6e · outbound

This paper cites $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization.

Towards Human-like Physical Intelligence: Lifelong Vision-Language-Action Learning for Robotic Manipulation $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-02T00:57:40.272666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:57:40.272666Z digest=sha256:bf7729027ef5571061040b602d7e33e109e8c95db1a51db346e7f4407ca5013f

Observation e7c4049a-4dde-4140-8e42-6f4f21c82016 · outbound

This paper cites VIMA: General Robot Manipulation with Multimodal Prompts.

Towards Human-like Physical Intelligence: Lifelong Vision-Language-Action Learning for Robotic Manipulation VIMA: General Robot Manipulation with Multimodal Prompts

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-02T00:57:40.434766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:57:40.434766Z digest=sha256:3c4ac2d65b3637ca1bdbe4a06f445ef75310f3a7a70db541d5ed7401f0295abd

Observation 671d6ba0-d7df-499c-9274-8b1d699d88fd · outbound

This paper cites Srt-h: A hierarchical framework for autonomous surgery via language-conditioned imitation learning.Science robotics, 10(104):eadt5254, 2025.

Towards Human-like Physical Intelligence: Lifelong Vision-Language-Action Learning for Robotic Manipulation Srt-h: A hierarchical framework for autonomous surgery via language-conditioned imitation learning.Science robotics, 10(104):eadt5254, 2025

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-02T00:57:40.557073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:57:40.557073Z digest=sha256:26445472893d989aede4da4a2b9c22ea5ef3c42bf73d3dfab7388d58ff037300

Observation 17855fb2-1f2f-45cd-96f9-767e61ea49ca · outbound

This paper cites OpenVLA: An Open-Source Vision-Language-Action Model.

Towards Human-like Physical Intelligence: Lifelong Vision-Language-Action Learning for Robotic Manipulation OpenVLA: An Open-Source Vision-Language-Action Model

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-02T00:57:40.564504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:57:40.564504Z digest=sha256:0bafb6e3e0ff39fec99771e335be1f36af00bd8030ed9d0583b50d66a8712bf9

Observation 7f1747ac-9dad-42dc-b167-da8070caa109 · outbound

This paper cites Overcoming catastrophic forgetting in neural networks.Proceedings of the national academy of sciences, 114(13): 3521–3526, 2017.

Towards Human-like Physical Intelligence: Lifelong Vision-Language-Action Learning for Robotic Manipulation Overcoming catastrophic forgetting in neural networks.Proceedings of the national academy of sciences, 114(13): 3521–3526, 2017

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-02T00:57:40.626531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:57:40.626531Z digest=sha256:2608ab6840c75e3b178ff2020315568ce18c6a48ba4edf5c5ebf04870151e135

Observation 973ce0bb-7239-40ee-8772-a2cee4eca53c · outbound

This paper cites Reinforcement Fine-Tuning Naturally Mitigates Forgetting in Continual Post-Training.

Towards Human-like Physical Intelligence: Lifelong Vision-Language-Action Learning for Robotic Manipulation Reinforcement Fine-Tuning Naturally Mitigates Forgetting in Continual Post-Training

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-02T00:57:40.746075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:57:40.746075Z digest=sha256:b8abc81ea739b2804971aff01eb9f553c943add477e927fe44ba7cbebe5e2bcd

Observation 920df1b8-d243-4890-b5e7-f1a1efd5aa4b · outbound

This paper cites The power of scale for parameter-efficient prompt tuning.

Towards Human-like Physical Intelligence: Lifelong Vision-Language-Action Learning for Robotic Manipulation The power of scale for parameter-efficient prompt tuning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-02T00:57:40.923216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:57:40.923216Z digest=sha256:63c6e2d0de6661eb59cf6504964d3efe5e7e859e2cc1868413bed3305a3e0a42

Observation 1c832505-4417-4d0b-ab31-5a3909d5fd35 · outbound

This paper cites Remem-vla: Empowering vision-language-action model with memory via dual-level recurrent queries.arXiv preprint arXiv:2603.12942, 2026.

Towards Human-like Physical Intelligence: Lifelong Vision-Language-Action Learning for Robotic Manipulation Remem-vla: Empowering vision-language-action model with memory via dual-level recurrent queries.arXiv preprint arXiv:2603.12942, 2026

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-02T00:57:41.063842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:57:41.063842Z digest=sha256:2673a93779178dec942f7568c166b5db3540d9b7a80882f7e22a37c1e5d76546

Observation 1db9480a-222b-4401-a7ef-7963e83fa241 · outbound

This paper cites Prefix-tuning: Optimizing continuous prompts for generation.

Towards Human-like Physical Intelligence: Lifelong Vision-Language-Action Learning for Robotic Manipulation Prefix-tuning: Optimizing continuous prompts for generation

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-02T00:57:41.166042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:57:41.166042Z digest=sha256:4a1e90499729214a6db18bd93e10a592aeef6e6f61764bbfdef68342bcee660f

Observation df862715-64fe-47c8-a6f2-2a3e0c02350e · outbound

This paper cites Learning without forgetting.IEEE transactions on pattern analysis and machine intelligence, 40(12):2935–2947, 2017.

Towards Human-like Physical Intelligence: Lifelong Vision-Language-Action Learning for Robotic Manipulation Learning without forgetting.IEEE transactions on pattern analysis and machine intelligence, 40(12):2935–2947, 2017

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-02T00:57:41.271781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:57:41.271781Z digest=sha256:f9612c65acc7bcee2aacf15665b920f88a0d82e4b8c39789c161452b52a9d700

Observation 6d394b78-c309-43d3-aaa3-a0f41b8b07f1 · outbound

This paper cites Learning without forgetting.

Towards Human-like Physical Intelligence: Lifelong Vision-Language-Action Learning for Robotic Manipulation Learning without forgetting

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-02T00:57:41.376843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:57:41.376843Z digest=sha256:6c66a6bd4d1580da41d3d099c3dc263f30e7941db59fb442fe254f34387e96ea

Observation d083236b-4763-4028-b504-68e83c95dba4 · outbound

This paper cites Code as policies: Language model programs for embodied control.

Towards Human-like Physical Intelligence: Lifelong Vision-Language-Action Learning for Robotic Manipulation Code as policies: Language model programs for embodied control

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-02T00:57:41.519701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:57:41.519701Z digest=sha256:f664ab6e611d28115391ebcd82736c2e3b3f91220481dff3f27cd7a9ed99eed5

Observation 8a6204cf-c9a6-406a-81ed-fce26d6c0c00 · outbound

This paper cites Never-ending behavior-cloning agent for robotic manipulation.arXiv preprint arXiv:2403.00336, 2024.

Towards Human-like Physical Intelligence: Lifelong Vision-Language-Action Learning for Robotic Manipulation Never-ending behavior-cloning agent for robotic manipulation.arXiv preprint arXiv:2403.00336, 2024

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-02T00:57:41.679900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:57:41.679900Z digest=sha256:2c3d131200066a6027cda5b2cb29ff9e782d60d8b507450cb61eeb89bca6a0e5

Observation ae038217-681c-40b4-934c-8fbd01cfe4a2 · outbound

This paper cites Pixelvla: Advancing pixel-level understanding in vision-language-action model.arXiv preprint arXiv:2511.01571, 2025.

Towards Human-like Physical Intelligence: Lifelong Vision-Language-Action Learning for Robotic Manipulation Pixelvla: Advancing pixel-level understanding in vision-language-action model.arXiv preprint arXiv:2511.01571, 2025

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-02T00:57:41.790039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:57:41.790039Z digest=sha256:10d4efe009ab7208d92ef312415f6617b8fcbea97725779927917fd6bed50556

Observation f4e12658-d377-4652-acaf-556b08b76456 · outbound

This paper cites Discrete Diffusion VLA: Bringing Discrete Diffusion to Action Decoding in Vision-Language-Action Policies.

Towards Human-like Physical Intelligence: Lifelong Vision-Language-Action Learning for Robotic Manipulation Discrete Diffusion VLA: Bringing Discrete Diffusion to Action Decoding in Vision-Language-Action Policies

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-02T00:57:41.909520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:57:41.909520Z digest=sha256:c99856c030ebc880526ed8f2a41f018bd22625f62dede8a5e90994c6ca24b117

Observation 75d14dfb-e1ee-489d-903a-193620f4e0f4 · outbound

This paper cites Showui: One vision-language-action model for gui visual agent.

Towards Human-like Physical Intelligence: Lifelong Vision-Language-Action Learning for Robotic Manipulation Showui: One vision-language-action model for gui visual agent

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-02T00:57:42.023492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:57:42.023492Z digest=sha256:14067f529f016efdf19032e22b86ba2eefc347f49c28b803992670a567c4a462

Observation c9dda9e5-a346-4835-9596-a5462827eb2a · outbound

This paper cites Libero: Benchmarking knowledge transfer for lifelong robot learning.Advances in Neural Information Processing Systems, 36:44776–44791, 2023.

Towards Human-like Physical Intelligence: Lifelong Vision-Language-Action Learning for Robotic Manipulation Libero: Benchmarking knowledge transfer for lifelong robot learning.Advances in Neural Information Processing Systems, 36:44776–44791, 2023

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-02T00:57:42.163804Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:57:42.163804Z digest=sha256:03601e944162b0cd1d4e8000144743aee426c07d6457f2b377fb2f0bb35af8c4

Observation 2861effb-1148-4105-b9e3-6352c2a770b6 · outbound

This paper cites Pretrained vision-language-action models are surprisingly resistant to forgetting in continual learning.arXiv preprint arXiv:2603.03818, 2026.

Towards Human-like Physical Intelligence: Lifelong Vision-Language-Action Learning for Robotic Manipulation Pretrained vision-language-action models are surprisingly resistant to forgetting in continual learning.arXiv preprint arXiv:2603.03818, 2026

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-02T00:57:42.309881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:57:42.309881Z digest=sha256:f800e203e9c50f4a364db38c72bc32580bf8e7a6beb7336110f41c76ddfba51a

Observation ad0e677d-ee68-4ae8-b596-9ab348fca764 · outbound

This paper cites Spatial-Temporal Aware Visuomotor Diffusion Policy Learning.

Towards Human-like Physical Intelligence: Lifelong Vision-Language-Action Learning for Robotic Manipulation Spatial-Temporal Aware Visuomotor Diffusion Policy Learning

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-02T00:57:42.384758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:57:42.384758Z digest=sha256:2fae186cdde2c4e86a65701c69f07b1c941152a9666363d4d6ba45257355fd12

Observation 76a1a350-ba0f-4403-b458-12b73c023994 · outbound

This paper cites Packnet: Adding multiple tasks to a single network by iterative pruning.

Towards Human-like Physical Intelligence: Lifelong Vision-Language-Action Learning for Robotic Manipulation Packnet: Adding multiple tasks to a single network by iterative pruning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-02T00:57:42.499558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:57:42.499558Z digest=sha256:7b9bf969fe057d9f726a038bfdd08b7b1571f5cfaebc9a602dc6a6ee519c9734

Observation f859b398-a502-45a9-b0de-7c514b6607a0 · outbound

This paper cites Preserving and combining knowledge in robotic lifelong reinforcement learning.Nature Machine Intelligence, 7(2):256–269, 2025.

Towards Human-like Physical Intelligence: Lifelong Vision-Language-Action Learning for Robotic Manipulation Preserving and combining knowledge in robotic lifelong reinforcement learning.Nature Machine Intelligence, 7(2):256–269, 2025

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-02T00:57:42.643648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:57:42.643648Z digest=sha256:8813de8d74b6e5cfc9b98fcd42d1331c03efc5fa04c4c7b7a4c26f47373f9084

Observation e21013a9-57c6-4ceb-8be4-35d30bdc961b · outbound

This paper cites LLARVA: Vision-Action Instruction Tuning Enhances Robot Learning.

Towards Human-like Physical Intelligence: Lifelong Vision-Language-Action Learning for Robotic Manipulation LLARVA: Vision-Action Instruction Tuning Enhances Robot Learning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-02T00:57:42.768376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:57:42.768376Z digest=sha256:aff523bc4139f344b60a7db10589d8482daa27cfc528f6e5628a14c59753ba21

Observation b7e19fac-c9c1-4f25-801d-4be669680dc4 · outbound

This paper cites Open X-Embodiment: Robotic Learning Datasets and RT-X Models.

Towards Human-like Physical Intelligence: Lifelong Vision-Language-Action Learning for Robotic Manipulation Open X-Embodiment: Robotic Learning Datasets and RT-X Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-02T00:57:42.876938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:57:42.876938Z digest=sha256:5c1b1e3efe5dc5246dc8e3563a3e4c6dcb2573f756c8223b0f73f28471d7c910

Observation 8e5de4bf-5370-457c-bf11-394968808e29 · outbound

This paper cites FAST: Efficient Action Tokenization for Vision-Language-Action Models.

Towards Human-like Physical Intelligence: Lifelong Vision-Language-Action Learning for Robotic Manipulation FAST: Efficient Action Tokenization for Vision-Language-Action Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-02T00:57:42.983225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:57:42.983225Z digest=sha256:5fe0664d7f0f9c7ceb33f77fe94c8578f5c69ff32510505b1a7a625d41d593e4

Observation f093673d-4b63-4894-9d1d-2b3ffc21433c · outbound

This paper cites ${\pi}_{0.7}$: a Steerable Generalist Robotic Foundation Model with Emergent Capabilities.

Towards Human-like Physical Intelligence: Lifelong Vision-Language-Action Learning for Robotic Manipulation ${\pi}_{0.7}$: a Steerable Generalist Robotic Foundation Model with Emergent Capabilities

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-02T00:57:43.044566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:57:43.044566Z digest=sha256:7adaa583c494f18ad9465fe0f89f097bb8195b6baa9603da19d1f3f6bb95fc5a

Observation deffecab-8279-43d4-a60f-555ffe976a3b · outbound

This paper cites Incremental classifier and representation learning.

Towards Human-like Physical Intelligence: Lifelong Vision-Language-Action Learning for Robotic Manipulation Incremental classifier and representation learning

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-02T00:57:43.119826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:57:43.119826Z digest=sha256:b79f9345fed0f09fbbf7ece6cff24f199de68911512fdfaad788230ab5d1f2d6

Observation 86645130-7d20-4dd4-838d-caa7f7cfebf8 · outbound

This paper cites A Generalist Agent.

Towards Human-like Physical Intelligence: Lifelong Vision-Language-Action Learning for Robotic Manipulation A Generalist Agent

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-02T00:57:43.200599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:57:43.200599Z digest=sha256:6943f8177e333ece2bbac2563e48f3e955a570c0231fe5378d5d67c56e5dc270

Observation e9c286bb-d113-4d4a-bf70-b76e4fcdae19 · outbound

This paper cites Progressive Neural Networks.

Towards Human-like Physical Intelligence: Lifelong Vision-Language-Action Learning for Robotic Manipulation Progressive Neural Networks

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-02T00:57:43.302391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:57:43.302391Z digest=sha256:2be89b147bbf0f37db628a25842488c3a64d7cab3c4c7078fad775eabd83bf8b

Observation 1d130e34-7e7e-4643-96c8-fc4b7eb9824f · outbound

This paper cites RL's Razor: Why Online Reinforcement Learning Forgets Less.

Towards Human-like Physical Intelligence: Lifelong Vision-Language-Action Learning for Robotic Manipulation RL's Razor: Why Online Reinforcement Learning Forgets Less

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-02T00:57:43.368877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:57:43.368877Z digest=sha256:c0563ff90bb943e8ef60339301f84e202844ec31447ad5d659580e0a92b9ec4b

Observation 4091b805-2fee-4570-bc78-e314dad59af3 · outbound

This paper cites Cliport: What and where pathways for robotic manipulation.

Towards Human-like Physical Intelligence: Lifelong Vision-Language-Action Learning for Robotic Manipulation Cliport: What and where pathways for robotic manipulation

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-02T00:57:43.430502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:57:43.430502Z digest=sha256:a13f7b02345182090b4d362bd313e31822ab9804af388722dff2d94af8d7b4a0

Observation 30b3ec67-9175-4544-8a41-6fa4ed8db12b · outbound

This paper cites Perceiver-actor: A multi-task transformer for robotic manipulation.

Towards Human-like Physical Intelligence: Lifelong Vision-Language-Action Learning for Robotic Manipulation Perceiver-actor: A multi-task transformer for robotic manipulation

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-02T00:57:43.556355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:57:43.556355Z digest=sha256:47ef43ffbaa1b96cf49bea29666b9361e875797cfb214bab16a61a380874a655

Observation 2b2a4d7c-00e1-47e2-a65c-e5d731c23ac4 · outbound

This paper cites SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics.

Towards Human-like Physical Intelligence: Lifelong Vision-Language-Action Learning for Robotic Manipulation SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-02T00:57:43.651170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:57:43.651170Z digest=sha256:b32bd6ba5d393040f8fe1e4d11807f4aa88a67576e52ef2cbede237844adb8cd

Observation 68528695-96c2-4fa0-a50c-627e6815c1ab · outbound

This paper cites Coda-prompt: Continual decomposed attention- based prompting for rehearsal-free continual learning.

Towards Human-like Physical Intelligence: Lifelong Vision-Language-Action Learning for Robotic Manipulation Coda-prompt: Continual decomposed attention- based prompting for rehearsal-free continual learning

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-02T00:57:43.723645Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:57:43.723645Z digest=sha256:d677c0a012a2dbbc3938d4fde2ec164c78bcf2e01477a61f19cad6ba731e67db

Observation 58294f10-753f-46e2-895f-e3ea0d455fe2 · outbound

This paper cites Continual Learning in Vision-Language Models via Aligned Model Merging.

Towards Human-like Physical Intelligence: Lifelong Vision-Language-Action Learning for Robotic Manipulation Continual Learning in Vision-Language Models via Aligned Model Merging

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-02T00:57:43.881234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:57:43.881234Z digest=sha256:31ee4f99921c0f0b147afdf8367c75164f5a42394206f0be9cc496c263938b2d

Observation 3449da03-ae32-4162-bdf3-0cbecd78b540 · outbound

This paper cites Octo: An Open-Source Generalist Robot Policy.

Towards Human-like Physical Intelligence: Lifelong Vision-Language-Action Learning for Robotic Manipulation Octo: An Open-Source Generalist Robot Policy

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-02T00:57:44.038690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:57:44.038690Z digest=sha256:6cfcc3a98453378d41b1175e23983e401706dd63702522705f119c7dd785a58d

Observation a5dcc6ca-0af3-4495-8190-012c45d24af2 · outbound

This paper cites A comprehensive survey of continual learning: Theory, method and application.IEEE transactions on pattern analysis and machine intelligence, 46(8): 5362–5383, 2024.

Towards Human-like Physical Intelligence: Lifelong Vision-Language-Action Learning for Robotic Manipulation A comprehensive survey of continual learning: Theory, method and application.IEEE transactions on pattern analysis and machine intelligence, 46(8): 5362–5383, 2024

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-02T00:57:44.188337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:57:44.188337Z digest=sha256:ad98757e82c4a39a2520709a81d6b6f431b408440f211dd8cd9833b43430b7db

Observation 697e9aba-0960-4c83-8869-10f31a36f1bd · outbound

This paper cites DualPrompt: Complementary Prompting for Rehearsal-free Continual Learning.

Towards Human-like Physical Intelligence: Lifelong Vision-Language-Action Learning for Robotic Manipulation DualPrompt: Complementary Prompting for Rehearsal-free Continual Learning

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-02T00:57:44.354469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:57:44.354469Z digest=sha256:1d5ae7f9bbbcf4761a8d0ead50df743eac08058fcddae017f36740e78055bae0

Observation c29e4a62-2969-42e5-b3e9-618575a75da8 · outbound

This paper cites Tinyvla: Towards fast, data-efficient vision-language-action models for robotic manipulation.IEEE Robotics and Automation Letters, 2025.

Towards Human-like Physical Intelligence: Lifelong Vision-Language-Action Learning for Robotic Manipulation Tinyvla: Towards fast, data-efficient vision-language-action models for robotic manipulation.IEEE Robotics and Automation Letters, 2025

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-02T00:57:44.474480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:57:44.474480Z digest=sha256:caa2b2bfd1f694d5b1e0b0e311df4dd8c20e8f4707c4b3cd8e998e96bf133163

Observation 69001641-1349-411b-8022-4794ddb7017b · outbound

This paper cites Continual world: A robotic benchmark for continual reinforcement learning.

Towards Human-like Physical Intelligence: Lifelong Vision-Language-Action Learning for Robotic Manipulation Continual world: A robotic benchmark for continual reinforcement learning

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-02T00:57:44.733577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:57:44.733577Z digest=sha256:1d499f98671bc9dd5711c37b31a01a3a355d81d2e098b9eaad31ab430905c471

Observation 543c0b8f-1bf6-4ffa-b97d-a34062faa093 · outbound

This paper cites Long-horizon language-conditioned imitation learning for robotic manipulation.

Towards Human-like Physical Intelligence: Lifelong Vision-Language-Action Learning for Robotic Manipulation Long-horizon language-conditioned imitation learning for robotic manipulation

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-02T00:57:44.809271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:57:44.809271Z digest=sha256:9e30afd7e23976e85c8bdabcfe6a5dc7da2aebeefe7d28d156b98c87d877e9e7

Observation 9c3e7bb1-682a-4721-b50b-ea52e0df196e · outbound

This paper cites Boosting continual learning of vision-language models via mixture-of-experts adapters.

Towards Human-like Physical Intelligence: Lifelong Vision-Language-Action Learning for Robotic Manipulation Boosting continual learning of vision-language models via mixture-of-experts adapters

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-02T00:57:44.957923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:57:44.957923Z digest=sha256:6711a92ab3c7e01cc5ae1addb8f1d08a6bf2b47c9d8f0a30825a40c7533a9b20

Observation 6a83ff64-d458-4670-9690-e4b773621fde · outbound

This paper cites Atomicvla: Unlocking the potential of atomic skill learning in robots.arXiv preprint arXiv:2603.07648, 2026.

Towards Human-like Physical Intelligence: Lifelong Vision-Language-Action Learning for Robotic Manipulation Atomicvla: Unlocking the potential of atomic skill learning in robots.arXiv preprint arXiv:2603.07648, 2026

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-02T00:57:45.006248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:57:45.006248Z digest=sha256:7d952d4a8ebfa15a994584178451f970fd043f3c5c2132294055439aa5d66fa4

Observation 20ba1cbf-ecb0-48d1-9ff7-2422168cd8fa · outbound

This paper cites Mllm-cl: Continual learning for multimodal large language models.arXiv preprint arXiv:2506.05453, 2025.

Towards Human-like Physical Intelligence: Lifelong Vision-Language-Action Learning for Robotic Manipulation Mllm-cl: Continual learning for multimodal large language models.arXiv preprint arXiv:2506.05453, 2025

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-02T00:57:45.048809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:57:45.048809Z digest=sha256:15ac8785eb1b16765f3b5f9fdc6fedc0174c2e9a56631d5a34e148b149ea7df2

Observation 3bf46d08-2adf-45ed-b526-6ccba4889e8b · outbound

This paper cites Information-theoretic constraints for continual vision-language-action alignment.arXiv preprint arXiv:2603.13335, 2026.

Towards Human-like Physical Intelligence: Lifelong Vision-Language-Action Learning for Robotic Manipulation Information-theoretic constraints for continual vision-language-action alignment.arXiv preprint arXiv:2603.13335, 2026

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-02T00:57:45.102602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:57:45.102602Z digest=sha256:b521c6be658a499de959356c7f32dfb9fe197338326ec821aa7e7f8af0653b95

Observation 965b4692-6aed-46b2-801c-c004dcfbc883 · outbound

This paper cites imanip: Skill-incremental learning for robotic manipulation.

Towards Human-like Physical Intelligence: Lifelong Vision-Language-Action Learning for Robotic Manipulation imanip: Skill-incremental learning for robotic manipulation

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-02T00:57:45.185859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:57:45.185859Z digest=sha256:f5adc46baa6ab20aa253fc3738ca002acb28fd8a0fbd3fbd260dea5a7c56153d

Pith citing papers

No inbound Pith citation observations are available.