Pith. sign in

Paper Citation Record · LEDGER

MoLe-VLA: Dynamic Layer-skipping Vision Language Action Model via Mixture-of-Layers for Efficient Robot Manipulation

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 25 inbound Pith citation observations for arXiv:2503.20384.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2503.20384 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 25 of 25 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 25 of 25 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:44:06.343233Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T11:18:03.368562Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 6f11c451-0ad3-4ec4-acac-2e7d5d3fa5d5 · inbound

Towards a Generalizable Bimanual Foundation Policy via Flow-based Video Prediction cites this paper.

Towards a Generalizable Bimanual Foundation Policy via Flow-based Video Prediction MoLe-VLA: Dynamic Layer-skipping Vision Language Action Model via Mixture-of-Layers for Efficient Robot Manipulation

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:06.343233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:06.343233Z digest=sha256:af965bfdd3240e4c1662b0b43742812f7d404a30aaf61419cb1abc189938b66b

Observation 25aa3d8d-ae22-4f6a-993b-bf3401d4953a · inbound

EfficientVLA: Training-Free Acceleration and Compression for Vision-Language-Action Models cites this paper.

EfficientVLA: Training-Free Acceleration and Compression for Vision-Language-Action Models MoLe-VLA: Dynamic Layer-skipping Vision Language Action Model via Mixture-of-Layers for Efficient Robot Manipulation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T04:42:22.434287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:42:22.434287Z digest=sha256:44d8a9ee3929ad3f43214f84c221d53510ddaa3a4e24777c1afa6be6daab8ca6

Observation fbe82f9c-440f-4a16-9397-b85d4f3eb02a · inbound

ROSA: Harnessing Robot States for Vision-Language and Action Alignment cites this paper.

ROSA: Harnessing Robot States for Vision-Language and Action Alignment MoLe-VLA: Dynamic Layer-skipping Vision Language Action Model via Mixture-of-Layers for Efficient Robot Manipulation

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:31.820504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:31.820504Z digest=sha256:52d97476902c0823fa2e69f0a14dd25e128559109d785668dddfc94b73e408c3

Observation dbda5eef-5153-4d3e-96b8-c310951c7fc8 · inbound

Large Model Empowered Embodied AI: A Survey on Decision-Making and Embodied Learning cites this paper.

Large Model Empowered Embodied AI: A Survey on Decision-Making and Embodied Learning MoLe-VLA: Dynamic Layer-skipping Vision Language Action Model via Mixture-of-Layers for Efficient Robot Manipulation

Reference 225

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:56.219219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:31:56.219219Z digest=sha256:fed07a85c5e2be3a795c4d64936258d7a804ce6d623211384eac4ec60f200801

Observation 43cff229-7ab5-4968-8f65-e7429a7f0a18 · inbound

Large VLM-based Vision-Language-Action Models for Robotic Manipulation: A Survey cites this paper.

Large VLM-based Vision-Language-Action Models for Robotic Manipulation: A Survey MoLe-VLA: Dynamic Layer-skipping Vision Language Action Model via Mixture-of-Layers for Efficient Robot Manipulation

Reference 114

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T20:28:16.234775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T20:28:15.818016Z digest=sha256:c006e50b9571212bef91ce0236f9ba324df4abe25217443b8551af5daa9a62cf

Observation 648c7342-4248-4cc0-8237-93756c7b40a4 · inbound

CogVLA: Cognition-Aligned Vision-Language-Action Model via Instruction-Driven Routing & Sparsification cites this paper.

CogVLA: Cognition-Aligned Vision-Language-Action Model via Instruction-Driven Routing & Sparsification MoLe-VLA: Dynamic Layer-skipping Vision Language Action Model via Mixture-of-Layers for Efficient Robot Manipulation

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-05T14:42:32.834882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:42:32.834882Z digest=sha256:dc09cb59f76a6e776a84344102fb72431302f25818a0b9e60d9d7ee47e61e9f5

Observation 0e1c2132-cd29-4bf8-b471-2df10c350181 · inbound

Align-Then-stEer: Adapting the Vision-Language Action Models through Unified Latent Guidance cites this paper.

Align-Then-stEer: Adapting the Vision-Language Action Models through Unified Latent Guidance MoLe-VLA: Dynamic Layer-skipping Vision Language Action Model via Mixture-of-Layers for Efficient Robot Manipulation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-05T12:02:28.319138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:02:28.319138Z digest=sha256:d79ac2bbee33fc48e366416cf2d5b34898e6450ea44e0da7d57b9c3b142dac27

Observation e264d137-52b2-49f5-a909-fb6c40e667f3 · inbound

Efficient Vision-Language-Action Models for Embodied Manipulation: A Systematic Survey cites this paper.

Efficient Vision-Language-Action Models for Embodied Manipulation: A Systematic Survey MoLe-VLA: Dynamic Layer-skipping Vision Language Action Model via Mixture-of-Layers for Efficient Robot Manipulation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-04T09:08:10.150950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:08:10.150950Z digest=sha256:991453c6760310f8168cec5357cdd177bcf746c9635154e27a876c317d6b27bd

Observation 9d5bd1c3-fafa-46c5-929f-a125e385d4ae · inbound

ActDistill: General Action-Guided Self-Derived Distillation for Efficient Vision-Language-Action Models cites this paper.

ActDistill: General Action-Guided Self-Derived Distillation for Efficient Vision-Language-Action Models MoLe-VLA: Dynamic Layer-skipping Vision Language Action Model via Mixture-of-Layers for Efficient Robot Manipulation

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-17T06:09:09.414448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T06:07:42.311608Z digest=sha256:55cd80fbd7a040be11ad685088f87127e308a4a0a468ca3a6b6494976e28b394

Observation d1649ef9-99ab-4c83-8cad-070952e4b976 · inbound

Continually Evolving Skill Knowledge in Vision Language Action Model cites this paper.

Continually Evolving Skill Knowledge in Vision Language Action Model MoLe-VLA: Dynamic Layer-skipping Vision Language Action Model via Mixture-of-Layers for Efficient Robot Manipulation

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-17T06:04:09.148437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T06:02:31.638120Z digest=sha256:4e10ec3a0bc43ee60817e6217176a964d288333de78e181cbefa2948d57d278a

Observation 54e94133-9306-4a81-9099-c7a388037d25 · inbound

TIDAL: Temporally Interleaved Diffusion and Action Loop for High-Frequency VLA Control cites this paper.

TIDAL: Temporally Interleaved Diffusion and Action Loop for High-Frequency VLA Control MoLe-VLA: Dynamic Layer-skipping Vision Language Action Model via Mixture-of-Layers for Efficient Robot Manipulation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-03T09:07:40.908773Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:07:40.908773Z digest=sha256:6fa1680f5e8f8d551ffa1ddbdc8d360f72e3b9bfe90693c7f5ca27e0065c0385

Observation b3c46f3f-6801-4a6a-a226-9396b404342d · inbound

QuantVLA: Scale-Calibrated Post-Training Quantization for Vision-Language-Action Models cites this paper.

QuantVLA: Scale-Calibrated Post-Training Quantization for Vision-Language-Action Models MoLe-VLA: Dynamic Layer-skipping Vision Language Action Model via Mixture-of-Layers for Efficient Robot Manipulation

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:20:17.369582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T20:20:10.435886Z digest=sha256:df99cd65f5c6373628ff7972c9762a60cac482fe64fbc8c518629d9c9f2fb0c9

Observation ff03193e-5367-4280-8aa8-b112ab5cfea7 · inbound

KERV: Kinematic-Rectified Speculative Decoding for Embodied VLA Models cites this paper.

KERV: Kinematic-Rectified Speculative Decoding for Embodied VLA Models MoLe-VLA: Dynamic Layer-skipping Vision Language Action Model via Mixture-of-Layers for Efficient Robot Manipulation

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-15T18:40:14.548333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T18:39:58.095112Z digest=sha256:81c1f73f2fc4ea32042a7907a4095a3d0801ec23aff27c4504037fcd64784d84

Observation b6d4d615-0113-4f1d-86a0-97987c6e12b1 · inbound

The Latent Color Subspace: Emergent Order in High-Dimensional Chaos cites this paper.

The Latent Color Subspace: Emergent Order in High-Dimensional Chaos MoLe-VLA: Dynamic Layer-skipping Vision Language Action Model via Mixture-of-Layers for Efficient Robot Manipulation

Reference 31

Resolution
unresolved
no resolver link, observed 2026-07-14T22:24:08.310224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T22:24:08.310224Z digest=sha256:dc7641e8622790dd5cb0bc31f8658672e50d9bc177d232d992470074f0445e18

Observation 56fc1280-524d-4639-b2fe-4b25a6d07616 · inbound

OxyGen: Unified KV Cache Management for VLA Inference under Multi-Task Parallelism cites this paper.

OxyGen: Unified KV Cache Management for VLA Inference under Multi-Task Parallelism MoLe-VLA: Dynamic Layer-skipping Vision Language Action Model via Mixture-of-Layers for Efficient Robot Manipulation

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-21T11:50:03.795413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T11:46:12.134869Z digest=sha256:b697ba4f20da10977f5e40d90a93e429247a47c278cd7ac7ab9d86988610c5c0

Observation 92a861db-b6ed-4abe-b07a-e6cbf71bbaeb · inbound

HeiSD: Hybrid Speculative Decoding for Embodied Vision-Language-Action Models with Kinematic Awareness cites this paper.

HeiSD: Hybrid Speculative Decoding for Embodied Vision-Language-Action Models with Kinematic Awareness MoLe-VLA: Dynamic Layer-skipping Vision Language Action Model via Mixture-of-Layers for Efficient Robot Manipulation

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-15T09:19:54.385906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T09:15:50.963123Z digest=sha256:8c589d234cb57773e6d1981c97f3a8b2c6b9a86555f12ffa80d8f667134a400b

Observation dc99f12e-66d8-486b-b2fc-75191ba4d7a0 · inbound

RoboECC: Multi-Factor-Aware Edge-Cloud Collaborative Deployment for VLA Models cites this paper.

RoboECC: Multi-Factor-Aware Edge-Cloud Collaborative Deployment for VLA Models MoLe-VLA: Dynamic Layer-skipping Vision Language Action Model via Mixture-of-Layers for Efficient Robot Manipulation

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T07:39:50.295294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T07:37:34.182285Z digest=sha256:f1e0667a4e48130891caac93c7c2a8fdb832608bf22d72cfb335a1561e956ba0

Observation b24c322c-e9a0-4f78-b45c-bd7d7c10fb82 · inbound

A1: A Fully Transparent Open-Source, Adaptive and Efficient Truncated Vision-Language-Action Model cites this paper.

A1: A Fully Transparent Open-Source, Adaptive and Efficient Truncated Vision-Language-Action Model MoLe-VLA: Dynamic Layer-skipping Vision Language Action Model via Mixture-of-Layers for Efficient Robot Manipulation

Reference 55

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T22:50:51.326127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T19:31:23.255452Z digest=sha256:c9b7ab924b1eb2d527459132971d9e30bbd282dfd2c7e37851c96ca075470881

Observation c52d2cea-cf37-4101-b497-6cd5ef54132f · inbound

ConsisVLA-4D: Advancing Spatiotemporal Consistency in Efficient 3D-Perception and 4D-Reasoning for Robotic Manipulation cites this paper.

ConsisVLA-4D: Advancing Spatiotemporal Consistency in Efficient 3D-Perception and 4D-Reasoning for Robotic Manipulation MoLe-VLA: Dynamic Layer-skipping Vision Language Action Model via Mixture-of-Layers for Efficient Robot Manipulation

Reference 85

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T18:06:06.378679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T16:40:19.057979Z digest=sha256:b34dcbf89bf4e1a92a955d3ec735ba424706de0dcd3adaec10251af71cc1c79e

Observation f5edacb1-c78f-48e2-bf86-7ecf19eb2920 · inbound

ElegantVLA: Learning When to Think for Efficient Vision-Language-Action Models cites this paper.

ElegantVLA: Learning When to Think for Efficient Vision-Language-Action Models MoLe-VLA: Dynamic Layer-skipping Vision Language Action Model via Mixture-of-Layers for Efficient Robot Manipulation

Reference 36

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T07:23:13.020544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T07:16:27.229603Z digest=sha256:98f205370c84eb926f0ed0bf922c062e8803482cff6f03b1d236a5b8f1a95f01

Observation e02b5006-061d-4e9d-ae8d-862a0840f7e9 · inbound

SAFE-Pruner: Semantic Attention-Guided Future-Aware Token Pruning for Efficient Vision-Language-Action Manipulation cites this paper.

SAFE-Pruner: Semantic Attention-Guided Future-Aware Token Pruning for Efficient Vision-Language-Action Manipulation MoLe-VLA: Dynamic Layer-skipping Vision Language Action Model via Mixture-of-Layers for Efficient Robot Manipulation

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-06-29T08:43:14.958648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T08:40:10.152344Z digest=sha256:e27a99399b15fada592e63c908b3260e13ceb9820795bedfaf3337d375a1cf73

Observation 8693073a-602d-45a2-b9df-7a1f2b9d3e3c · inbound

SAFE-Pruner: Semantic Attention-Guided Future-Aware Token Pruning for Efficient Vision-Language-Action Manipulation cites this paper.

SAFE-Pruner: Semantic Attention-Guided Future-Aware Token Pruning for Efficient Vision-Language-Action Manipulation MoLe-VLA: Dynamic Layer-skipping Vision Language Action Model via Mixture-of-Layers for Efficient Robot Manipulation

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-02T12:54:43.278929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T12:54:43.278929Z digest=sha256:3dd0d9b28e84a96a14bb27f288236152f2c981f18bb42449505fb92f4fbd8cd6

Observation 88aa6e0c-32d1-4b32-ac6c-c3ecd25f707b · inbound

World Pilot: Steering Vision-Language-Action Models with World-Action Priors cites this paper.

World Pilot: Steering Vision-Language-Action Models with World-Action Priors MoLe-VLA: Dynamic Layer-skipping Vision Language Action Model via Mixture-of-Layers for Efficient Robot Manipulation

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-03T11:18:03.369857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-27T09:40:02.137152Z digest=sha256:3dd7fbf03e0b83d97f89963670c6b4b82febb9aee860e72371361f30b55b51cc

Observation edade014-2e2b-4fda-9e2a-88d3288e1b49 · inbound

Revisiting Parameter Redundancy in Vision-Language-Action Models: Insights from VLM-to-VLA Adaptation cites this paper.

Revisiting Parameter Redundancy in Vision-Language-Action Models: Insights from VLM-to-VLA Adaptation MoLe-VLA: Dynamic Layer-skipping Vision Language Action Model via Mixture-of-Layers for Efficient Robot Manipulation

Reference 38

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T10:45:42.895181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-01T05:07:45.792132Z digest=sha256:90c44fba5ca1e5f142d1da61ca3d086238064b776fddc7243282760892d54fa3

Observation a14612d1-b8d3-4c1b-9c13-4cdad4a1e87a · inbound

ActionCache: Training-Free Acceleration for Vision-Language-Action Models with Action Caching and Refinement cites this paper.

ActionCache: Training-Free Acceleration for Vision-Language-Action Models with Action Caching and Refinement MoLe-VLA: Dynamic Layer-skipping Vision Language Action Model via Mixture-of-Layers for Efficient Robot Manipulation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T04:31:42.149217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T04:31:42.149217Z digest=sha256:77752a75bf721f0b27760e232fe12bf224789733861016612f7767ac2e50ba04