Pith. sign in

Paper Citation Record · LEDGER

MoLe-VLA: Dynamic Layer-skipping Vision Language Action Model via Mixture-of-Layers for Efficient Robot Manipulation

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 25 inbound Pith citation observations for arXiv:2503.20384.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2503.20384 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 25 of 25 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 25 of 25 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:44:06.343233Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T11:18:03.368562Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 6f11c451-0ad3-4ec4-acac-2e7d5d3fa5d5 · inbound

Towards a Generalizable Bimanual Foundation Policy via Flow-based Video Prediction cites this paper.

Towards a Generalizable Bimanual Foundation Policy via Flow-based Video Prediction MoLe-VLA: Dynamic Layer-skipping Vision Language Action Model via Mixture-of-Layers for Efficient Robot Manipulation

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:06.343233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:06.343233Z digest=sha256:af965bfdd3240e4c1662b0b43742812f7d404a30aaf61419cb1abc189938b66b

Observation 25aa3d8d-ae22-4f6a-993b-bf3401d4953a · inbound

EfficientVLA: Training-Free Acceleration and Compression for Vision-Language-Action Models cites this paper.

EfficientVLA: Training-Free Acceleration and Compression for Vision-Language-Action Models MoLe-VLA: Dynamic Layer-skipping Vision Language Action Model via Mixture-of-Layers for Efficient Robot Manipulation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T04:42:22.434287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:42:22.434287Z digest=sha256:cc96816924486916745aedb0dd9fbd77fee1db823204481cc1912083858fc801

Observation fbe82f9c-440f-4a16-9397-b85d4f3eb02a · inbound

ROSA: Harnessing Robot States for Vision-Language and Action Alignment cites this paper.

ROSA: Harnessing Robot States for Vision-Language and Action Alignment MoLe-VLA: Dynamic Layer-skipping Vision Language Action Model via Mixture-of-Layers for Efficient Robot Manipulation

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:31.820504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:31.820504Z digest=sha256:9a3567b3142235457090c3db2e17649ee6a457590087dfc17a1e6b66062bd614

Observation dbda5eef-5153-4d3e-96b8-c310951c7fc8 · inbound

Large Model Empowered Embodied AI: A Survey on Decision-Making and Embodied Learning cites this paper.

Large Model Empowered Embodied AI: A Survey on Decision-Making and Embodied Learning MoLe-VLA: Dynamic Layer-skipping Vision Language Action Model via Mixture-of-Layers for Efficient Robot Manipulation

Reference 225

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:56.219219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:31:56.219219Z digest=sha256:378e5e9a04e0605f89339338202ae564f4398fdbf1d2823b622a059d4f9e2d0f

Observation 43cff229-7ab5-4968-8f65-e7429a7f0a18 · inbound

Large VLM-based Vision-Language-Action Models for Robotic Manipulation: A Survey cites this paper.

Large VLM-based Vision-Language-Action Models for Robotic Manipulation: A Survey MoLe-VLA: Dynamic Layer-skipping Vision Language Action Model via Mixture-of-Layers for Efficient Robot Manipulation

Reference 114

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T20:28:16.234775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-17T20:28:15.818016Z digest=sha256:c6151331eb579dc4b74c8b13ac672e102c95b5c7956c7c37302cba6bd20edd7e

Observation 648c7342-4248-4cc0-8237-93756c7b40a4 · inbound

CogVLA: Cognition-Aligned Vision-Language-Action Model via Instruction-Driven Routing & Sparsification cites this paper.

CogVLA: Cognition-Aligned Vision-Language-Action Model via Instruction-Driven Routing & Sparsification MoLe-VLA: Dynamic Layer-skipping Vision Language Action Model via Mixture-of-Layers for Efficient Robot Manipulation

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-05T14:42:32.834882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:42:32.834882Z digest=sha256:dc09cb59f76a6e776a84344102fb72431302f25818a0b9e60d9d7ee47e61e9f5

Observation 0e1c2132-cd29-4bf8-b471-2df10c350181 · inbound

Align-Then-stEer: Adapting the Vision-Language Action Models through Unified Latent Guidance cites this paper.

Align-Then-stEer: Adapting the Vision-Language Action Models through Unified Latent Guidance MoLe-VLA: Dynamic Layer-skipping Vision Language Action Model via Mixture-of-Layers for Efficient Robot Manipulation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-05T12:02:28.319138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:02:28.319138Z digest=sha256:ab7e34b306963e7a55f8b54e91b21b32135a2ddc96988992531ceb967a99625b

Observation e264d137-52b2-49f5-a909-fb6c40e667f3 · inbound

Efficient Vision-Language-Action Models for Embodied Manipulation: A Systematic Survey cites this paper.

Efficient Vision-Language-Action Models for Embodied Manipulation: A Systematic Survey MoLe-VLA: Dynamic Layer-skipping Vision Language Action Model via Mixture-of-Layers for Efficient Robot Manipulation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-04T09:08:10.150950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:08:10.150950Z digest=sha256:991453c6760310f8168cec5357cdd177bcf746c9635154e27a876c317d6b27bd

Observation 9d5bd1c3-fafa-46c5-929f-a125e385d4ae · inbound

ActDistill: General Action-Guided Self-Derived Distillation for Efficient Vision-Language-Action Models cites this paper.

ActDistill: General Action-Guided Self-Derived Distillation for Efficient Vision-Language-Action Models MoLe-VLA: Dynamic Layer-skipping Vision Language Action Model via Mixture-of-Layers for Efficient Robot Manipulation

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-17T06:09:09.414448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-17T06:07:42.311608Z digest=sha256:2ef21134245a2197aecf832e4fadbbda623662fbc7b73bbc909328833057600e

Observation d1649ef9-99ab-4c83-8cad-070952e4b976 · inbound

Continually Evolving Skill Knowledge in Vision Language Action Model cites this paper.

Continually Evolving Skill Knowledge in Vision Language Action Model MoLe-VLA: Dynamic Layer-skipping Vision Language Action Model via Mixture-of-Layers for Efficient Robot Manipulation

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-17T06:04:09.148437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-17T06:02:31.638120Z digest=sha256:9168af2abf1001dd19a0ae7053ac52e61e3cabb946edcccc5be55932924de50a

Observation 54e94133-9306-4a81-9099-c7a388037d25 · inbound

TIDAL: Temporally Interleaved Diffusion and Action Loop for High-Frequency VLA Control cites this paper.

TIDAL: Temporally Interleaved Diffusion and Action Loop for High-Frequency VLA Control MoLe-VLA: Dynamic Layer-skipping Vision Language Action Model via Mixture-of-Layers for Efficient Robot Manipulation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-03T09:07:40.908773Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:07:40.908773Z digest=sha256:6fa1680f5e8f8d551ffa1ddbdc8d360f72e3b9bfe90693c7f5ca27e0065c0385

Observation b3c46f3f-6801-4a6a-a226-9396b404342d · inbound

QuantVLA: Scale-Calibrated Post-Training Quantization for Vision-Language-Action Models cites this paper.

QuantVLA: Scale-Calibrated Post-Training Quantization for Vision-Language-Action Models MoLe-VLA: Dynamic Layer-skipping Vision Language Action Model via Mixture-of-Layers for Efficient Robot Manipulation

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:20:17.369582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T20:20:10.435886Z digest=sha256:9352e668b016b939b0581c3b400d57899f607955514b9a76ee5c946e395d1934

Observation ff03193e-5367-4280-8aa8-b112ab5cfea7 · inbound

KERV: Kinematic-Rectified Speculative Decoding for Embodied VLA Models cites this paper.

KERV: Kinematic-Rectified Speculative Decoding for Embodied VLA Models MoLe-VLA: Dynamic Layer-skipping Vision Language Action Model via Mixture-of-Layers for Efficient Robot Manipulation

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-15T18:40:14.548333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T18:39:58.095112Z digest=sha256:79307f20b404ef6c892fa770c54e06392fbca35988dd8c0c10387ab73da21f5b

Observation b6d4d615-0113-4f1d-86a0-97987c6e12b1 · inbound

The Latent Color Subspace: Emergent Order in High-Dimensional Chaos cites this paper.

The Latent Color Subspace: Emergent Order in High-Dimensional Chaos MoLe-VLA: Dynamic Layer-skipping Vision Language Action Model via Mixture-of-Layers for Efficient Robot Manipulation

Reference 31

Resolution
unresolved
no resolver link, observed 2026-07-14T22:24:08.310224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T22:24:08.310224Z digest=sha256:dc7641e8622790dd5cb0bc31f8658672e50d9bc177d232d992470074f0445e18

Observation 56fc1280-524d-4639-b2fe-4b25a6d07616 · inbound

OxyGen: Unified KV Cache Management for VLA Inference under Multi-Task Parallelism cites this paper.

OxyGen: Unified KV Cache Management for VLA Inference under Multi-Task Parallelism MoLe-VLA: Dynamic Layer-skipping Vision Language Action Model via Mixture-of-Layers for Efficient Robot Manipulation

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-21T11:50:03.795413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-21T11:46:12.134869Z digest=sha256:26cd5c556d961536ee3046864ee05579495adb58512bd8552e7821da19229cc5

Observation 92a861db-b6ed-4abe-b07a-e6cbf71bbaeb · inbound

HeiSD: Hybrid Speculative Decoding for Embodied Vision-Language-Action Models with Kinematic Awareness cites this paper.

HeiSD: Hybrid Speculative Decoding for Embodied Vision-Language-Action Models with Kinematic Awareness MoLe-VLA: Dynamic Layer-skipping Vision Language Action Model via Mixture-of-Layers for Efficient Robot Manipulation

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-15T09:19:54.385906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T09:15:50.963123Z digest=sha256:d6f764439de142ae6967822229d15a7aee630a00434401e8f9dc2fdc8b4eecad

Observation dc99f12e-66d8-486b-b2fc-75191ba4d7a0 · inbound

RoboECC: Multi-Factor-Aware Edge-Cloud Collaborative Deployment for VLA Models cites this paper.

RoboECC: Multi-Factor-Aware Edge-Cloud Collaborative Deployment for VLA Models MoLe-VLA: Dynamic Layer-skipping Vision Language Action Model via Mixture-of-Layers for Efficient Robot Manipulation

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T07:39:50.295294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T07:37:34.182285Z digest=sha256:c4469e481273c224d1a544f60ffd49bdce88fc69f97900f8a3c685c76e3084e7

Observation b24c322c-e9a0-4f78-b45c-bd7d7c10fb82 · inbound

A1: A Fully Transparent Open-Source, Adaptive and Efficient Truncated Vision-Language-Action Model cites this paper.

A1: A Fully Transparent Open-Source, Adaptive and Efficient Truncated Vision-Language-Action Model MoLe-VLA: Dynamic Layer-skipping Vision Language Action Model via Mixture-of-Layers for Efficient Robot Manipulation

Reference 55

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T22:50:51.326127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T19:31:23.255452Z digest=sha256:2a56d573e2fae15e120ab1293963e05177bdea1edb0aeda77cc5d2e3dea72aa4

Observation c52d2cea-cf37-4101-b497-6cd5ef54132f · inbound

ConsisVLA-4D: Advancing Spatiotemporal Consistency in Efficient 3D-Perception and 4D-Reasoning for Robotic Manipulation cites this paper.

ConsisVLA-4D: Advancing Spatiotemporal Consistency in Efficient 3D-Perception and 4D-Reasoning for Robotic Manipulation MoLe-VLA: Dynamic Layer-skipping Vision Language Action Model via Mixture-of-Layers for Efficient Robot Manipulation

Reference 85

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T18:06:06.378679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T16:40:19.057979Z digest=sha256:080232bd984d76dba3b683dbbb423c32e69d6839df6ca6677daf9200318b6a1a

Observation f5edacb1-c78f-48e2-bf86-7ecf19eb2920 · inbound

ElegantVLA: Learning When to Think for Efficient Vision-Language-Action Models cites this paper.

ElegantVLA: Learning When to Think for Efficient Vision-Language-Action Models MoLe-VLA: Dynamic Layer-skipping Vision Language Action Model via Mixture-of-Layers for Efficient Robot Manipulation

Reference 36

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T07:23:13.020544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-29T07:16:27.229603Z digest=sha256:17b608c263df5a927f02fa5591377f53fd32d60a14bd61e74d3a584428b3f401

Observation e02b5006-061d-4e9d-ae8d-862a0840f7e9 · inbound

SAFE-Pruner: Semantic Attention-Guided Future-Aware Token Pruning for Efficient Vision-Language-Action Manipulation cites this paper.

SAFE-Pruner: Semantic Attention-Guided Future-Aware Token Pruning for Efficient Vision-Language-Action Manipulation MoLe-VLA: Dynamic Layer-skipping Vision Language Action Model via Mixture-of-Layers for Efficient Robot Manipulation

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-06-29T08:43:14.958648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-29T08:40:10.152344Z digest=sha256:a798e0cf43202aa3e4b8e8698da0aee74094a8fa10687dbbe38b7f8bfb3fb3bf

Observation 8693073a-602d-45a2-b9df-7a1f2b9d3e3c · inbound

SAFE-Pruner: Semantic Attention-Guided Future-Aware Token Pruning for Efficient Vision-Language-Action Manipulation cites this paper.

SAFE-Pruner: Semantic Attention-Guided Future-Aware Token Pruning for Efficient Vision-Language-Action Manipulation MoLe-VLA: Dynamic Layer-skipping Vision Language Action Model via Mixture-of-Layers for Efficient Robot Manipulation

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-02T12:54:43.278929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T12:54:43.278929Z digest=sha256:3dd0d9b28e84a96a14bb27f288236152f2c981f18bb42449505fb92f4fbd8cd6

Observation 88aa6e0c-32d1-4b32-ac6c-c3ecd25f707b · inbound

World Pilot: Steering Vision-Language-Action Models with World-Action Priors cites this paper.

World Pilot: Steering Vision-Language-Action Models with World-Action Priors MoLe-VLA: Dynamic Layer-skipping Vision Language Action Model via Mixture-of-Layers for Efficient Robot Manipulation

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-03T11:18:03.369857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T09:40:02.137152Z digest=sha256:a911aac790eaa8e094d34ebe2268ece3f77444f3c203e2abe7a160b2eb1179a9

Observation edade014-2e2b-4fda-9e2a-88d3288e1b49 · inbound

Revisiting Parameter Redundancy in Vision-Language-Action Models: Insights from VLM-to-VLA Adaptation cites this paper.

Revisiting Parameter Redundancy in Vision-Language-Action Models: Insights from VLM-to-VLA Adaptation MoLe-VLA: Dynamic Layer-skipping Vision Language Action Model via Mixture-of-Layers for Efficient Robot Manipulation

Reference 38

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T10:45:42.895181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-01T05:07:45.792132Z digest=sha256:6f5d49af5572353b5816ff189302e24796d10498dd970f23bb1a7ce31fc4571d

Observation a14612d1-b8d3-4c1b-9c13-4cdad4a1e87a · inbound

ActionCache: Training-Free Acceleration for Vision-Language-Action Models with Action Caching and Refinement cites this paper.

ActionCache: Training-Free Acceleration for Vision-Language-Action Models with Action Caching and Refinement MoLe-VLA: Dynamic Layer-skipping Vision Language Action Model via Mixture-of-Layers for Efficient Robot Manipulation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T04:31:42.149217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T04:31:42.149217Z digest=sha256:77752a75bf721f0b27760e232fe12bf224789733861016612f7767ac2e50ba04