Pith. sign in

Paper Citation Record · LEDGER

Structure-Aware Robust Fine-Tuning: Defending Vision-Language-Action Robots Against Physical Attention Hijacking

As of 9 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 0 inbound Pith citation observations for arXiv:2608.03231.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.03231 v1

Coverage vector

measured 29 of 29 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T23:06:14.759570Z

measured 29 of 29 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

29 of 29 outbound references displayed

  • verified exact0
  • verified fuzzy13
  • unresolved16
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a39f8970-2279-4ac0-88d7-1dc642e30583 · outbound

This paper cites Rt-2: Vision-language-action models transfer web knowledge to robotic control,.

Structure-Aware Robust Fine-Tuning: Defending Vision-Language-Action Robots Against Physical Attention Hijacking Rt-2: Vision-language-action models transfer web knowledge to robotic control,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T23:06:15.420535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T23:06:14.686176Z digest=sha256:567ee87495e63506ac7b6d4b29027e72bc5bc3df77de4d625e4c54e7af9a5ce6

Observation 998d6341-7a24-4fa0-b90d-22e3db3156df · outbound

This paper cites OpenVLA: An Open-Source Vision-Language-Action Model.

Structure-Aware Robust Fine-Tuning: Defending Vision-Language-Action Robots Against Physical Attention Hijacking OpenVLA: An Open-Source Vision-Language-Action Model

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-05T23:06:14.689326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:06:14.689326Z digest=sha256:59a1d99ae952b06c153829658aa2b3ca76bacc0e2dc7555c98c38c9536c96825

Observation 393e0d5c-73d9-480f-9eff-4f5853cb0402 · outbound

This paper cites Octo: An Open-Source Generalist Robot Policy.

Structure-Aware Robust Fine-Tuning: Defending Vision-Language-Action Robots Against Physical Attention Hijacking Octo: An Open-Source Generalist Robot Policy

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T23:06:14.692364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:06:14.692364Z digest=sha256:7360234665daee9ed7fb74fbbd5eeee1344580e36cfe1c9fb5d53cda631d6153

Observation 6e8d5955-b9a4-442f-9860-0194f51174a7 · outbound

This paper cites $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control.

Structure-Aware Robust Fine-Tuning: Defending Vision-Language-Action Robots Against Physical Attention Hijacking $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T23:06:14.695292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:06:14.695292Z digest=sha256:313e5550b01762319442b6aff691495b4095eee64dd400fc8035dd2d3a0c3c6f

Observation feee247a-bdd0-4aa3-9b6a-493696c6891a · outbound

This paper cites Adversarial Patch.

Structure-Aware Robust Fine-Tuning: Defending Vision-Language-Action Robots Against Physical Attention Hijacking Adversarial Patch

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T23:06:14.698236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:06:14.698236Z digest=sha256:dc61f35f18f3be4802826ae7af4914809b62e0072c44a2a742e65b4a52381007

Observation 19b1536e-87ef-4e5c-b3da-a120db88d505 · outbound

This paper cites Synthesizing robust adversar- ial examples,.

Structure-Aware Robust Fine-Tuning: Defending Vision-Language-Action Robots Against Physical Attention Hijacking Synthesizing robust adversar- ial examples,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T23:06:15.413009Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T23:06:14.700973Z digest=sha256:6b479ba91d3ce54784e711da3716e039f9adbcdfa887d6817d3beae7374dc10d

Observation e97eb49e-5296-4fc2-844f-6a5f943a21bf · outbound

This paper cites Manipulation facing threats: Evaluating physical vulnerabilities in end-to-end vision language action models,.

Structure-Aware Robust Fine-Tuning: Defending Vision-Language-Action Robots Against Physical Attention Hijacking Manipulation facing threats: Evaluating physical vulnerabilities in end-to-end vision language action models,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T23:06:14.703646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:06:14.703646Z digest=sha256:9768813b8d4d0bf27a3af9ca10150c8098cb116116e3b82de8d60789bd8452c0

Observation 84f5fcfe-1e54-483a-a82b-c0e5f696e2dc · outbound

This paper cites Eva-vla: Evaluating vision-language- action models’ robustness under real-world physical variations,.

Structure-Aware Robust Fine-Tuning: Defending Vision-Language-Action Robots Against Physical Attention Hijacking Eva-vla: Evaluating vision-language- action models’ robustness under real-world physical variations,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T23:06:14.706033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:06:14.706033Z digest=sha256:51d3990eadc0fe1e5a232d0c5a91ba09e3a791b653b0338dc9cbf785ee8f1851

Observation 48426eb5-ee36-4468-84b0-a80486a00611 · outbound

This paper cites Exploring the adversarial vulnera- bilities of vision-language-action models in robotics,.

Structure-Aware Robust Fine-Tuning: Defending Vision-Language-Action Robots Against Physical Attention Hijacking Exploring the adversarial vulnera- bilities of vision-language-action models in robotics,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T23:06:15.405547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T23:06:14.708393Z digest=sha256:69c5b41db105537c046d326e628529b1f8d92295c9d2568f2a67e18c6a24bf58

Observation 807f5217-a23e-487b-a2ab-25ecb6b9fe89 · outbound

This paper cites Model-agnostic adversarial attack and defense for vision-language-action models,.

Structure-Aware Robust Fine-Tuning: Defending Vision-Language-Action Robots Against Physical Attention Hijacking Model-agnostic adversarial attack and defense for vision-language-action models,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T23:06:14.710688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:06:14.710688Z digest=sha256:e8ab7c7671d58db17d718f0c46e93b088d6fc194babe17111f221f058ce53e96

Observation 767a8bfc-1281-4899-bc4d-cf22ffa671d7 · outbound

This paper cites Libero: Benchmarking knowledge transfer for lifelong robot learning,.

Structure-Aware Robust Fine-Tuning: Defending Vision-Language-Action Robots Against Physical Attention Hijacking Libero: Benchmarking knowledge transfer for lifelong robot learning,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T23:06:15.398053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T23:06:14.712967Z digest=sha256:6734752776c571eda9619d19e17dc4075504a52bfe874dc8642ad5bd7ad85471

Observation a5a8ebe4-eecb-4940-a540-487cff11fba4 · outbound

This paper cites Diffusion policy: Visuomotor policy learning via action diffusion,.

Structure-Aware Robust Fine-Tuning: Defending Vision-Language-Action Robots Against Physical Attention Hijacking Diffusion policy: Visuomotor policy learning via action diffusion,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T23:06:15.390459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T23:06:14.715441Z digest=sha256:3e841819eedf663a4bdcd911a7c4f190de58b05ed7d2e97e97194a75f4693839

Observation 284defe1-7ffa-4eb0-a72c-761eb9a53b6c · outbound

This paper cites Mft: Modal fusion transformer for cross- modal fusion in 3d object detection,.

Structure-Aware Robust Fine-Tuning: Defending Vision-Language-Action Robots Against Physical Attention Hijacking Mft: Modal fusion transformer for cross- modal fusion in 3d object detection,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T23:06:15.383088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T23:06:14.717800Z digest=sha256:5e641275eada5f3919964c23cbc322bbb17dd7210c0c83e7d4505ecc4daac794

Observation afc7abe0-8fef-4768-ae24-61d97a120988 · outbound

This paper cites Dstr: Dual scenes transformer for cross-modal fusion in 3d object detection,.

Structure-Aware Robust Fine-Tuning: Defending Vision-Language-Action Robots Against Physical Attention Hijacking Dstr: Dual scenes transformer for cross-modal fusion in 3d object detection,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T23:06:15.375780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T23:06:14.720082Z digest=sha256:841ccd72f6c682b69e65ee901006c8c412d728a6e974309e128eb1b872500720

Observation 8ba9e4b6-88b7-46fc-becf-7f9867d48c38 · outbound

This paper cites When robots obey the patch: Universal transferable patch attacks on vision-language-action models,.

Structure-Aware Robust Fine-Tuning: Defending Vision-Language-Action Robots Against Physical Attention Hijacking When robots obey the patch: Universal transferable patch attacks on vision-language-action models,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T23:06:14.722301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:06:14.722301Z digest=sha256:12f7714fd95ccb7905f2283c792d8f56fe4bf7d36a19baaa3f49d8610190a63f

Observation 0b3525f3-c5d2-4f4f-91fa-a6e4f217afe1 · outbound

This paper cites Attention-guided patch-wise sparse adversarial attacks on vision-language-action models,.

Structure-Aware Robust Fine-Tuning: Defending Vision-Language-Action Robots Against Physical Attention Hijacking Attention-guided patch-wise sparse adversarial attacks on vision-language-action models,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T23:06:14.724813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:06:14.724813Z digest=sha256:c1578d2f7806dfd32345dc65c7f9e8187b683e60c92351d21eb5e50e5de27048

Observation 49d8984b-36cb-4172-baf8-0950d6a21ae8 · outbound

This paper cites Attackvla: Benchmarking adversarial and backdoor attacks on vision-language-action models,.

Structure-Aware Robust Fine-Tuning: Defending Vision-Language-Action Robots Against Physical Attention Hijacking Attackvla: Benchmarking adversarial and backdoor attacks on vision-language-action models,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T23:06:14.727346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:06:14.727346Z digest=sha256:f798e5bb79c1110b2deb14aa40273f03f081037b4bed3686a285f8db5e47374e

Observation d8f99bcd-fea5-4bf4-b756-0e9d0583d466 · outbound

This paper cites Towards Deep Learning Models Resistant to Adversarial Attacks.

Structure-Aware Robust Fine-Tuning: Defending Vision-Language-Action Robots Against Physical Attention Hijacking Towards Deep Learning Models Resistant to Adversarial Attacks

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T23:06:14.729891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:06:14.729891Z digest=sha256:213ef5b36b80b550d0ce0a695c59ca806093edf3bb46e5c4f4b15fec8124ef32

Observation 0ea06ddf-7872-4150-9c24-17e0baa8b7cc · outbound

This paper cites Theoretically principled trade-off be- tween robustness and accuracy,.

Structure-Aware Robust Fine-Tuning: Defending Vision-Language-Action Robots Against Physical Attention Hijacking Theoretically principled trade-off be- tween robustness and accuracy,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T23:06:15.368012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T23:06:14.732865Z digest=sha256:d50a104a53dca50ee4d07622191ecb7406feb28527eed6917e1d055b4a5c6936

Observation 9a1ced44-9188-45e3-bd27-43a297ea07ab · outbound

This paper cites Domain randomization for transferring deep neural networks from simulation to the real world,.

Structure-Aware Robust Fine-Tuning: Defending Vision-Language-Action Robots Against Physical Attention Hijacking Domain randomization for transferring deep neural networks from simulation to the real world,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T23:06:15.359352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T23:06:14.735473Z digest=sha256:fc5a3b0f4e52f9a4a616fba7920270830186df3b89496848f378c19339e6cc90

Observation 6c8ae63e-61ba-4009-9244-e7d4e2947e81 · outbound

This paper cites Diffusion Models for Adversarial Purification.

Structure-Aware Robust Fine-Tuning: Defending Vision-Language-Action Robots Against Physical Attention Hijacking Diffusion Models for Adversarial Purification

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T23:06:14.738049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:06:14.738049Z digest=sha256:20fca9af777447d961bb54c0e1ac8ad833abbcd1332e9e00493d8a6d303d1420

Observation bc99b205-cf3f-4840-bf59-05acb970b94a · outbound

This paper cites Defense-GAN: Protecting Classifiers Against Adversarial Attacks Using Generative Models.

Structure-Aware Robust Fine-Tuning: Defending Vision-Language-Action Robots Against Physical Attention Hijacking Defense-GAN: Protecting Classifiers Against Adversarial Attacks Using Generative Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-05T23:06:14.740820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:06:14.740820Z digest=sha256:8e8bdde67662201159ed554b5f9c9522bf9a42b496baae1345bba8a7c6fc4e12

Observation ae3302a1-186a-4f09-a8ce-412f7a986002 · outbound

This paper cites PixelDefend: Leveraging Generative Models to Understand and Defend against Adversarial Examples.

Structure-Aware Robust Fine-Tuning: Defending Vision-Language-Action Robots Against Physical Attention Hijacking PixelDefend: Leveraging Generative Models to Understand and Defend against Adversarial Examples

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-05T23:06:14.743502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:06:14.743502Z digest=sha256:902eab9b6b7483dd589a586c263e7136f117c56a4e8b6735a6e5f9ddcca05121

Observation 5bc0a4bd-5d3a-4a00-9995-de5d41b1131d · outbound

This paper cites {PatchGuard}: A provably robust defense against adversarial patches via small receptive fields and masking,.

Structure-Aware Robust Fine-Tuning: Defending Vision-Language-Action Robots Against Physical Attention Hijacking {PatchGuard}: A provably robust defense against adversarial patches via small receptive fields and masking,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T23:06:15.351567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T23:06:14.746180Z digest=sha256:0c379d03af3d033557e07f92ca2e521944e8de811da2cc98c1a09f17884f5073

Observation 57c8423d-e9ca-4fa9-a846-a72f802c3c1e · outbound

This paper cites Obfuscated gradients give a false sense of security: Circumventing defenses to adversarial examples,.

Structure-Aware Robust Fine-Tuning: Defending Vision-Language-Action Robots Against Physical Attention Hijacking Obfuscated gradients give a false sense of security: Circumventing defenses to adversarial examples,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T23:06:15.343797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T23:06:14.748858Z digest=sha256:a94e1cd1c4ccde1253412fa3aaa2d71cfeb6340c4ac86a39077a750ae46b04d7

Observation e5709ec9-7152-43a4-a1d5-2762132b9330 · outbound

This paper cites Overcoming catas- trophic forgetting in neural networks,.

Structure-Aware Robust Fine-Tuning: Defending Vision-Language-Action Robots Against Physical Attention Hijacking Overcoming catas- trophic forgetting in neural networks,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T23:06:15.336236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T23:06:14.751347Z digest=sha256:116a4f4ba9ae24c27d3d0565e91544a9e8ddf9c59294a56a097acf1a21a2678c

Observation c5b97522-3c4f-4842-9b6a-4a4c77fb37d9 · outbound

This paper cites Learning without forgetting,.

Structure-Aware Robust Fine-Tuning: Defending Vision-Language-Action Robots Against Physical Attention Hijacking Learning without forgetting,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T23:06:15.326827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T23:06:14.754003Z digest=sha256:7a4e56b14c1dd2472fc7090fb11a2ecaa9ac0ef990d6a5f38b6f9f31e2ac7652

Observation 9d766829-16c6-49ec-a2eb-a53a516646af · outbound

This paper cites Paying More Attention to Attention: Improving the Performance of Convolutional Neural Networks via Attention Transfer.

Structure-Aware Robust Fine-Tuning: Defending Vision-Language-Action Robots Against Physical Attention Hijacking Paying More Attention to Attention: Improving the Performance of Convolutional Neural Networks via Attention Transfer

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-05T23:06:14.756828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:06:14.756828Z digest=sha256:6a00ebae86e558837c03ff900d7e01dda159eec4a622f9d944284e6fe04f6907

Observation 8be20555-9bce-4205-b18b-922824e007bb · outbound

This paper cites Robust finetuning of vision- language-action robot policies via parameter merging,.

Structure-Aware Robust Fine-Tuning: Defending Vision-Language-Action Robots Against Physical Attention Hijacking Robust finetuning of vision- language-action robot policies via parameter merging,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-05T23:06:14.759570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:06:14.759570Z digest=sha256:afa537b255f4dc189d8ecd72b9833219b26ed6a65e007f93fc05d48f711222e4

Pith citing papers

No inbound Pith citation observations are available.