Pith. sign in

Paper Citation Record · LEDGER

Structure-Aware Robust Fine-Tuning: Defending Vision-Language-Action Robots Against Physical Attention Hijacking

As of 10 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 0 inbound Pith citation observations for arXiv:2608.03231.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.03231 v1

Coverage vector

measured 29 of 29 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T23:06:14.759570Z

measured 29 of 29 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

29 of 29 outbound references displayed

  • verified exact0
  • verified fuzzy13
  • unresolved16
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a39f8970-2279-4ac0-88d7-1dc642e30583 · outbound

This paper cites Rt-2: Vision-language-action models transfer web knowledge to robotic control,.

Structure-Aware Robust Fine-Tuning: Defending Vision-Language-Action Robots Against Physical Attention Hijacking Rt-2: Vision-language-action models transfer web knowledge to robotic control,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T23:06:15.420535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T23:06:14.686176Z digest=sha256:aa5e02417aadb6e1e8eb3c39cce538daf280eb02999c80dd5fbaa937289a4902

Observation 998d6341-7a24-4fa0-b90d-22e3db3156df · outbound

This paper cites OpenVLA: An Open-Source Vision-Language-Action Model.

Structure-Aware Robust Fine-Tuning: Defending Vision-Language-Action Robots Against Physical Attention Hijacking OpenVLA: An Open-Source Vision-Language-Action Model

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-05T23:06:14.689326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:06:14.689326Z digest=sha256:9b4bc7a9af4601f035cc182338bf6c250f0411e7c25154c9e4df28a998f60e98

Observation 393e0d5c-73d9-480f-9eff-4f5853cb0402 · outbound

This paper cites Octo: An Open-Source Generalist Robot Policy.

Structure-Aware Robust Fine-Tuning: Defending Vision-Language-Action Robots Against Physical Attention Hijacking Octo: An Open-Source Generalist Robot Policy

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T23:06:14.692364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:06:14.692364Z digest=sha256:b7441af562192905030c0cffb9ce27861ab5f2ae48812631ac23595917b0b067

Observation 6e8d5955-b9a4-442f-9860-0194f51174a7 · outbound

This paper cites $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control.

Structure-Aware Robust Fine-Tuning: Defending Vision-Language-Action Robots Against Physical Attention Hijacking $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T23:06:14.695292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:06:14.695292Z digest=sha256:9a1ea37f72cf10b66fa8c54f5d77d37e001263ffbe68a054212d1e69ea91cbdd

Observation feee247a-bdd0-4aa3-9b6a-493696c6891a · outbound

This paper cites Adversarial Patch.

Structure-Aware Robust Fine-Tuning: Defending Vision-Language-Action Robots Against Physical Attention Hijacking Adversarial Patch

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T23:06:14.698236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:06:14.698236Z digest=sha256:39a218730949a1d26d447f1cf1f3c6bad6e2a6a093c1765f1075d131588a658d

Observation 19b1536e-87ef-4e5c-b3da-a120db88d505 · outbound

This paper cites Synthesizing robust adversar- ial examples,.

Structure-Aware Robust Fine-Tuning: Defending Vision-Language-Action Robots Against Physical Attention Hijacking Synthesizing robust adversar- ial examples,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T23:06:15.413009Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T23:06:14.700973Z digest=sha256:bcc67ad5a9902baafbabb64e9d8ef8b578c4e7d37fa0852d937ac87f3867df92

Observation e97eb49e-5296-4fc2-844f-6a5f943a21bf · outbound

This paper cites Manipulation facing threats: Evaluating physical vulnerabilities in end-to-end vision language action models,.

Structure-Aware Robust Fine-Tuning: Defending Vision-Language-Action Robots Against Physical Attention Hijacking Manipulation facing threats: Evaluating physical vulnerabilities in end-to-end vision language action models,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T23:06:14.703646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:06:14.703646Z digest=sha256:7896b8d6da6e4af2bcf75e42acc2cff38410a614fd1b647daaf3cfe44e68aa79

Observation 84f5fcfe-1e54-483a-a82b-c0e5f696e2dc · outbound

This paper cites Eva-vla: Evaluating vision-language- action models’ robustness under real-world physical variations,.

Structure-Aware Robust Fine-Tuning: Defending Vision-Language-Action Robots Against Physical Attention Hijacking Eva-vla: Evaluating vision-language- action models’ robustness under real-world physical variations,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T23:06:14.706033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:06:14.706033Z digest=sha256:0790c24ef31ad751b20f1af0247a8323b5102d37aedeb21bfe33eceed6eb737e

Observation 48426eb5-ee36-4468-84b0-a80486a00611 · outbound

This paper cites Exploring the adversarial vulnera- bilities of vision-language-action models in robotics,.

Structure-Aware Robust Fine-Tuning: Defending Vision-Language-Action Robots Against Physical Attention Hijacking Exploring the adversarial vulnera- bilities of vision-language-action models in robotics,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T23:06:15.405547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T23:06:14.708393Z digest=sha256:20c3ca7550c8a97cbffa941cf82ebf38819ddcfb77f3fc4d1f4e223b5dc80756

Observation 807f5217-a23e-487b-a2ab-25ecb6b9fe89 · outbound

This paper cites Model-agnostic adversarial attack and defense for vision-language-action models,.

Structure-Aware Robust Fine-Tuning: Defending Vision-Language-Action Robots Against Physical Attention Hijacking Model-agnostic adversarial attack and defense for vision-language-action models,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T23:06:14.710688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:06:14.710688Z digest=sha256:40a24285942c8f9b2c74a7fe6f44a30182e1b96885b10a9a0d102daca7fe6ef9

Observation 767a8bfc-1281-4899-bc4d-cf22ffa671d7 · outbound

This paper cites Libero: Benchmarking knowledge transfer for lifelong robot learning,.

Structure-Aware Robust Fine-Tuning: Defending Vision-Language-Action Robots Against Physical Attention Hijacking Libero: Benchmarking knowledge transfer for lifelong robot learning,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T23:06:15.398053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T23:06:14.712967Z digest=sha256:58f9d0669ce76c9f565367fb5f73d956de3968cf0661871affde6ac1fca77318

Observation a5a8ebe4-eecb-4940-a540-487cff11fba4 · outbound

This paper cites Diffusion policy: Visuomotor policy learning via action diffusion,.

Structure-Aware Robust Fine-Tuning: Defending Vision-Language-Action Robots Against Physical Attention Hijacking Diffusion policy: Visuomotor policy learning via action diffusion,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T23:06:15.390459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T23:06:14.715441Z digest=sha256:562289a9390c248b2fa7d13aaf1ae8a22c33023d9b0fb509f5d20710008f4c96

Observation 284defe1-7ffa-4eb0-a72c-761eb9a53b6c · outbound

This paper cites Mft: Modal fusion transformer for cross- modal fusion in 3d object detection,.

Structure-Aware Robust Fine-Tuning: Defending Vision-Language-Action Robots Against Physical Attention Hijacking Mft: Modal fusion transformer for cross- modal fusion in 3d object detection,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T23:06:15.383088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T23:06:14.717800Z digest=sha256:b672144cc98979602025b6f5e88834c46e4a74083de75af8a6dd26a8d4cf3061

Observation afc7abe0-8fef-4768-ae24-61d97a120988 · outbound

This paper cites Dstr: Dual scenes transformer for cross-modal fusion in 3d object detection,.

Structure-Aware Robust Fine-Tuning: Defending Vision-Language-Action Robots Against Physical Attention Hijacking Dstr: Dual scenes transformer for cross-modal fusion in 3d object detection,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T23:06:15.375780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T23:06:14.720082Z digest=sha256:96cc3435a2db513151c288e787d069a78f03f2afd06b8de55a2902b48a077479

Observation 8ba9e4b6-88b7-46fc-becf-7f9867d48c38 · outbound

This paper cites When robots obey the patch: Universal transferable patch attacks on vision-language-action models,.

Structure-Aware Robust Fine-Tuning: Defending Vision-Language-Action Robots Against Physical Attention Hijacking When robots obey the patch: Universal transferable patch attacks on vision-language-action models,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T23:06:14.722301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:06:14.722301Z digest=sha256:c6b7359d03e78acaba738751bf4d53707f580f810e4138c108d1f1c7fda5c3b4

Observation 0b3525f3-c5d2-4f4f-91fa-a6e4f217afe1 · outbound

This paper cites Attention-guided patch-wise sparse adversarial attacks on vision-language-action models,.

Structure-Aware Robust Fine-Tuning: Defending Vision-Language-Action Robots Against Physical Attention Hijacking Attention-guided patch-wise sparse adversarial attacks on vision-language-action models,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T23:06:14.724813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:06:14.724813Z digest=sha256:1e7ed7822ee3c13ea8a80268403599f1b124da3b991473a3097931bdf22a44b3

Observation 49d8984b-36cb-4172-baf8-0950d6a21ae8 · outbound

This paper cites Attackvla: Benchmarking adversarial and backdoor attacks on vision-language-action models,.

Structure-Aware Robust Fine-Tuning: Defending Vision-Language-Action Robots Against Physical Attention Hijacking Attackvla: Benchmarking adversarial and backdoor attacks on vision-language-action models,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T23:06:14.727346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:06:14.727346Z digest=sha256:08a30c2e3efb3b5e25011f14aee9fdbbdaee55158ca5f3c8088d697ed926080a

Observation d8f99bcd-fea5-4bf4-b756-0e9d0583d466 · outbound

This paper cites Towards Deep Learning Models Resistant to Adversarial Attacks.

Structure-Aware Robust Fine-Tuning: Defending Vision-Language-Action Robots Against Physical Attention Hijacking Towards Deep Learning Models Resistant to Adversarial Attacks

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T23:06:14.729891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:06:14.729891Z digest=sha256:60e8cf9e18aef794d6244c6d93eb00d4d47697419ed7e769fad1ce90b57a848d

Observation 0ea06ddf-7872-4150-9c24-17e0baa8b7cc · outbound

This paper cites Theoretically principled trade-off be- tween robustness and accuracy,.

Structure-Aware Robust Fine-Tuning: Defending Vision-Language-Action Robots Against Physical Attention Hijacking Theoretically principled trade-off be- tween robustness and accuracy,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T23:06:15.368012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T23:06:14.732865Z digest=sha256:08f0a104f58714c6ef0508b555439713ef733c62ba68127581ae2bec2d3c4dcb

Observation 9a1ced44-9188-45e3-bd27-43a297ea07ab · outbound

This paper cites Domain randomization for transferring deep neural networks from simulation to the real world,.

Structure-Aware Robust Fine-Tuning: Defending Vision-Language-Action Robots Against Physical Attention Hijacking Domain randomization for transferring deep neural networks from simulation to the real world,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T23:06:15.359352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T23:06:14.735473Z digest=sha256:a73343f52be5bb9d4c3b4da02b6e9f2a3707e7c0fa2bd5838a31952b3d8290b8

Observation 6c8ae63e-61ba-4009-9244-e7d4e2947e81 · outbound

This paper cites Diffusion Models for Adversarial Purification.

Structure-Aware Robust Fine-Tuning: Defending Vision-Language-Action Robots Against Physical Attention Hijacking Diffusion Models for Adversarial Purification

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T23:06:14.738049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:06:14.738049Z digest=sha256:0afefaa5bf63fd0f3c54f48932cd6b84c02a806b48215c71a2646420eb4f7a7e

Observation bc99b205-cf3f-4840-bf59-05acb970b94a · outbound

This paper cites Defense-GAN: Protecting Classifiers Against Adversarial Attacks Using Generative Models.

Structure-Aware Robust Fine-Tuning: Defending Vision-Language-Action Robots Against Physical Attention Hijacking Defense-GAN: Protecting Classifiers Against Adversarial Attacks Using Generative Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-05T23:06:14.740820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:06:14.740820Z digest=sha256:b18c448d6386d66ccc8e2c84552b732145818eb1e835cf3efea3741fd9d5f060

Observation ae3302a1-186a-4f09-a8ce-412f7a986002 · outbound

This paper cites PixelDefend: Leveraging Generative Models to Understand and Defend against Adversarial Examples.

Structure-Aware Robust Fine-Tuning: Defending Vision-Language-Action Robots Against Physical Attention Hijacking PixelDefend: Leveraging Generative Models to Understand and Defend against Adversarial Examples

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-05T23:06:14.743502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:06:14.743502Z digest=sha256:50b5f6e10dfe64cf56bc12d16e76809e5c5b3303e1f9daffedd159d554aea4b5

Observation 5bc0a4bd-5d3a-4a00-9995-de5d41b1131d · outbound

This paper cites {PatchGuard}: A provably robust defense against adversarial patches via small receptive fields and masking,.

Structure-Aware Robust Fine-Tuning: Defending Vision-Language-Action Robots Against Physical Attention Hijacking {PatchGuard}: A provably robust defense against adversarial patches via small receptive fields and masking,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T23:06:15.351567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T23:06:14.746180Z digest=sha256:4018707d4eb81a0c418add6e25329c9385e8f6d03c6817105d75536bca8b2771

Observation 57c8423d-e9ca-4fa9-a846-a72f802c3c1e · outbound

This paper cites Obfuscated gradients give a false sense of security: Circumventing defenses to adversarial examples,.

Structure-Aware Robust Fine-Tuning: Defending Vision-Language-Action Robots Against Physical Attention Hijacking Obfuscated gradients give a false sense of security: Circumventing defenses to adversarial examples,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T23:06:15.343797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T23:06:14.748858Z digest=sha256:6888438951419d3ea50f7eaeac3a2b266b441b1b9e0432ace5f3395a4cb28d03

Observation e5709ec9-7152-43a4-a1d5-2762132b9330 · outbound

This paper cites Overcoming catas- trophic forgetting in neural networks,.

Structure-Aware Robust Fine-Tuning: Defending Vision-Language-Action Robots Against Physical Attention Hijacking Overcoming catas- trophic forgetting in neural networks,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T23:06:15.336236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T23:06:14.751347Z digest=sha256:2205954fc74ffc3aa8f51b4ca4b3fb9267252a300e1495d2cad4d82bbbc96f5f

Observation c5b97522-3c4f-4842-9b6a-4a4c77fb37d9 · outbound

This paper cites Learning without forgetting,.

Structure-Aware Robust Fine-Tuning: Defending Vision-Language-Action Robots Against Physical Attention Hijacking Learning without forgetting,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T23:06:15.326827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T23:06:14.754003Z digest=sha256:f12b44b41fc2b9ac5b17ce72877d945fee456c5519f0344a2c075293c91c472b

Observation 9d766829-16c6-49ec-a2eb-a53a516646af · outbound

This paper cites Paying More Attention to Attention: Improving the Performance of Convolutional Neural Networks via Attention Transfer.

Structure-Aware Robust Fine-Tuning: Defending Vision-Language-Action Robots Against Physical Attention Hijacking Paying More Attention to Attention: Improving the Performance of Convolutional Neural Networks via Attention Transfer

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-05T23:06:14.756828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:06:14.756828Z digest=sha256:652424139e9b116f15c5c8ebdf7bfa8b45775e0f3b276d3b09a14efbd51c5688

Observation 8be20555-9bce-4205-b18b-922824e007bb · outbound

This paper cites Robust finetuning of vision- language-action robot policies via parameter merging,.

Structure-Aware Robust Fine-Tuning: Defending Vision-Language-Action Robots Against Physical Attention Hijacking Robust finetuning of vision- language-action robot policies via parameter merging,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-05T23:06:14.759570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:06:14.759570Z digest=sha256:66406dc90fe50ccb8c098fb22d5e52bc42b8ee9575f7fdb1a7a151b61b6c37c3

Pith citing papers

No inbound Pith citation observations are available.