Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T04:33:43.018059Z
Paper Citation Record · LEDGER
As of 10 August 2026, this Paper Citation Record lists 59 of 59 outbound references and 0 inbound Pith citation observations for arXiv:2506.10416.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T04:33:43.018059Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
59 of 59 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 1eeec5c1-e129-4f26-8102-1c04208b468f · outbound
Can Sound Replace Vision in LLaVA With Token Substitution? Deep Variational Information Bottleneck
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 399d9b91-dab0-4331-bc52-2ecd06612ede · outbound
Can Sound Replace Vision in LLaVA With Token Substitution? Arandjelovic and P
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 25b4e1c7-bac9-4129-9302-39a680047730 · outbound
Can Sound Replace Vision in LLaVA With Token Substitution? Eagle: Egocentric aggregated language-video engine
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation f7582180-498f-4591-a8a5-08b2f53abd49 · outbound
Can Sound Replace Vision in LLaVA With Token Substitution? Vggsound: A large-scale audio-visual dataset
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 1ec82599-025e-4684-806a-cccee916717b · outbound
Can Sound Replace Vision in LLaVA With Token Substitution? Clap learning audio concepts from natural language supervision
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 0b3aa060-f582-4a32-afff-b3bce5803e62 · outbound
Can Sound Replace Vision in LLaVA With Token Substitution? Imagebind: One embedding space to bind them all
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation f8ce3aa5-20d5-45e3-874e-87172bf9c4c6 · outbound
Can Sound Replace Vision in LLaVA With Token Substitution? Audioset, 2017
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 2fb0fe18-772e-4e59-9f9d-e2cb2b8b7012 · outbound
Can Sound Replace Vision in LLaVA With Token Substitution? Ego4d: Around the world in 3,000 hours of egocentric video
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 688d5303-08a6-4b08-8916-95101b4fe48c · outbound
Can Sound Replace Vision in LLaVA With Token Substitution? Audioclip: Extending clip to image, text and audio
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 229b3f0b-af1d-45cc-a8dd-ff58ed09cc74 · outbound
Can Sound Replace Vision in LLaVA With Token Substitution? chirp" from the
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 7bf345b4-1253-4cd4-a972-2834263afd53 · outbound
Can Sound Replace Vision in LLaVA With Token Substitution? The Kinetics Human Action Video Dataset
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 33696968-3c3e-4fae-b024-5ba2912e6da6 · outbound
Can Sound Replace Vision in LLaVA With Token Substitution? Audiocaps: Generating captions for audios in the wild
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 0ece4407-90ed-45c0-bb84-e19ffd824ba0 · outbound
Can Sound Replace Vision in LLaVA With Token Substitution? Align before fuse: Vision and language representation learning with momentum distillation
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e7805c19-772b-4ef9-8c6f-e235e47ab875 · outbound
Can Sound Replace Vision in LLaVA With Token Substitution? Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 61577345-de57-47c0-bd6c-798f9fa993a9 · outbound
Can Sound Replace Vision in LLaVA With Token Substitution? Visual instruction tuning
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7faa1bbc-c9d6-4356-8203-b052561945b9 · outbound
Can Sound Replace Vision in LLaVA With Token Substitution? Oscar: Object state captioning and state change representation
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 99f20e5a-5bf3-46a7-a8e7-18406e0b812f · outbound
Can Sound Replace Vision in LLaVA With Token Substitution? Learning transferable visual models from natural language supervision
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 22f4aa2f-dee6-4f6d-8559-e3b3135d532a · outbound
Can Sound Replace Vision in LLaVA With Token Substitution? Robust speech recognition via large-scale weak supervision
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 64313291-d0c2-47a0-8951-a8fcf4bf8884 · outbound
Can Sound Replace Vision in LLaVA With Token Substitution? Tvsum: Summarizing web videos using titles
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 5714cb3f-b4da-4cda-9efe-709e06c74dfc · outbound
Can Sound Replace Vision in LLaVA With Token Substitution? From vision to audio and beyond: A unified model for audio-visual representation and generation
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation f68a795f-4049-49d5-8e57-4c0ac710b825 · outbound
Can Sound Replace Vision in LLaVA With Token Substitution? Deep learning and the information bottleneck principle
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation aceba966-5d20-446b-af05-fdfcf9689d2c · outbound
Can Sound Replace Vision in LLaVA With Token Substitution? Learning audio concepts from counterfactual natural language
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation a6d8834f-01e8-4935-9b08-ee2af8f4e48c · outbound
Can Sound Replace Vision in LLaVA With Token Substitution? Quality over quantity? LLM -based curation for a data-efficient audio-video foundation model
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 17c23739-ba88-4dc4-ad22-fbe49fe4f1c2 · outbound
Can Sound Replace Vision in LLaVA With Token Substitution? Wav2clip: Learning robust audio representations from clip
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 5faf1619-94bc-447f-a128-e519893e1c40 · outbound
Can Sound Replace Vision in LLaVA With Token Substitution? Rangevit: Towards vision transformers for 3d semantic segmentation in autonomous driving
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation c19962b8-8fbf-4fb8-8e2a-d447e947a18f · outbound
Can Sound Replace Vision in LLaVA With Token Substitution? Square Attack: A Query-efficient Black-box Adversarial Attack via Random Search
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 5e2bd0f7-e70c-40f6-8d5c-acb3f304682c · outbound
Can Sound Replace Vision in LLaVA With Token Substitution? Adversarial example games
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation ffeaad62-4b5e-4f16-bdf0-2a40badd41e2 · outbound
Can Sound Replace Vision in LLaVA With Token Substitution? Towards Evaluating the Robustness of Neural Networks
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation b163a716-518e-43e0-bf60-941595ba0217 · outbound
Can Sound Replace Vision in LLaVA With Token Substitution? Boosting Decision-based Black-box Adversarial Attacks with Random Sign Flip
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 162bc593-88b9-4252-aa74-b9e015fd55f0 · outbound
Can Sound Replace Vision in LLaVA With Token Substitution? Boosting Adversarial Attacks with Momentum
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 5de57de4-7aff-4281-b5d6-04573fea7f09 · outbound
Can Sound Replace Vision in LLaVA With Token Substitution? Evading Defenses to Transferable Adversarial Examples by Translation-invariant Attacks
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 691a9c33-bac7-42d8-bc04-796ebf1ff2ef · outbound
Can Sound Replace Vision in LLaVA With Token Substitution? An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation b7a3cfb9-8ae9-4db3-934c-c7f8b83ddd59 · outbound
Can Sound Replace Vision in LLaVA With Token Substitution? Patch-wise Attack for Fooling Deep Neural Network
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 911f3898-f876-4803-827a-5927ea5e3425 · outbound
Can Sound Replace Vision in LLaVA With Token Substitution? Explaining and Harnessing Adversarial Examples
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 99a82fea-1f55-4916-b397-54a4a5c8839b · outbound
Can Sound Replace Vision in LLaVA With Token Substitution? Lgv: Boosting Adversarial Example Transferability from Large Geometric Vicinity
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 7ce17b0a-e15f-420e-8080-63a941d23db1 · outbound
Can Sound Replace Vision in LLaVA With Token Substitution? Deep Residual Learning for Image Recognition
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 9549b075-33a5-4ad4-a4b7-6f400c0def6b · outbound
Can Sound Replace Vision in LLaVA With Token Substitution? Rethinking spatial dimensions of vision transformers
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 2b82a15c-a747-4962-b1dd-e75f0d75fa96 · outbound
Can Sound Replace Vision in LLaVA With Token Substitution? Densely Connected Convolutional Networks
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 0fab9f73-1bad-4f08-93e2-7d8b6fd9dde3 · outbound
Can Sound Replace Vision in LLaVA With Token Substitution? Adversarial Examples in the Physical World
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 0c51461e-4abf-46ed-9e5b-a428033f9080 · outbound
Can Sound Replace Vision in LLaVA With Token Substitution? Decision-based Adversarial Attack with Frequency Mixup
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 902a1a39-5021-42e5-a129-e76cf6fb6b34 · outbound
Can Sound Replace Vision in LLaVA With Token Substitution? Learning Transferable Adversarial Examples via Ghost Networks
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 1b884826-fdc7-483d-8d24-4e11505f6783 · outbound
Can Sound Replace Vision in LLaVA With Token Substitution? Nesterov Accelerated Gradient and Scale Invariance for Adversarial Attacks
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 10a42f93-030b-4607-8c3f-6f6bfde9b06d · outbound
Can Sound Replace Vision in LLaVA With Token Substitution? Delving into Transferable Adversarial Examples and Black-box Attacks
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 003a6926-40dc-4584-8f12-4310e49c0681 · outbound
Can Sound Replace Vision in LLaVA With Token Substitution? Swin Transformer: Hierarchical Vision Transformer using Shifted Windows
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation b8f5898f-aae8-47a4-a262-99784bd424ad · outbound
Can Sound Replace Vision in LLaVA With Token Substitution? Frequency Domain Model Augmentation for Adversarial Attack
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 37486144-b4a5-4d59-a0dd-0606ec1aa59b · outbound
Can Sound Replace Vision in LLaVA With Token Substitution? Hierarchical vision transformers for disease progression detection in chest x-ray images
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation db8fc341-514c-4879-b35a-eb2e998ce721 · outbound
Can Sound Replace Vision in LLaVA With Token Substitution? Deepfool: A Simple and Accurate Method to Fool Deep Neural Networks
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 8185f638-4e69-49e9-896b-730c5825501f · outbound
Can Sound Replace Vision in LLaVA With Token Substitution? Intriguing properties of neural networks
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f7aeaafc-7e5c-4128-9721-c41694a4a926 · outbound
Can Sound Replace Vision in LLaVA With Token Substitution? Boosting the Transferability of Adversarial Attacks with Global Momentum Initialization
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c3357ed4-f904-4556-aaab-43385d4c34b4 · outbound
Can Sound Replace Vision in LLaVA With Token Substitution? Enhancing the Transferability of Adversarial Attacks through Variance Tuning
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 906fdac7-92a9-4306-8443-262cdec68fa9 · outbound
Can Sound Replace Vision in LLaVA With Token Substitution? Admix: Enhancing the Transferability of Adversarial Attacks
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation fea9f415-d335-4f43-b80f-178a3a154bf3 · outbound
Can Sound Replace Vision in LLaVA With Token Substitution? Boosting Adversarial Transferability through Enhanced Momentum
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 739907ee-f7b9-4e18-9ec6-d6426b2e6076 · outbound
Can Sound Replace Vision in LLaVA With Token Substitution? Triangle Attack: A Query-efficient Decision-based Adversarial Attack
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation b92c5d52-1045-4180-87e1-d56247199dca · outbound
Can Sound Replace Vision in LLaVA With Token Substitution? Unresolved cited work
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation d11d0327-8d0e-44bb-b31d-15da0ad856ec · outbound
Can Sound Replace Vision in LLaVA With Token Substitution? Aggregated residual transformations for deep neural networks
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation d575648a-a125-4569-a48b-5f88f5be1e5e · outbound
Can Sound Replace Vision in LLaVA With Token Substitution? Stochastic Variance Reduced Ensemble Adversarial Attack for Boosting the Adversarial Transferability
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 97d36d4d-76f2-42ea-8622-fab888ba1ffd · outbound
Can Sound Replace Vision in LLaVA With Token Substitution? Meta-learning the Search Distribution of Black-box Random Search Based Adversarial Attacks
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 1351d327-b91a-4381-beff-3dc01ef964ef · outbound
Can Sound Replace Vision in LLaVA With Token Substitution? Learning to transform dynamically for better adversarial transferability
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation d561b2b5-bade-424b-9c6f-6b6253c02d99 · outbound
Can Sound Replace Vision in LLaVA With Token Substitution? Uia-vit: Unsupervised inconsistency-aware method based on vision transformer for face forgery detection
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
No inbound Pith citation observations are available.