Pith. sign in

Paper Citation Record · LEDGER

Can Sound Replace Vision in LLaVA With Token Substitution?

As of 10 August 2026, this Paper Citation Record lists 59 of 59 outbound references and 0 inbound Pith citation observations for arXiv:2506.10416.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.10416 v2

Coverage vector

measured 59 of 59 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T04:33:43.018059Z

measured 59 of 59 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

59 of 59 outbound references displayed

  • verified exact0
  • verified fuzzy52
  • unresolved7
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1eeec5c1-e129-4f26-8102-1c04208b468f · outbound

This paper cites Deep Variational Information Bottleneck.

Can Sound Replace Vision in LLaVA With Token Substitution? Deep Variational Information Bottleneck

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T04:33:34.550596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:33:34.550596Z digest=sha256:e05f099bb472417f8e03cc4728d8b2ab2f2da730eb1329e8c1760b78daea26c8

Observation 399d9b91-dab0-4331-bc52-2ecd06612ede · outbound

This paper cites Arandjelovic and P.

Can Sound Replace Vision in LLaVA With Token Substitution? Arandjelovic and P

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:33:57.983039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T04:33:34.722897Z digest=sha256:fb8e065073f6a5c4a9189db4bcda95bb1ba183153982bb97d35ed57371eef889

Observation 25b4e1c7-bac9-4129-9302-39a680047730 · outbound

This paper cites Eagle: Egocentric aggregated language-video engine.

Can Sound Replace Vision in LLaVA With Token Substitution? Eagle: Egocentric aggregated language-video engine

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:33:57.700719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T04:33:34.900943Z digest=sha256:fcfcde3e36d1b78ff67b4e09339ddcc74e6b0ea377c66452e8ed52632924d21c

Observation f7582180-498f-4591-a8a5-08b2f53abd49 · outbound

This paper cites Vggsound: A large-scale audio-visual dataset.

Can Sound Replace Vision in LLaVA With Token Substitution? Vggsound: A large-scale audio-visual dataset

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:33:57.349878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T04:33:35.074043Z digest=sha256:77aba540897918a96bce562e262d88136c3e31c645b5e95a1c09daf472c2d0eb

Observation 1ec82599-025e-4684-806a-cccee916717b · outbound

This paper cites Clap learning audio concepts from natural language supervision.

Can Sound Replace Vision in LLaVA With Token Substitution? Clap learning audio concepts from natural language supervision

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:33:57.060553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T04:33:35.222849Z digest=sha256:9369a088ca1d88ba3ecf64e1b983044dd43ecec8f4df0001f5996034c0a9da95

Observation 0b3aa060-f582-4a32-afff-b3bce5803e62 · outbound

This paper cites Imagebind: One embedding space to bind them all.

Can Sound Replace Vision in LLaVA With Token Substitution? Imagebind: One embedding space to bind them all

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:33:56.756949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T04:33:35.398977Z digest=sha256:0f0962d06f451d50aa3b004e5645ab0684a121ecde5abc72671704181e75e144

Observation f8ce3aa5-20d5-45e3-874e-87172bf9c4c6 · outbound

This paper cites Audioset, 2017.

Can Sound Replace Vision in LLaVA With Token Substitution? Audioset, 2017

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:33:56.482953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T04:33:35.537791Z digest=sha256:c55ba797531614095ab90f64bccb24064f0cd60ccf4db68be97ad32d4e5dc806

Observation 2fb0fe18-772e-4e59-9f9d-e2cb2b8b7012 · outbound

This paper cites Ego4d: Around the world in 3,000 hours of egocentric video.

Can Sound Replace Vision in LLaVA With Token Substitution? Ego4d: Around the world in 3,000 hours of egocentric video

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:33:56.215372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T04:33:35.719073Z digest=sha256:85332c78052a2e3f9015dd2f004df283a18bb7c536051108740ee2267c6bcedd

Observation 688d5303-08a6-4b08-8916-95101b4fe48c · outbound

This paper cites Audioclip: Extending clip to image, text and audio.

Can Sound Replace Vision in LLaVA With Token Substitution? Audioclip: Extending clip to image, text and audio

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:33:55.881108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T04:33:35.852888Z digest=sha256:30b55465cf915f2948077e084f2e04dd60fc184fc0e67e180f68a24066b17a25

Observation 229b3f0b-af1d-45cc-a8dd-ff58ed09cc74 · outbound

This paper cites chirp" from the.

Can Sound Replace Vision in LLaVA With Token Substitution? chirp" from the

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:33:55.559027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T04:33:35.996086Z digest=sha256:cea429d62f28a3a13e3f019a1f87831c3eb86ceba28f5a13abf96e525f089b4e

Observation 7bf345b4-1253-4cd4-a972-2834263afd53 · outbound

This paper cites The Kinetics Human Action Video Dataset.

Can Sound Replace Vision in LLaVA With Token Substitution? The Kinetics Human Action Video Dataset

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T04:33:36.134515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:33:36.134515Z digest=sha256:adcb4821246e33eaf2cdc0b56dd82ee1dc002612158887aff0f601025cbd247b

Observation 33696968-3c3e-4fae-b024-5ba2912e6da6 · outbound

This paper cites Audiocaps: Generating captions for audios in the wild.

Can Sound Replace Vision in LLaVA With Token Substitution? Audiocaps: Generating captions for audios in the wild

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:33:55.221888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T04:33:36.315150Z digest=sha256:ba4b74bb067611072fe05fb3b14b5d154e7db47625eedd8f974668e79e98cd0d

Observation 0ece4407-90ed-45c0-bb84-e19ffd824ba0 · outbound

This paper cites Align before fuse: Vision and language representation learning with momentum distillation.

Can Sound Replace Vision in LLaVA With Token Substitution? Align before fuse: Vision and language representation learning with momentum distillation

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T04:33:36.437297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:33:36.437297Z digest=sha256:718d6221ae5adf992eb4195fa7021f66f29a07c8e548373231987e192dd05f19

Observation e7805c19-772b-4ef9-8c6f-e235e47ab875 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

Can Sound Replace Vision in LLaVA With Token Substitution? Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:33:54.948158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T04:33:36.573178Z digest=sha256:fa4134b77fe281150897a0fc5c74377a162d5d84cbaca80d02c9051ab946357d

Observation 61577345-de57-47c0-bd6c-798f9fa993a9 · outbound

This paper cites Visual instruction tuning.

Can Sound Replace Vision in LLaVA With Token Substitution? Visual instruction tuning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T04:33:36.731603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:33:36.731603Z digest=sha256:f27714a976dc68ee217f17af71e628e78901ac99ecdd87f9b662e81d9093d27b

Observation 7faa1bbc-c9d6-4356-8203-b052561945b9 · outbound

This paper cites Oscar: Object state captioning and state change representation.

Can Sound Replace Vision in LLaVA With Token Substitution? Oscar: Object state captioning and state change representation

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:33:54.581357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T04:33:36.882611Z digest=sha256:f3180835b3e572fbaa1ffe53ed69577ec6f717f1981bbaf17beb088b95963e23

Observation 99f20e5a-5bf3-46a7-a8e7-18406e0b812f · outbound

This paper cites Learning transferable visual models from natural language supervision.

Can Sound Replace Vision in LLaVA With Token Substitution? Learning transferable visual models from natural language supervision

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:33:54.274477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T04:33:36.985961Z digest=sha256:7d9c70a51ed0d8b21e51f44be20bb249513a20fcd90a7c71f565f1e696becaa4

Observation 22f4aa2f-dee6-4f6d-8559-e3b3135d532a · outbound

This paper cites Robust speech recognition via large-scale weak supervision.

Can Sound Replace Vision in LLaVA With Token Substitution? Robust speech recognition via large-scale weak supervision

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:33:53.979587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T04:33:37.111056Z digest=sha256:f66a35ec10e514234aff7ce94619c6bb41d2fd5be65249b0c351fe4fa1d99792

Observation 64313291-d0c2-47a0-8951-a8fcf4bf8884 · outbound

This paper cites Tvsum: Summarizing web videos using titles.

Can Sound Replace Vision in LLaVA With Token Substitution? Tvsum: Summarizing web videos using titles

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:33:53.640365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T04:33:37.269770Z digest=sha256:176ccf3c5091e33b7cb1f8e53ac37d51524a126a5a12a8e4735d3001ff243307

Observation 5714cb3f-b4da-4cda-9efe-709e06c74dfc · outbound

This paper cites From vision to audio and beyond: A unified model for audio-visual representation and generation.

Can Sound Replace Vision in LLaVA With Token Substitution? From vision to audio and beyond: A unified model for audio-visual representation and generation

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:33:53.338208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T04:33:37.392366Z digest=sha256:c0034de1e05d0f20161deb35f0e1e9026eefcb6be0aa62b0f7a96f4060d38acc

Observation f68a795f-4049-49d5-8e57-4c0ac710b825 · outbound

This paper cites Deep learning and the information bottleneck principle.

Can Sound Replace Vision in LLaVA With Token Substitution? Deep learning and the information bottleneck principle

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:33:53.083973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T04:33:37.518237Z digest=sha256:665a8d34beed843d64cc17c7f7a095f3ddf7ccae7df7b3d1f79071215b614167

Observation aceba966-5d20-446b-af05-fdfcf9689d2c · outbound

This paper cites Learning audio concepts from counterfactual natural language.

Can Sound Replace Vision in LLaVA With Token Substitution? Learning audio concepts from counterfactual natural language

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:33:52.816523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T04:33:37.637945Z digest=sha256:407c773675a79f59b86aec6730324f3f80ef554b3ed6a76f8a292bba668a24de

Observation a6d8834f-01e8-4935-9b08-ee2af8f4e48c · outbound

This paper cites Quality over quantity? LLM -based curation for a data-efficient audio-video foundation model.

Can Sound Replace Vision in LLaVA With Token Substitution? Quality over quantity? LLM -based curation for a data-efficient audio-video foundation model

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:33:52.532284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T04:33:37.808323Z digest=sha256:50889e1d56e76e8f9d77e976f78885c9a54476d7f4166ffec49e6d7725814db4

Observation 17c23739-ba88-4dc4-ad22-fbe49fe4f1c2 · outbound

This paper cites Wav2clip: Learning robust audio representations from clip.

Can Sound Replace Vision in LLaVA With Token Substitution? Wav2clip: Learning robust audio representations from clip

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:33:52.272157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T04:33:37.984137Z digest=sha256:84bcbcb8de47a4453b66e45a13a366424124f66f1e902698cd238edcdb5abc50

Observation 5faf1619-94bc-447f-a128-e519893e1c40 · outbound

This paper cites Rangevit: Towards vision transformers for 3d semantic segmentation in autonomous driving.

Can Sound Replace Vision in LLaVA With Token Substitution? Rangevit: Towards vision transformers for 3d semantic segmentation in autonomous driving

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:33:51.978934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T04:33:38.150187Z digest=sha256:687009751e9653ec1bcfe5052c2d7015abb3eb8dfb4b2faddcf14b5ea6ba6267

Observation c19962b8-8fbf-4fb8-8e2a-d447e947a18f · outbound

This paper cites Square Attack: A Query-efficient Black-box Adversarial Attack via Random Search.

Can Sound Replace Vision in LLaVA With Token Substitution? Square Attack: A Query-efficient Black-box Adversarial Attack via Random Search

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:33:51.734693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T04:33:38.314332Z digest=sha256:36152c2eb066ea91bf6a3131b90d0ea797e6aa43ac1ca9c77c9c12b0b0cad4a5

Observation 5e2bd0f7-e70c-40f6-8d5c-acb3f304682c · outbound

This paper cites Adversarial example games.

Can Sound Replace Vision in LLaVA With Token Substitution? Adversarial example games

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:33:51.410476Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T04:33:38.456762Z digest=sha256:6b67c70f40439ab0b6fbaa1ec2e2462ccdd5d02c6f1ece9077e4af69997ed4ad

Observation ffeaad62-4b5e-4f16-bdf0-2a40badd41e2 · outbound

This paper cites Towards Evaluating the Robustness of Neural Networks.

Can Sound Replace Vision in LLaVA With Token Substitution? Towards Evaluating the Robustness of Neural Networks

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:33:51.165188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T04:33:38.600312Z digest=sha256:d21787d7dd560137b9cc448c2377457a56ef44b4414973eb688a7cc0bf87f28a

Observation b163a716-518e-43e0-bf60-941595ba0217 · outbound

This paper cites Boosting Decision-based Black-box Adversarial Attacks with Random Sign Flip.

Can Sound Replace Vision in LLaVA With Token Substitution? Boosting Decision-based Black-box Adversarial Attacks with Random Sign Flip

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:33:50.880922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T04:33:38.762489Z digest=sha256:1f6f12987905d63c338ce3c2cea3452195278c9ad0d93ad83a9bd99856ce947d

Observation 162bc593-88b9-4252-aa74-b9e015fd55f0 · outbound

This paper cites Boosting Adversarial Attacks with Momentum.

Can Sound Replace Vision in LLaVA With Token Substitution? Boosting Adversarial Attacks with Momentum

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:33:50.570310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T04:33:38.896691Z digest=sha256:882252918d9b00ab22b2b01f9afc6bb3a57d6c77cc0fe38d681af45ce21423be

Observation 5de57de4-7aff-4281-b5d6-04573fea7f09 · outbound

This paper cites Evading Defenses to Transferable Adversarial Examples by Translation-invariant Attacks.

Can Sound Replace Vision in LLaVA With Token Substitution? Evading Defenses to Transferable Adversarial Examples by Translation-invariant Attacks

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:33:50.247426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T04:33:39.042728Z digest=sha256:3dd5289f0b30d9b5fa243dbbf3246203c0ed20ec34cbfe76d41e3076dfdeb45e

Observation 691a9c33-bac7-42d8-bc04-796ebf1ff2ef · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

Can Sound Replace Vision in LLaVA With Token Substitution? An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:33:49.965747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T04:33:39.154387Z digest=sha256:dd096960c11cc1804288d949af470f91c155b8c34a621de7f04417df436205df

Observation b7a3cfb9-8ae9-4db3-934c-c7f8b83ddd59 · outbound

This paper cites Patch-wise Attack for Fooling Deep Neural Network.

Can Sound Replace Vision in LLaVA With Token Substitution? Patch-wise Attack for Fooling Deep Neural Network

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:33:49.638510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T04:33:39.269174Z digest=sha256:041567e97cbb3ab360af6dd4835d7ffd5a430a880a321437fcfcad66b5a84e8f

Observation 911f3898-f876-4803-827a-5927ea5e3425 · outbound

This paper cites Explaining and Harnessing Adversarial Examples.

Can Sound Replace Vision in LLaVA With Token Substitution? Explaining and Harnessing Adversarial Examples

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:33:49.380073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T04:33:39.420094Z digest=sha256:900f60edfa145137317979e5719083605ce170506f1ca937ed86c93b46b649b2

Observation 99a82fea-1f55-4916-b397-54a4a5c8839b · outbound

This paper cites Lgv: Boosting Adversarial Example Transferability from Large Geometric Vicinity.

Can Sound Replace Vision in LLaVA With Token Substitution? Lgv: Boosting Adversarial Example Transferability from Large Geometric Vicinity

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:33:49.164734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T04:33:39.571232Z digest=sha256:8b071fcd6ffd1e794a1bb6d22c7438e32f53533fc27a97f586460044d564a6a1

Observation 7ce17b0a-e15f-420e-8080-63a941d23db1 · outbound

This paper cites Deep Residual Learning for Image Recognition.

Can Sound Replace Vision in LLaVA With Token Substitution? Deep Residual Learning for Image Recognition

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:33:48.894312Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T04:33:39.684366Z digest=sha256:8fec31aa1d3ae5a65542bd26e57f47f9126e2d97d7c7e6a687d8687ca93bff95

Observation 9549b075-33a5-4ad4-a4b7-6f400c0def6b · outbound

This paper cites Rethinking spatial dimensions of vision transformers.

Can Sound Replace Vision in LLaVA With Token Substitution? Rethinking spatial dimensions of vision transformers

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:33:48.571177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T04:33:39.799507Z digest=sha256:27d1d2d2e543d13ffc9d99ec921468038b5680842902518b22135692a254e169

Observation 2b82a15c-a747-4962-b1dd-e75f0d75fa96 · outbound

This paper cites Densely Connected Convolutional Networks.

Can Sound Replace Vision in LLaVA With Token Substitution? Densely Connected Convolutional Networks

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:33:48.326317Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T04:33:39.931642Z digest=sha256:aac2b1de97c0339a91737552ba75cbc090b89168250bf074cb0d9f354f7a62a9

Observation 0fab9f73-1bad-4f08-93e2-7d8b6fd9dde3 · outbound

This paper cites Adversarial Examples in the Physical World.

Can Sound Replace Vision in LLaVA With Token Substitution? Adversarial Examples in the Physical World

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:33:48.056715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T04:33:40.081411Z digest=sha256:9d41c2b7ac936bc8e53dfe5ac132b51f117a852f70a36246e7f241bc14a6c258

Observation 0c51461e-4abf-46ed-9e5b-a428033f9080 · outbound

This paper cites Decision-based Adversarial Attack with Frequency Mixup.

Can Sound Replace Vision in LLaVA With Token Substitution? Decision-based Adversarial Attack with Frequency Mixup

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:33:47.758584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T04:33:40.227759Z digest=sha256:4eb6e3f9543fcb6501a9738b5dc48ba2dc8784761629be322d5b2271ec66f3f2

Observation 902a1a39-5021-42e5-a129-e76cf6fb6b34 · outbound

This paper cites Learning Transferable Adversarial Examples via Ghost Networks.

Can Sound Replace Vision in LLaVA With Token Substitution? Learning Transferable Adversarial Examples via Ghost Networks

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:33:47.457825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T04:33:40.380157Z digest=sha256:9e011a52237ef897a8523a0b19bbb802b146b0366a90846b62af8e5ac99d9b9f

Observation 1b884826-fdc7-483d-8d24-4e11505f6783 · outbound

This paper cites Nesterov Accelerated Gradient and Scale Invariance for Adversarial Attacks.

Can Sound Replace Vision in LLaVA With Token Substitution? Nesterov Accelerated Gradient and Scale Invariance for Adversarial Attacks

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:33:47.203450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T04:33:40.507235Z digest=sha256:c89cfec41bef8b38fce9b9a73a623b8eb70fec50bfb07665d4856a02950e769f

Observation 10a42f93-030b-4607-8c3f-6f6bfde9b06d · outbound

This paper cites Delving into Transferable Adversarial Examples and Black-box Attacks.

Can Sound Replace Vision in LLaVA With Token Substitution? Delving into Transferable Adversarial Examples and Black-box Attacks

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:33:46.947867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T04:33:40.663487Z digest=sha256:1f2d817eecf34c90dcd7b58e05e9a3af2da398009907ebd349ab21143d38200c

Observation 003a6926-40dc-4584-8f12-4310e49c0681 · outbound

This paper cites Swin Transformer: Hierarchical Vision Transformer using Shifted Windows.

Can Sound Replace Vision in LLaVA With Token Substitution? Swin Transformer: Hierarchical Vision Transformer using Shifted Windows

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:33:46.690725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T04:33:40.778669Z digest=sha256:a64f684432f06ba2493b80ce1dc19b8860c84b4f143a8da125a538068fdfbf48

Observation b8f5898f-aae8-47a4-a262-99784bd424ad · outbound

This paper cites Frequency Domain Model Augmentation for Adversarial Attack.

Can Sound Replace Vision in LLaVA With Token Substitution? Frequency Domain Model Augmentation for Adversarial Attack

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:33:46.455924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T04:33:40.899836Z digest=sha256:743a214f6464c6860356f5da94619846ee1d4affb70d6699c55a6281dc8dc066

Observation 37486144-b4a5-4d59-a0dd-0606ec1aa59b · outbound

This paper cites Hierarchical vision transformers for disease progression detection in chest x-ray images.

Can Sound Replace Vision in LLaVA With Token Substitution? Hierarchical vision transformers for disease progression detection in chest x-ray images

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:33:46.196221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T04:33:41.060840Z digest=sha256:6e0b3f90c1999d8beb06d0b878ba3832a6d81ed72d73a6970ff6a3f71abae3e7

Observation db8fc341-514c-4879-b35a-eb2e998ce721 · outbound

This paper cites Deepfool: A Simple and Accurate Method to Fool Deep Neural Networks.

Can Sound Replace Vision in LLaVA With Token Substitution? Deepfool: A Simple and Accurate Method to Fool Deep Neural Networks

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:33:45.905787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T04:33:41.220911Z digest=sha256:9a88da4a9a5e523b1bbc57436f83e4e2fb62b8e1aedb9a5ad074921f470bc30d

Observation 8185f638-4e69-49e9-896b-730c5825501f · outbound

This paper cites Intriguing properties of neural networks.

Can Sound Replace Vision in LLaVA With Token Substitution? Intriguing properties of neural networks

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T04:33:41.406253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:33:41.406253Z digest=sha256:9fd84a87ee85be3f67d8d5549580b9e5eb3481f5a443c181fb93c466defb344a

Observation f7aeaafc-7e5c-4128-9721-c41694a4a926 · outbound

This paper cites Boosting the Transferability of Adversarial Attacks with Global Momentum Initialization.

Can Sound Replace Vision in LLaVA With Token Substitution? Boosting the Transferability of Adversarial Attacks with Global Momentum Initialization

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T04:33:41.562333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:33:41.562333Z digest=sha256:4bade788f81fcee19783cd9bdb808bf8243d2bc4fe4ffdc999f0284bb018c62a

Observation c3357ed4-f904-4556-aaab-43385d4c34b4 · outbound

This paper cites Enhancing the Transferability of Adversarial Attacks through Variance Tuning.

Can Sound Replace Vision in LLaVA With Token Substitution? Enhancing the Transferability of Adversarial Attacks through Variance Tuning

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:33:45.653630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T04:33:41.695567Z digest=sha256:564d711a52115d1947207b8f2e503f51f1e9dddffa519d4a7a2b53dbd432f602

Observation 906fdac7-92a9-4306-8443-262cdec68fa9 · outbound

This paper cites Admix: Enhancing the Transferability of Adversarial Attacks.

Can Sound Replace Vision in LLaVA With Token Substitution? Admix: Enhancing the Transferability of Adversarial Attacks

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:33:45.420097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T04:33:41.865033Z digest=sha256:a91ab039ff13c84436a1bf525dc1f1c2152c6655badf6ddd49364336831ce4b8

Observation fea9f415-d335-4f43-b80f-178a3a154bf3 · outbound

This paper cites Boosting Adversarial Transferability through Enhanced Momentum.

Can Sound Replace Vision in LLaVA With Token Substitution? Boosting Adversarial Transferability through Enhanced Momentum

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:33:45.144411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T04:33:41.978382Z digest=sha256:c11ce5b35ae876625edc5801a05a3c9ea67e52fbc7518a49bebcead29b9444c2

Observation 739907ee-f7b9-4e18-9ec6-d6426b2e6076 · outbound

This paper cites Triangle Attack: A Query-efficient Decision-based Adversarial Attack.

Can Sound Replace Vision in LLaVA With Token Substitution? Triangle Attack: A Query-efficient Decision-based Adversarial Attack

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:33:44.796408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T04:33:42.096650Z digest=sha256:8c0a146869826bb2c075a7292e9a5a4f583c0ca1d53e2d10f278c0ecfa409d5c

Observation b92c5d52-1045-4180-87e1-d56247199dca · outbound

This paper cites an unresolved cited work.

Can Sound Replace Vision in LLaVA With Token Substitution? Unresolved cited work

Reference 54

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:33:44.541442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T04:33:42.263455Z digest=sha256:38536c82d55508d6d56dbc0235104bfa81f13c0de74c9d49f2700834ad0a3080

Observation d11d0327-8d0e-44bb-b31d-15da0ad856ec · outbound

This paper cites Aggregated residual transformations for deep neural networks.

Can Sound Replace Vision in LLaVA With Token Substitution? Aggregated residual transformations for deep neural networks

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:33:44.306326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T04:33:42.432082Z digest=sha256:5aff9d76af64a1218ab0996cddcd46ab8f4e5b3625fccfb3e731860c114e41e6

Observation d575648a-a125-4569-a48b-5f88f5be1e5e · outbound

This paper cites Stochastic Variance Reduced Ensemble Adversarial Attack for Boosting the Adversarial Transferability.

Can Sound Replace Vision in LLaVA With Token Substitution? Stochastic Variance Reduced Ensemble Adversarial Attack for Boosting the Adversarial Transferability

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:33:44.029146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T04:33:42.562240Z digest=sha256:38647c14ce747b8cfcef952cb92150245ca1dd5ffaa8c8e7a6abc620c140f850

Observation 97d36d4d-76f2-42ea-8622-fab888ba1ffd · outbound

This paper cites Meta-learning the Search Distribution of Black-box Random Search Based Adversarial Attacks.

Can Sound Replace Vision in LLaVA With Token Substitution? Meta-learning the Search Distribution of Black-box Random Search Based Adversarial Attacks

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:33:43.749934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T04:33:42.713169Z digest=sha256:2f9db74499e32bbe0c756fb9dde8c72bf82383f60f09973f671efa469c35dbb5

Observation 1351d327-b91a-4381-beff-3dc01ef964ef · outbound

This paper cites Learning to transform dynamically for better adversarial transferability.

Can Sound Replace Vision in LLaVA With Token Substitution? Learning to transform dynamically for better adversarial transferability

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:33:43.472480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T04:33:42.894400Z digest=sha256:014613235d059586a95fd301b3adebd4dc9006c0a468f76f34e1cc0dd19d494c

Observation d561b2b5-bade-424b-9c6f-6b6253c02d99 · outbound

This paper cites Uia-vit: Unsupervised inconsistency-aware method based on vision transformer for face forgery detection.

Can Sound Replace Vision in LLaVA With Token Substitution? Uia-vit: Unsupervised inconsistency-aware method based on vision transformer for face forgery detection

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:33:43.260754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T04:33:43.018059Z digest=sha256:437786ee1bb2e79c8c5d322fed53c39de4faec5969853ab6c36cc9ae0adac111

Pith citing papers

No inbound Pith citation observations are available.