Pith. sign in

Paper Citation Record · LEDGER

Can Sound Replace Vision in LLaVA With Token Substitution?

As of 18 August 2026, this Paper Citation Record lists 59 of 59 outbound references and 0 inbound Pith citation observations for arXiv:2506.10416.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.10416 v2

Coverage vector

measured 59 of 59 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T04:33:43.018059Z

measured 59 of 59 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

59 of 59 outbound references displayed

  • verified exact0
  • verified fuzzy52
  • unresolved7
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1eeec5c1-e129-4f26-8102-1c04208b468f · outbound

This paper cites Deep Variational Information Bottleneck.

Can Sound Replace Vision in LLaVA With Token Substitution? Deep Variational Information Bottleneck

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T04:33:34.550596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:33:34.550596Z digest=sha256:2e8815b2457e6fd9ffb8a211a14429a84a46e0ad587e5809660a96ca44e1a8ef

Observation 399d9b91-dab0-4331-bc52-2ecd06612ede · outbound

This paper cites Arandjelovic and P.

Can Sound Replace Vision in LLaVA With Token Substitution? Arandjelovic and P

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:33:57.983039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T04:33:34.722897Z digest=sha256:2ae0052988d49212e1d199bb85d64b561155523c7aeee039c79d3783e0961214

Observation 25b4e1c7-bac9-4129-9302-39a680047730 · outbound

This paper cites Eagle: Egocentric aggregated language-video engine.

Can Sound Replace Vision in LLaVA With Token Substitution? Eagle: Egocentric aggregated language-video engine

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:33:57.700719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T04:33:34.900943Z digest=sha256:f737f33d6d2f7ee30fb38c65ffb5a36be127655e88eac509ae4223f13317823c

Observation f7582180-498f-4591-a8a5-08b2f53abd49 · outbound

This paper cites Vggsound: A large-scale audio-visual dataset.

Can Sound Replace Vision in LLaVA With Token Substitution? Vggsound: A large-scale audio-visual dataset

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:33:57.349878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T04:33:35.074043Z digest=sha256:3885b8ffb6c1bc3e8c3fd35eed7db85892c6dcc07e1c875e07ce56e8de584581

Observation 1ec82599-025e-4684-806a-cccee916717b · outbound

This paper cites Clap learning audio concepts from natural language supervision.

Can Sound Replace Vision in LLaVA With Token Substitution? Clap learning audio concepts from natural language supervision

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:33:57.060553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T04:33:35.222849Z digest=sha256:d0abc1dd12914cca017d0b065179024d2d00ac6ff265e199478ec177694ce8b0

Observation 0b3aa060-f582-4a32-afff-b3bce5803e62 · outbound

This paper cites Imagebind: One embedding space to bind them all.

Can Sound Replace Vision in LLaVA With Token Substitution? Imagebind: One embedding space to bind them all

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:33:56.756949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T04:33:35.398977Z digest=sha256:6f1b650d67b4de3c8ea9e65a6b755a8d534694a5e763d10abb9edf337089de01

Observation f8ce3aa5-20d5-45e3-874e-87172bf9c4c6 · outbound

This paper cites Audioset, 2017.

Can Sound Replace Vision in LLaVA With Token Substitution? Audioset, 2017

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:33:56.482953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T04:33:35.537791Z digest=sha256:ee39d46dd87c3a213b7a0ebd4a4394f341c3caa853634fab2378b48a451b6258

Observation 2fb0fe18-772e-4e59-9f9d-e2cb2b8b7012 · outbound

This paper cites Ego4d: Around the world in 3,000 hours of egocentric video.

Can Sound Replace Vision in LLaVA With Token Substitution? Ego4d: Around the world in 3,000 hours of egocentric video

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:33:56.215372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T04:33:35.719073Z digest=sha256:d8f90305571a68ec34a5433a536cf0afaee12ec4140f87f690bba38603007857

Observation 688d5303-08a6-4b08-8916-95101b4fe48c · outbound

This paper cites Audioclip: Extending clip to image, text and audio.

Can Sound Replace Vision in LLaVA With Token Substitution? Audioclip: Extending clip to image, text and audio

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:33:55.881108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T04:33:35.852888Z digest=sha256:f0b09a26f1362f6242359639457b60eb1c6edc95400bf4b796ce57746d310c80

Observation 229b3f0b-af1d-45cc-a8dd-ff58ed09cc74 · outbound

This paper cites chirp" from the.

Can Sound Replace Vision in LLaVA With Token Substitution? chirp" from the

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:33:55.559027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T04:33:35.996086Z digest=sha256:e9238eeec9dceb11d6cfe308492a5d1ca1709641c8bc8528309430acc08afc17

Observation 7bf345b4-1253-4cd4-a972-2834263afd53 · outbound

This paper cites The Kinetics Human Action Video Dataset.

Can Sound Replace Vision in LLaVA With Token Substitution? The Kinetics Human Action Video Dataset

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T04:33:36.134515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:33:36.134515Z digest=sha256:c31e71ee365114e6552f4af230a82edc54487f55fea6945b9ef59ce32cc53a73

Observation 33696968-3c3e-4fae-b024-5ba2912e6da6 · outbound

This paper cites Audiocaps: Generating captions for audios in the wild.

Can Sound Replace Vision in LLaVA With Token Substitution? Audiocaps: Generating captions for audios in the wild

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:33:55.221888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T04:33:36.315150Z digest=sha256:8c4d6438ab92ca662ee27872c059d64134489a8af7975ef307dcd84eade92583

Observation 0ece4407-90ed-45c0-bb84-e19ffd824ba0 · outbound

This paper cites Align before fuse: Vision and language representation learning with momentum distillation.

Can Sound Replace Vision in LLaVA With Token Substitution? Align before fuse: Vision and language representation learning with momentum distillation

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T04:33:36.437297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:33:36.437297Z digest=sha256:b4d05a3d9801d1d58a4996c772fce576d08195b8d3a49acc572cb27c17936f05

Observation e7805c19-772b-4ef9-8c6f-e235e47ab875 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

Can Sound Replace Vision in LLaVA With Token Substitution? Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:33:54.948158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T04:33:36.573178Z digest=sha256:247ba4f0635bd7f2e161c33d4be8fff683f68ad57a8d9ef88794410730dd5ae5

Observation 61577345-de57-47c0-bd6c-798f9fa993a9 · outbound

This paper cites Visual instruction tuning.

Can Sound Replace Vision in LLaVA With Token Substitution? Visual instruction tuning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T04:33:36.731603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:33:36.731603Z digest=sha256:8ca6e415544c5cc96727925eb06ad77ea9d8510d9962880518152618c72eda02

Observation 7faa1bbc-c9d6-4356-8203-b052561945b9 · outbound

This paper cites Oscar: Object state captioning and state change representation.

Can Sound Replace Vision in LLaVA With Token Substitution? Oscar: Object state captioning and state change representation

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:33:54.581357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T04:33:36.882611Z digest=sha256:46da536df36c2cd2db6a9090a9d5ae420e8ff1880dab35fb0ded097da5e6f4d8

Observation 99f20e5a-5bf3-46a7-a8e7-18406e0b812f · outbound

This paper cites Learning transferable visual models from natural language supervision.

Can Sound Replace Vision in LLaVA With Token Substitution? Learning transferable visual models from natural language supervision

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:33:54.274477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T04:33:36.985961Z digest=sha256:cb228f5efb023973ac2c7bb451b8e2cf18f81a0a4773bc7f62d514537e275132

Observation 22f4aa2f-dee6-4f6d-8559-e3b3135d532a · outbound

This paper cites Robust speech recognition via large-scale weak supervision.

Can Sound Replace Vision in LLaVA With Token Substitution? Robust speech recognition via large-scale weak supervision

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:33:53.979587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T04:33:37.111056Z digest=sha256:fc8bb3c74364a0d3c54afd1a2758dd07927c91be6ca944539f7bee868400aeca

Observation 64313291-d0c2-47a0-8951-a8fcf4bf8884 · outbound

This paper cites Tvsum: Summarizing web videos using titles.

Can Sound Replace Vision in LLaVA With Token Substitution? Tvsum: Summarizing web videos using titles

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:33:53.640365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T04:33:37.269770Z digest=sha256:b3f12e505e8bff4f9ee1b4fa3e3a212c455bddb47f275e653b0e117b5ceee373

Observation 5714cb3f-b4da-4cda-9efe-709e06c74dfc · outbound

This paper cites From vision to audio and beyond: A unified model for audio-visual representation and generation.

Can Sound Replace Vision in LLaVA With Token Substitution? From vision to audio and beyond: A unified model for audio-visual representation and generation

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:33:53.338208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T04:33:37.392366Z digest=sha256:50052e8c9225e228070e6bc968bb5ae72e7c2539ccf9607f197f8036f6dc5eb6

Observation f68a795f-4049-49d5-8e57-4c0ac710b825 · outbound

This paper cites Deep learning and the information bottleneck principle.

Can Sound Replace Vision in LLaVA With Token Substitution? Deep learning and the information bottleneck principle

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:33:53.083973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T04:33:37.518237Z digest=sha256:1eb62d949d8b186b3f9112b87989f934166948c9fdde4fbbe316ddc99f2871b2

Observation aceba966-5d20-446b-af05-fdfcf9689d2c · outbound

This paper cites Learning audio concepts from counterfactual natural language.

Can Sound Replace Vision in LLaVA With Token Substitution? Learning audio concepts from counterfactual natural language

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:33:52.816523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T04:33:37.637945Z digest=sha256:b79f4ea4987647c34d117c3c646dbfbe495a108ea984acccf734c9cadccb4c10

Observation a6d8834f-01e8-4935-9b08-ee2af8f4e48c · outbound

This paper cites Quality over quantity? LLM -based curation for a data-efficient audio-video foundation model.

Can Sound Replace Vision in LLaVA With Token Substitution? Quality over quantity? LLM -based curation for a data-efficient audio-video foundation model

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:33:52.532284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T04:33:37.808323Z digest=sha256:5a1f2a018d0371672868e7fce24f36c6fe143921637f31d558d56651103bfa2d

Observation 17c23739-ba88-4dc4-ad22-fbe49fe4f1c2 · outbound

This paper cites Wav2clip: Learning robust audio representations from clip.

Can Sound Replace Vision in LLaVA With Token Substitution? Wav2clip: Learning robust audio representations from clip

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:33:52.272157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T04:33:37.984137Z digest=sha256:f19349a14a853791f0b437c4b21804929f4c814c4bf19ab983ebb9736f0c0147

Observation 5faf1619-94bc-447f-a128-e519893e1c40 · outbound

This paper cites Rangevit: Towards vision transformers for 3d semantic segmentation in autonomous driving.

Can Sound Replace Vision in LLaVA With Token Substitution? Rangevit: Towards vision transformers for 3d semantic segmentation in autonomous driving

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:33:51.978934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T04:33:38.150187Z digest=sha256:c116891fc49e56522bc450959ede7d8b1ca596b7df6d3593d7b1380454a1d9e7

Observation c19962b8-8fbf-4fb8-8e2a-d447e947a18f · outbound

This paper cites Square Attack: A Query-efficient Black-box Adversarial Attack via Random Search.

Can Sound Replace Vision in LLaVA With Token Substitution? Square Attack: A Query-efficient Black-box Adversarial Attack via Random Search

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:33:51.734693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T04:33:38.314332Z digest=sha256:befc18b470c17f5e9f5dfc031b15ba6305d400d0c0c714ab9de18c05d999d950

Observation 5e2bd0f7-e70c-40f6-8d5c-acb3f304682c · outbound

This paper cites Adversarial example games.

Can Sound Replace Vision in LLaVA With Token Substitution? Adversarial example games

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:33:51.410476Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T04:33:38.456762Z digest=sha256:0f9dfbd3d58d15da47088e29003b069bfcff8f79ab005c384832f5ed2ce6a7e0

Observation ffeaad62-4b5e-4f16-bdf0-2a40badd41e2 · outbound

This paper cites Towards Evaluating the Robustness of Neural Networks.

Can Sound Replace Vision in LLaVA With Token Substitution? Towards Evaluating the Robustness of Neural Networks

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:33:51.165188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T04:33:38.600312Z digest=sha256:32d542c799ef78f2d7cd4349bb7f87397422e0da908b6ac320379626b4ae47b9

Observation b163a716-518e-43e0-bf60-941595ba0217 · outbound

This paper cites Boosting Decision-based Black-box Adversarial Attacks with Random Sign Flip.

Can Sound Replace Vision in LLaVA With Token Substitution? Boosting Decision-based Black-box Adversarial Attacks with Random Sign Flip

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:33:50.880922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T04:33:38.762489Z digest=sha256:1b124c9d33763e5c12b7ced073c315c11d5df9284bf7267d1c502a9b6d56e413

Observation 162bc593-88b9-4252-aa74-b9e015fd55f0 · outbound

This paper cites Boosting Adversarial Attacks with Momentum.

Can Sound Replace Vision in LLaVA With Token Substitution? Boosting Adversarial Attacks with Momentum

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:33:50.570310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T04:33:38.896691Z digest=sha256:b9451056dc8192ba9f6f3e01bbc9f2ee855d4538ec72a1be2ba06efbbb3c787b

Observation 5de57de4-7aff-4281-b5d6-04573fea7f09 · outbound

This paper cites Evading Defenses to Transferable Adversarial Examples by Translation-invariant Attacks.

Can Sound Replace Vision in LLaVA With Token Substitution? Evading Defenses to Transferable Adversarial Examples by Translation-invariant Attacks

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:33:50.247426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T04:33:39.042728Z digest=sha256:881248692f8f0ad37f7ad1c0adf8d354274403016b19549a51d05780240ab03a

Observation 691a9c33-bac7-42d8-bc04-796ebf1ff2ef · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

Can Sound Replace Vision in LLaVA With Token Substitution? An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:33:49.965747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T04:33:39.154387Z digest=sha256:8f66d7e309e2b87c91f9698f433f5eb8680e782bc2735266b22959e1f93084d5

Observation b7a3cfb9-8ae9-4db3-934c-c7f8b83ddd59 · outbound

This paper cites Patch-wise Attack for Fooling Deep Neural Network.

Can Sound Replace Vision in LLaVA With Token Substitution? Patch-wise Attack for Fooling Deep Neural Network

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:33:49.638510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T04:33:39.269174Z digest=sha256:5378827596fe8363d289790d011456bccb226950beb18e8c228a45e8ea3c0d02

Observation 911f3898-f876-4803-827a-5927ea5e3425 · outbound

This paper cites Explaining and Harnessing Adversarial Examples.

Can Sound Replace Vision in LLaVA With Token Substitution? Explaining and Harnessing Adversarial Examples

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:33:49.380073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T04:33:39.420094Z digest=sha256:e7778f8791047bc79697c7a21fff1ae6d2cdce07c37fd0e9270f1fcd651b0fb4

Observation 99a82fea-1f55-4916-b397-54a4a5c8839b · outbound

This paper cites Lgv: Boosting Adversarial Example Transferability from Large Geometric Vicinity.

Can Sound Replace Vision in LLaVA With Token Substitution? Lgv: Boosting Adversarial Example Transferability from Large Geometric Vicinity

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:33:49.164734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T04:33:39.571232Z digest=sha256:a8a96bbb9a607dd3b6b4572980b2371861eeb5cc69740941ef67ab6350fccee5

Observation 7ce17b0a-e15f-420e-8080-63a941d23db1 · outbound

This paper cites Deep Residual Learning for Image Recognition.

Can Sound Replace Vision in LLaVA With Token Substitution? Deep Residual Learning for Image Recognition

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:33:48.894312Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T04:33:39.684366Z digest=sha256:8c43e934f0f3d671ef423d2715a2e08c1f72260dfa3ad19b8396edb4407c028c

Observation 9549b075-33a5-4ad4-a4b7-6f400c0def6b · outbound

This paper cites Rethinking spatial dimensions of vision transformers.

Can Sound Replace Vision in LLaVA With Token Substitution? Rethinking spatial dimensions of vision transformers

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:33:48.571177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T04:33:39.799507Z digest=sha256:fdbc756c8929f77358ebd5d0f861a422bb50b3893e9756949ebe1a2aea0ed2b6

Observation 2b82a15c-a747-4962-b1dd-e75f0d75fa96 · outbound

This paper cites Densely Connected Convolutional Networks.

Can Sound Replace Vision in LLaVA With Token Substitution? Densely Connected Convolutional Networks

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:33:48.326317Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T04:33:39.931642Z digest=sha256:98741decd1d8a9a96ac00b680cf5a8a3f3a11187ff90c09804fa8ac28775c23f

Observation 0fab9f73-1bad-4f08-93e2-7d8b6fd9dde3 · outbound

This paper cites Adversarial Examples in the Physical World.

Can Sound Replace Vision in LLaVA With Token Substitution? Adversarial Examples in the Physical World

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:33:48.056715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T04:33:40.081411Z digest=sha256:bc1d2c6f574fb6f8f09b7fe0e87f57aba9e6844fb2e2a483fa24533a47552b16

Observation 0c51461e-4abf-46ed-9e5b-a428033f9080 · outbound

This paper cites Decision-based Adversarial Attack with Frequency Mixup.

Can Sound Replace Vision in LLaVA With Token Substitution? Decision-based Adversarial Attack with Frequency Mixup

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:33:47.758584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T04:33:40.227759Z digest=sha256:cd5c4d77c1288ed581080e7a1d10606179104dcfe0d455c7defc9b23888d3187

Observation 902a1a39-5021-42e5-a129-e76cf6fb6b34 · outbound

This paper cites Learning Transferable Adversarial Examples via Ghost Networks.

Can Sound Replace Vision in LLaVA With Token Substitution? Learning Transferable Adversarial Examples via Ghost Networks

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:33:47.457825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T04:33:40.380157Z digest=sha256:ed40e7bca231d938f269cae728527e78fab54daad53aa92eab06bd967ec7c3e0

Observation 1b884826-fdc7-483d-8d24-4e11505f6783 · outbound

This paper cites Nesterov Accelerated Gradient and Scale Invariance for Adversarial Attacks.

Can Sound Replace Vision in LLaVA With Token Substitution? Nesterov Accelerated Gradient and Scale Invariance for Adversarial Attacks

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:33:47.203450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T04:33:40.507235Z digest=sha256:30cf6e6eab45b4f0afd3a72c0c892d26f5180327ee10a97d686d9d534fb8f9e4

Observation 10a42f93-030b-4607-8c3f-6f6bfde9b06d · outbound

This paper cites Delving into Transferable Adversarial Examples and Black-box Attacks.

Can Sound Replace Vision in LLaVA With Token Substitution? Delving into Transferable Adversarial Examples and Black-box Attacks

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:33:46.947867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T04:33:40.663487Z digest=sha256:f535d810fd5eb7e16b53078fe40b33efa3d5c5644413648b794dd7c417e838c5

Observation 003a6926-40dc-4584-8f12-4310e49c0681 · outbound

This paper cites Swin Transformer: Hierarchical Vision Transformer using Shifted Windows.

Can Sound Replace Vision in LLaVA With Token Substitution? Swin Transformer: Hierarchical Vision Transformer using Shifted Windows

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:33:46.690725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T04:33:40.778669Z digest=sha256:b351ac1d6f069ded80f08136fcdfe26be227ee2a354eeaffea0cdcaa924d1eba

Observation b8f5898f-aae8-47a4-a262-99784bd424ad · outbound

This paper cites Frequency Domain Model Augmentation for Adversarial Attack.

Can Sound Replace Vision in LLaVA With Token Substitution? Frequency Domain Model Augmentation for Adversarial Attack

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:33:46.455924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T04:33:40.899836Z digest=sha256:cf7c3c0f2a197be14ddfc9187224abf549f06a0be568e63bc7095a32d8d70ab9

Observation 37486144-b4a5-4d59-a0dd-0606ec1aa59b · outbound

This paper cites Hierarchical vision transformers for disease progression detection in chest x-ray images.

Can Sound Replace Vision in LLaVA With Token Substitution? Hierarchical vision transformers for disease progression detection in chest x-ray images

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:33:46.196221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T04:33:41.060840Z digest=sha256:783863f7cfcc2cf9097dc2a1e2f2c11821e2b93afb9edd9f0262920d3aab302d

Observation db8fc341-514c-4879-b35a-eb2e998ce721 · outbound

This paper cites Deepfool: A Simple and Accurate Method to Fool Deep Neural Networks.

Can Sound Replace Vision in LLaVA With Token Substitution? Deepfool: A Simple and Accurate Method to Fool Deep Neural Networks

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:33:45.905787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T04:33:41.220911Z digest=sha256:b6b89cb20116b3492dfdef1231c6c0c1c37e4887ed3bd423d4c56e88de11daa9

Observation 8185f638-4e69-49e9-896b-730c5825501f · outbound

This paper cites Intriguing properties of neural networks.

Can Sound Replace Vision in LLaVA With Token Substitution? Intriguing properties of neural networks

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T04:33:41.406253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:33:41.406253Z digest=sha256:6a86813b8f10f6b04c42c36974ecd0e14d585f4f3a1d6e5cd00a8d6039299e6e

Observation f7aeaafc-7e5c-4128-9721-c41694a4a926 · outbound

This paper cites Boosting the Transferability of Adversarial Attacks with Global Momentum Initialization.

Can Sound Replace Vision in LLaVA With Token Substitution? Boosting the Transferability of Adversarial Attacks with Global Momentum Initialization

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T04:33:41.562333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:33:41.562333Z digest=sha256:f775af9ad94424f0d02ff5a94e6eb3e8ad70804478b42c278bdee648d7c64559

Observation c3357ed4-f904-4556-aaab-43385d4c34b4 · outbound

This paper cites Enhancing the Transferability of Adversarial Attacks through Variance Tuning.

Can Sound Replace Vision in LLaVA With Token Substitution? Enhancing the Transferability of Adversarial Attacks through Variance Tuning

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:33:45.653630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T04:33:41.695567Z digest=sha256:89358042a35ad9141947277fbcc853319d3fa0ab29c91a2de0b785c11e2ac363

Observation 906fdac7-92a9-4306-8443-262cdec68fa9 · outbound

This paper cites Admix: Enhancing the Transferability of Adversarial Attacks.

Can Sound Replace Vision in LLaVA With Token Substitution? Admix: Enhancing the Transferability of Adversarial Attacks

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:33:45.420097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T04:33:41.865033Z digest=sha256:9ef99a6a4c2885cb119b5147007d6da9e79b339cf986b215b629867e3ab4a200

Observation fea9f415-d335-4f43-b80f-178a3a154bf3 · outbound

This paper cites Boosting Adversarial Transferability through Enhanced Momentum.

Can Sound Replace Vision in LLaVA With Token Substitution? Boosting Adversarial Transferability through Enhanced Momentum

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:33:45.144411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T04:33:41.978382Z digest=sha256:9714f861678918425745e52630767ecd44aba502a7851e8399309acee5d5d410

Observation 739907ee-f7b9-4e18-9ec6-d6426b2e6076 · outbound

This paper cites Triangle Attack: A Query-efficient Decision-based Adversarial Attack.

Can Sound Replace Vision in LLaVA With Token Substitution? Triangle Attack: A Query-efficient Decision-based Adversarial Attack

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:33:44.796408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T04:33:42.096650Z digest=sha256:16d1e99a65e4aaeccf7c53dc04a80b51afa91ef82633d734e5b6ac26eed60d05

Observation b92c5d52-1045-4180-87e1-d56247199dca · outbound

This paper cites an unresolved cited work.

Can Sound Replace Vision in LLaVA With Token Substitution? Unresolved cited work

Reference 54

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:33:44.541442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T04:33:42.263455Z digest=sha256:9b87a1b2533a394e10327a81e5e4f4165c990e322202bc26742057c6d6267fcb

Observation d11d0327-8d0e-44bb-b31d-15da0ad856ec · outbound

This paper cites Aggregated residual transformations for deep neural networks.

Can Sound Replace Vision in LLaVA With Token Substitution? Aggregated residual transformations for deep neural networks

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:33:44.306326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T04:33:42.432082Z digest=sha256:94cbcf368e1128fb6c051cf3f54bcac4a8a4f179e51e05774ed45b3d152b0ed0

Observation d575648a-a125-4569-a48b-5f88f5be1e5e · outbound

This paper cites Stochastic Variance Reduced Ensemble Adversarial Attack for Boosting the Adversarial Transferability.

Can Sound Replace Vision in LLaVA With Token Substitution? Stochastic Variance Reduced Ensemble Adversarial Attack for Boosting the Adversarial Transferability

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:33:44.029146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T04:33:42.562240Z digest=sha256:7b335b5b38d5627720e7fe05959f83ad143404e2e33c3b56b0151e2b8ecbde9c

Observation 97d36d4d-76f2-42ea-8622-fab888ba1ffd · outbound

This paper cites Meta-learning the Search Distribution of Black-box Random Search Based Adversarial Attacks.

Can Sound Replace Vision in LLaVA With Token Substitution? Meta-learning the Search Distribution of Black-box Random Search Based Adversarial Attacks

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:33:43.749934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T04:33:42.713169Z digest=sha256:09a4d7b3ffa09ebf4e50a8c0a84d54fad030f423bafbec7e3378e5f85b97cd1b

Observation 1351d327-b91a-4381-beff-3dc01ef964ef · outbound

This paper cites Learning to transform dynamically for better adversarial transferability.

Can Sound Replace Vision in LLaVA With Token Substitution? Learning to transform dynamically for better adversarial transferability

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:33:43.472480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T04:33:42.894400Z digest=sha256:d7ea10ffceea64344e7392302111a55d1a8b287cacf6809338598f4a91cb09d0

Observation d561b2b5-bade-424b-9c6f-6b6253c02d99 · outbound

This paper cites Uia-vit: Unsupervised inconsistency-aware method based on vision transformer for face forgery detection.

Can Sound Replace Vision in LLaVA With Token Substitution? Uia-vit: Unsupervised inconsistency-aware method based on vision transformer for face forgery detection

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:33:43.260754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T04:33:43.018059Z digest=sha256:ff5fb6f8e51aac9ae2c9b02675f99056bd4b192cf6b47fab96f89041177ce5a9

Pith citing papers

No inbound Pith citation observations are available.