Pith. sign in

Paper Citation Record · LEDGER

Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models

As of 17 August 2026, this Paper Citation Record lists 37 of 37 outbound references and 1 inbound Pith citation observation for arXiv:2506.06537.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.06537 v1

Coverage vector

measured 37 of 37 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:58:02.357429Z

measured 38 of 38 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:58:02.208194Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T05:58:02.433864Z

Reference resolution

37 of 37 outbound references displayed

  • verified exact1
  • verified fuzzy30
  • unresolved6
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9c324de4-f448-45d4-9973-c134069f57ec · outbound

This paper cites This capability is essential for applica- tions in video understanding, surveillance, human-computer in- teraction, and autonomous systems.

Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models This capability is essential for applica- tions in video understanding, surveillance, human-computer in- teraction, and autonomous systems

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:58:03.137883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T05:58:02.197732Z digest=sha256:eaff7df11e2a604e209cc4f7b2d7cc627b8d12a463f1d6cc0e270bf5cefafd7e

Observation ba2b96cc-0daf-4343-abb5-a88b9eeb2afc · outbound

This paper cites Many approaches leverage cross-modal attention [1, 2, 3] combined with contrastive learning [4, 5] to align audio and visual features.

Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models Many approaches leverage cross-modal attention [1, 2, 3] combined with contrastive learning [4, 5] to align audio and visual features

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:58:03.124094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T05:58:02.203606Z digest=sha256:ad6898c2c56725f734336c68fd1ff87e8eac3e5667e4b004452c05806c9ac206

Observation 5b467b46-f159-4823-b8ab-669a8d04bc28 · outbound

This paper cites Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models.

Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-08-07T05:58:02.440658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T05:58:02.208194Z digest=sha256:299f07af4594e92d4c976a48d83199720f17a693ec3e039335d0d1a95ffc10ce

Observation 05360814-6668-4846-b41d-7cdc1d1a6efe · outbound

This paper cites Models Below, we elaborate on the final models built for each approach.

Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models Models Below, we elaborate on the final models built for each approach

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:58:03.110256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T05:58:02.213087Z digest=sha256:74be2df858492c1deb31b3c0cd121d21ad3079403d8e0fccb12a23000172b77a

Observation 7266a948-ce5a-4a2f-97ae-f12a5361c794 · outbound

This paper cites By systematically integrating audio, vision, and text representations, we bridge modality gaps and achieve fine-grained segmentation without requiring A VS-specific annotations.

Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models By systematically integrating audio, vision, and text representations, we bridge modality gaps and achieve fine-grained segmentation without requiring A VS-specific annotations

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:58:03.096865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T05:58:02.218012Z digest=sha256:324dba3cb714ff175322f64f17e2358cad57b53bb31f535838f66b76ce2dd984

Observation 45b681cf-81c4-47a9-8c43-0945a0339f6f · outbound

This paper cites an unresolved cited work.

Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:58:03.083300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T05:58:02.222340Z digest=sha256:d03bc5f5c9c70bfdc9184d42a1f10b35c45d25077b76523a9bf9e5a0569db9a5

Observation 51b6b9a9-a56e-40c4-a262-d9929e24d7fe · outbound

This paper cites Learning to localize sound sources in visual scenes: Analysis and applications,.

Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models Learning to localize sound sources in visual scenes: Analysis and applications,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:58:03.069950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T05:58:02.227428Z digest=sha256:a6e0fe0183597dace6564253b5b12c9e2992ad16857af1baaa662b7729350bc0

Observation dbc8d8e5-6e64-4c80-844f-d0602e249a49 · outbound

This paper cites Aligning Sight and Sound: Advanced Sound Source Localization Through Audio-Visual Alignment.

Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models Aligning Sight and Sound: Advanced Sound Source Localization Through Audio-Visual Alignment

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T05:58:02.231835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:58:02.231835Z digest=sha256:d5edc2c866bc3aff27d0426bd48de0433cf18ed2a68beb3eb0e11edb2f060630

Observation ea2187a2-3ab3-4113-9ef0-27d3b535398a · outbound

This paper cites Exploiting transformation in- variance and equivariance for self-supervised sound localisation,.

Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models Exploiting transformation in- variance and equivariance for self-supervised sound localisation,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:58:03.057441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T05:58:02.236552Z digest=sha256:5ee4705bf4ec13fb3bb0ba6f4a253b6bec5483edc5fe0430b1258d2605ee910d

Observation 61908570-0221-42fd-89a2-e50426373aeb · outbound

This paper cites Localizing visual sounds the easy way,.

Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models Localizing visual sounds the easy way,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:58:03.044572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T05:58:02.241086Z digest=sha256:11cc83caefa6901ab72e7ede50c22d1a2d0367fd5d6fa33d4b9f6cb8da162983

Observation 187b50d7-1b94-47c1-a2b1-a173a5565f20 · outbound

This paper cites Localizing visual sounds the hard way,.

Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models Localizing visual sounds the hard way,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:58:03.032265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T05:58:02.246202Z digest=sha256:e275bd23449c3dc444337f6cbcd49ccbf60304a89e11be2af8ac42c274cfd3ae

Observation 8330852f-ff89-4c02-b9ba-acd7124ae7d0 · outbound

This paper cites Learning audio-visual source local- ization via false negative aware contrastive learning,.

Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models Learning audio-visual source local- ization via false negative aware contrastive learning,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:58:03.019857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T05:58:02.250492Z digest=sha256:1f879c24652a8a1eb3cb2cc7f48a691b3f06a4c10ce19bcfefad95587114a773

Observation fe06a59f-5fdb-47ab-a449-d45ab740654f · outbound

This paper cites Marginnce: Robust sound localization with a negative margin,.

Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models Marginnce: Robust sound localization with a negative margin,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:58:03.006909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T05:58:02.255202Z digest=sha256:b8be3449712f188df84454e314a566f422492e304275f5f5be061f26633d878c

Observation c2d24cae-a9bd-49af-86f2-29dc969e99db · outbound

This paper cites Audio–visual segmen- tation,.

Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models Audio–visual segmen- tation,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:58:02.994261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T05:58:02.259409Z digest=sha256:90582dd7c5c6ddc3de18e657bf8e34d4243b2f368d9424b653af802b3e30a569

Observation a97a0451-fba9-402e-b8a0-3e358bed884f · outbound

This paper cites Improving audio-visual segmentation with bidirectional generation,.

Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models Improving audio-visual segmentation with bidirectional generation,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:58:02.981778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T05:58:02.263449Z digest=sha256:84ee16e1b0a9d7fdab55bdc42fe0a25bbbe4c30c308152e6e9464e23ac3cb86d

Observation ea52bc2d-7569-4054-80fd-e8d55ad2b4fb · outbound

This paper cites Selm: Selective mechanism based audio-visual segmentation,.

Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models Selm: Selective mechanism based audio-visual segmentation,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:58:02.910728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T05:58:02.267534Z digest=sha256:1865f19df8d0a44277a81640dc55221be662f4803cecfce87791828b9255cdfe

Observation 110a07f7-5a14-41a5-a5a0-4d25e7cd41e2 · outbound

This paper cites Bavs: Bootstrapping audio-visual segmentation by integrating foundation knowledge,.

Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models Bavs: Bootstrapping audio-visual segmentation by integrating foundation knowledge,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:58:02.801448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T05:58:02.271953Z digest=sha256:e5ccdd6dc044f73bc408a7f20ec8f804b899c04b70df3035417d2976b64c18ef

Observation b0f3a086-d99a-4b58-add1-057deb845d7d · outbound

This paper cites Learning transferable visual models from natural language supervision,.

Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models Learning transferable visual models from natural language supervision,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:58:02.755960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T05:58:02.276213Z digest=sha256:04efa2b7d4e5971957735fb68a4da72cb1b9ef5d640ffbae0cbd77f52e2d0ef7

Observation b97ce92b-dd08-458a-a942-8ba4a2e442af · outbound

This paper cites Natural Language Supervision for General-Purpose Audio Representations.

Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models Natural Language Supervision for General-Purpose Audio Representations

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T05:58:02.280500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:58:02.280500Z digest=sha256:eb9da3c7cd9e8df0505f003b8b7adc9f464f99d5efeab96cc671466f9fee80d7

Observation b957c4a4-d464-45c3-8a17-463ab759847e · outbound

This paper cites WavCaps: A ChatGPT-assisted weakly-labelled audio captioning dataset for audio-language mul- timodal research,.

Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models WavCaps: A ChatGPT-assisted weakly-labelled audio captioning dataset for audio-language mul- timodal research,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:58:02.714559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T05:58:02.285239Z digest=sha256:150aeef1c5b20386928be878b459092f20e37ada379bc8ce64f85bced5f43566

Observation b8a97d62-a16f-4e8d-9b2a-5cfb5cca0335 · outbound

This paper cites AudioCaps: Generating Captions for Audios in The Wild,.

Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models AudioCaps: Generating Captions for Audios in The Wild,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:58:02.682158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T05:58:02.289939Z digest=sha256:5ac2db2bb9b4fc764d5012021f140ec9bbdc766ef9b232de7e840a1aa0ca2881

Observation 5d9c8752-8a7a-40f4-a793-eda42895a494 · outbound

This paper cites Microsoft coco: Common objects in context,.

Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models Microsoft coco: Common objects in context,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:58:02.646312Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T05:58:02.294346Z digest=sha256:0c6cfe76903dfe77d321b0df01305e19460c010c8eca1ee456b99b8aeb9c6458

Observation c4704ecb-6a7c-4a03-bfd6-782595e87095 · outbound

This paper cites Blip: Bootstrapping language- image pre-training for unified vision-language understanding and generation,.

Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models Blip: Bootstrapping language- image pre-training for unified vision-language understanding and generation,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:58:02.626754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T05:58:02.298465Z digest=sha256:cc99999c4167beb6afbf0f4e628a4a6f1589af0481fcbe7176385cb0997106a9

Observation 3cead3a2-851f-41a0-9133-c6a0c9ce0374 · outbound

This paper cites An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual Inversion.

Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual Inversion

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T05:58:02.302544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:58:02.302544Z digest=sha256:3d88c774b86718f601914c3f918f3401961c171307680df96b8133164051f38b

Observation cdcf4a11-9322-438d-b2be-c3d27ffbf35f · outbound

This paper cites Learning to visually localize sound sources from mixtures without prior source knowl- edge,.

Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models Learning to visually localize sound sources from mixtures without prior source knowl- edge,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:58:02.613063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T05:58:02.307043Z digest=sha256:f801576d0cde64f4fd33a1d28623c781207c8c509f40c3c628632c775cf8ff28

Observation ef3a48c4-44d5-4ca7-923b-66227211d06f · outbound

This paper cites Adaptive selection based referring im- age segmentation,.

Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models Adaptive selection based referring im- age segmentation,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:58:02.599272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T05:58:02.310995Z digest=sha256:7603f05bca7a9ecafd7dc2b7e7a85d4627d8c0a2446ceacf9f2eec6f1bd98a94

Observation a1d42340-6c74-4b35-a8e2-918242b51871 · outbound

This paper cites Beats: Audio pre-training with acoustic tokenizers,.

Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models Beats: Audio pre-training with acoustic tokenizers,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:58:02.585955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T05:58:02.315084Z digest=sha256:11915a529d6c4b3d9aa82cc85674e5b5aa66b4139895f36b11cc71d1e2a1e96c

Observation 3955a45d-4ea2-4ede-af26-ed42f347f5f7 · outbound

This paper cites Audio set: An ontology and human-labeled dataset for audio events,.

Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models Audio set: An ontology and human-labeled dataset for audio events,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T05:58:02.319071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:58:02.319071Z digest=sha256:6f9438e1564dacec57fa21986e8f1e28052953662636ea04ca6d0ff7a034e9bf

Observation 7e948701-6ee2-4222-99c9-14d9ced8b76b · outbound

This paper cites Clotho: An audio cap- tioning dataset,.

Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models Clotho: An audio cap- tioning dataset,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:58:02.563661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T05:58:02.323482Z digest=sha256:bfe05932f33018912910e7c1474b32b579379b42d27fa0f05e1015488f09af1f

Observation 809d8bfc-ead5-43fe-8513-94e268eb35f4 · outbound

This paper cites spaCy 2: Natural language under- standing with Bloom embeddings, convolutional neural networks and incremental parsing,.

Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models spaCy 2: Natural language under- standing with Bloom embeddings, convolutional neural networks and incremental parsing,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:58:02.550108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T05:58:02.327698Z digest=sha256:776a2d94a507025ba728c46936fdc6965d0ca08bb6898c0e6cee16686c44efeb

Observation 99888b15-3398-4842-943c-7ca0b879f59c · outbound

This paper cites High-resolution image synthesis with latent diffusion models,.

Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models High-resolution image synthesis with latent diffusion models,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T05:58:02.331753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:58:02.331753Z digest=sha256:18bca56962fa90b52d73ee319344a3192b8da9b4201fe1699710500d6bd4d247

Observation 406847a3-bac8-415e-b23d-bce3dce0f5e7 · outbound

This paper cites Vggsound: A large-scale audio-visual dataset,.

Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models Vggsound: A large-scale audio-visual dataset,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:58:02.526857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T05:58:02.335799Z digest=sha256:6cf57828f04c17d0be2575f890847661c13f1e1bc2b2e9ee9214e5e475e368db

Observation 96a2c20c-d4bb-4263-b611-9c364e4e3414 · outbound

This paper cites Segment anything,.

Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models Segment anything,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:58:02.512415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T05:58:02.340053Z digest=sha256:a7b70492a5bffbc8941cacd42f359d56b23584bca720f53f7ec3cb865095652f

Observation 3dec38ed-1032-4ad8-8887-aca36f518f87 · outbound

This paper cites Unraveling instance associations: A closer look for audio-visual segmentation,.

Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models Unraveling instance associations: A closer look for audio-visual segmentation,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:58:02.497818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T05:58:02.344260Z digest=sha256:8c9a6cf9758e21690a37d7548498b1ca358665480223e573a028acc5af3cb35b

Observation 005345c1-4e4b-48e5-b82e-acfb4c9df8b5 · outbound

This paper cites A closer look at weakly-supervised audio-visual source localization,.

Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models A closer look at weakly-supervised audio-visual source localization,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:58:02.483950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T05:58:02.348596Z digest=sha256:e51984904154936e904b28105f6dc6d3c1fadc7fa246be6771a11c5a1d4ec6a8

Observation 9f67a6e9-214e-4beb-8c99-694d45256cc4 · outbound

This paper cites Bridg- ing vision and language encoders: Parameter-efficient tuning for referring image segmentation,.

Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models Bridg- ing vision and language encoders: Parameter-efficient tuning for referring image segmentation,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:58:02.469077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T05:58:02.352781Z digest=sha256:12ab248085b03b74328ee20020f71ae927fcdb91c0d2704e792455b56acc2087

Observation f16500da-de98-4eac-9740-46fc5134d510 · outbound

This paper cites Cris: Clip-driven referring image segmentation,.

Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models Cris: Clip-driven referring image segmentation,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:58:02.455112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T05:58:02.357429Z digest=sha256:8af36f538e5e8ca4114e6dda015b21d103f8dc438ed696b631306230fc18790e

Pith citing papers

Observation 5b467b46-f159-4823-b8ab-669a8d04bc28 · inbound

Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models cites this paper.

Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-08-07T05:58:02.440658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T05:58:02.208194Z digest=sha256:299f07af4594e92d4c976a48d83199720f17a693ec3e039335d0d1a95ffc10ce