Pith. sign in

Paper Citation Record · LEDGER

Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models

As of 9 August 2026, this Paper Citation Record lists 37 of 37 outbound references and 1 inbound Pith citation observation for arXiv:2506.06537.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.06537 v1

Coverage vector

measured 37 of 37 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:58:02.357429Z

measured 38 of 38 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:58:02.208194Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T05:58:02.433864Z

Reference resolution

37 of 37 outbound references displayed

  • verified exact1
  • verified fuzzy30
  • unresolved6
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9c324de4-f448-45d4-9973-c134069f57ec · outbound

This paper cites This capability is essential for applica- tions in video understanding, surveillance, human-computer in- teraction, and autonomous systems.

Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models This capability is essential for applica- tions in video understanding, surveillance, human-computer in- teraction, and autonomous systems

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:58:03.137883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:58:02.197732Z digest=sha256:b70cf7c0cc47cd13b4688f69da9f2adf24ae903307147261dca5bbd15d5f669f

Observation ba2b96cc-0daf-4343-abb5-a88b9eeb2afc · outbound

This paper cites Many approaches leverage cross-modal attention [1, 2, 3] combined with contrastive learning [4, 5] to align audio and visual features.

Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models Many approaches leverage cross-modal attention [1, 2, 3] combined with contrastive learning [4, 5] to align audio and visual features

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:58:03.124094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:58:02.203606Z digest=sha256:301bde84492001af07ea7f6b4a62c3f4423b8e364e275df921e2a5d7160bbab6

Observation 5b467b46-f159-4823-b8ab-669a8d04bc28 · outbound

This paper cites Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models.

Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-08-07T05:58:02.440658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:58:02.208194Z digest=sha256:8701c8d796f3d4af63ab16e44bee00c80eba0bebb2946543be8905d192f7cb61

Observation 05360814-6668-4846-b41d-7cdc1d1a6efe · outbound

This paper cites Models Below, we elaborate on the final models built for each approach.

Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models Models Below, we elaborate on the final models built for each approach

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:58:03.110256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:58:02.213087Z digest=sha256:eee58a0df4009f6e3cbfd3d2652e3f5b189403161dc5cc54a8f786999a504cd2

Observation 7266a948-ce5a-4a2f-97ae-f12a5361c794 · outbound

This paper cites By systematically integrating audio, vision, and text representations, we bridge modality gaps and achieve fine-grained segmentation without requiring A VS-specific annotations.

Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models By systematically integrating audio, vision, and text representations, we bridge modality gaps and achieve fine-grained segmentation without requiring A VS-specific annotations

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:58:03.096865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:58:02.218012Z digest=sha256:076c1079ac9bcc6bd77fb98b1fc7a6e6d7c037b66f4ebcdaf5af0547e3834edb

Observation 45b681cf-81c4-47a9-8c43-0945a0339f6f · outbound

This paper cites an unresolved cited work.

Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:58:03.083300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:58:02.222340Z digest=sha256:23fcc3aa7b8c064663db1ed9f9c8d282d992c44337a3dbe1c05a6d78f4efd86d

Observation 51b6b9a9-a56e-40c4-a262-d9929e24d7fe · outbound

This paper cites Learning to localize sound sources in visual scenes: Analysis and applications,.

Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models Learning to localize sound sources in visual scenes: Analysis and applications,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:58:03.069950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:58:02.227428Z digest=sha256:32fb7b09efaef8883d42396f9b9bcc8812d8457ff58cd866da9dd1e640438ed9

Observation dbc8d8e5-6e64-4c80-844f-d0602e249a49 · outbound

This paper cites Aligning Sight and Sound: Advanced Sound Source Localization Through Audio-Visual Alignment.

Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models Aligning Sight and Sound: Advanced Sound Source Localization Through Audio-Visual Alignment

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T05:58:02.231835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:58:02.231835Z digest=sha256:640b096b3de41d203ddfc6f3527b61f0608cafd1fc2b994cbdd7be938840bf4f

Observation ea2187a2-3ab3-4113-9ef0-27d3b535398a · outbound

This paper cites Exploiting transformation in- variance and equivariance for self-supervised sound localisation,.

Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models Exploiting transformation in- variance and equivariance for self-supervised sound localisation,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:58:03.057441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:58:02.236552Z digest=sha256:a9a9f6d83f2570d4a56f15e43643f852a179ddbe17c2fe6126eb79883135c074

Observation 61908570-0221-42fd-89a2-e50426373aeb · outbound

This paper cites Localizing visual sounds the easy way,.

Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models Localizing visual sounds the easy way,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:58:03.044572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:58:02.241086Z digest=sha256:fd2e49353935dab77a68446a5fb92af2d24725047c5b1b217a0a022d0d263dcd

Observation 187b50d7-1b94-47c1-a2b1-a173a5565f20 · outbound

This paper cites Localizing visual sounds the hard way,.

Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models Localizing visual sounds the hard way,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:58:03.032265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:58:02.246202Z digest=sha256:8cc9f153c4457397c2e41941c145a96fd2e43a11a276f0a2e9058b1da52f0d81

Observation 8330852f-ff89-4c02-b9ba-acd7124ae7d0 · outbound

This paper cites Learning audio-visual source local- ization via false negative aware contrastive learning,.

Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models Learning audio-visual source local- ization via false negative aware contrastive learning,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:58:03.019857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:58:02.250492Z digest=sha256:9c6b52034c96026154c90fe8b414b96095d0699ec63eabdd600e5c99ae0e3260

Observation fe06a59f-5fdb-47ab-a449-d45ab740654f · outbound

This paper cites Marginnce: Robust sound localization with a negative margin,.

Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models Marginnce: Robust sound localization with a negative margin,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:58:03.006909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:58:02.255202Z digest=sha256:5a751e6c29b87036fd01f59788d8cbf729fd1f7f24b787712ccebce2902a8240

Observation c2d24cae-a9bd-49af-86f2-29dc969e99db · outbound

This paper cites Audio–visual segmen- tation,.

Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models Audio–visual segmen- tation,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:58:02.994261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:58:02.259409Z digest=sha256:e7bd563fce1b1fd403411658fd2443ad99ddafac7a499a69f977c6f7e7221fa0

Observation a97a0451-fba9-402e-b8a0-3e358bed884f · outbound

This paper cites Improving audio-visual segmentation with bidirectional generation,.

Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models Improving audio-visual segmentation with bidirectional generation,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:58:02.981778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:58:02.263449Z digest=sha256:f2f382fe69a6d4f2362eb1b69e3cd92ac2b70d8f29fad5b9844134983b662215

Observation ea52bc2d-7569-4054-80fd-e8d55ad2b4fb · outbound

This paper cites Selm: Selective mechanism based audio-visual segmentation,.

Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models Selm: Selective mechanism based audio-visual segmentation,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:58:02.910728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:58:02.267534Z digest=sha256:0495a48aa3af2e0b94469fb8fb9900026df5a760d89cd2bb2088eef5d175a344

Observation 110a07f7-5a14-41a5-a5a0-4d25e7cd41e2 · outbound

This paper cites Bavs: Bootstrapping audio-visual segmentation by integrating foundation knowledge,.

Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models Bavs: Bootstrapping audio-visual segmentation by integrating foundation knowledge,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:58:02.801448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:58:02.271953Z digest=sha256:c1f533c9f49de08934afea878f398e7bb59ba406f4a7375d58b889bb66a27198

Observation b0f3a086-d99a-4b58-add1-057deb845d7d · outbound

This paper cites Learning transferable visual models from natural language supervision,.

Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models Learning transferable visual models from natural language supervision,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:58:02.755960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:58:02.276213Z digest=sha256:63e7179e1bd874c1dd7def3327955f33aaf6d41e87e0fb16857095945dafd380

Observation b97ce92b-dd08-458a-a942-8ba4a2e442af · outbound

This paper cites Natural Language Supervision for General-Purpose Audio Representations.

Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models Natural Language Supervision for General-Purpose Audio Representations

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T05:58:02.280500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:58:02.280500Z digest=sha256:ec55eaa8621c54a15d63ba56c4a969d242ffcc3091a936ab6e30632a6224105a

Observation b957c4a4-d464-45c3-8a17-463ab759847e · outbound

This paper cites WavCaps: A ChatGPT-assisted weakly-labelled audio captioning dataset for audio-language mul- timodal research,.

Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models WavCaps: A ChatGPT-assisted weakly-labelled audio captioning dataset for audio-language mul- timodal research,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:58:02.714559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:58:02.285239Z digest=sha256:456fd9f0a2e87f256b42ef2d378b1827b87f336959c0e96a7b1241803355c5d7

Observation b8a97d62-a16f-4e8d-9b2a-5cfb5cca0335 · outbound

This paper cites AudioCaps: Generating Captions for Audios in The Wild,.

Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models AudioCaps: Generating Captions for Audios in The Wild,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:58:02.682158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:58:02.289939Z digest=sha256:0d0b184dfca1b61791a16136be689f0ab320b94056c2ec9b4153c440be38f469

Observation 5d9c8752-8a7a-40f4-a793-eda42895a494 · outbound

This paper cites Microsoft coco: Common objects in context,.

Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models Microsoft coco: Common objects in context,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:58:02.646312Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:58:02.294346Z digest=sha256:d8e3da92cbbc0f37a4ea19587aacb3e92ccec4ef929f44795f8108a10e37ae70

Observation c4704ecb-6a7c-4a03-bfd6-782595e87095 · outbound

This paper cites Blip: Bootstrapping language- image pre-training for unified vision-language understanding and generation,.

Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models Blip: Bootstrapping language- image pre-training for unified vision-language understanding and generation,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:58:02.626754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:58:02.298465Z digest=sha256:8689cca82376da11ce61217922d5184b860a66eebe9cc58da45ac26d5d62c54a

Observation 3cead3a2-851f-41a0-9133-c6a0c9ce0374 · outbound

This paper cites An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual Inversion.

Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual Inversion

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T05:58:02.302544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:58:02.302544Z digest=sha256:92a6a2c698795d1512a1754e5f3b45fc41078332b825aed07449ac343c858873

Observation cdcf4a11-9322-438d-b2be-c3d27ffbf35f · outbound

This paper cites Learning to visually localize sound sources from mixtures without prior source knowl- edge,.

Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models Learning to visually localize sound sources from mixtures without prior source knowl- edge,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:58:02.613063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:58:02.307043Z digest=sha256:ac0e4114d281782afde0e8a77846c6ed52f834418d233cf55cd4e1317013a045

Observation ef3a48c4-44d5-4ca7-923b-66227211d06f · outbound

This paper cites Adaptive selection based referring im- age segmentation,.

Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models Adaptive selection based referring im- age segmentation,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:58:02.599272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:58:02.310995Z digest=sha256:65a431e4c2bd5160af3fad22df24b02979368b1dd5b883beb6423bef9d96c7ba

Observation a1d42340-6c74-4b35-a8e2-918242b51871 · outbound

This paper cites Beats: Audio pre-training with acoustic tokenizers,.

Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models Beats: Audio pre-training with acoustic tokenizers,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:58:02.585955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:58:02.315084Z digest=sha256:7e946247396c4c2cdd5d9a052714fe57cfb875e1520b43b307ec6af654b73051

Observation 3955a45d-4ea2-4ede-af26-ed42f347f5f7 · outbound

This paper cites Audio set: An ontology and human-labeled dataset for audio events,.

Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models Audio set: An ontology and human-labeled dataset for audio events,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T05:58:02.319071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:58:02.319071Z digest=sha256:a6ce38b9f9a51daa186ae6abfee5865876438c8041a1928aa2ea49dadbc9be6d

Observation 7e948701-6ee2-4222-99c9-14d9ced8b76b · outbound

This paper cites Clotho: An audio cap- tioning dataset,.

Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models Clotho: An audio cap- tioning dataset,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:58:02.563661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:58:02.323482Z digest=sha256:5b3837599c1914bc8f2b27b2c5a50d473e1c2b4c0efdb9e0d0a57567fc3a5c6c

Observation 809d8bfc-ead5-43fe-8513-94e268eb35f4 · outbound

This paper cites spaCy 2: Natural language under- standing with Bloom embeddings, convolutional neural networks and incremental parsing,.

Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models spaCy 2: Natural language under- standing with Bloom embeddings, convolutional neural networks and incremental parsing,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:58:02.550108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:58:02.327698Z digest=sha256:70f19f9dbcb5f6991c0ddafec1078ce65e575234b2ab0bb91c0df2830f76cd14

Observation 99888b15-3398-4842-943c-7ca0b879f59c · outbound

This paper cites High-resolution image synthesis with latent diffusion models,.

Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models High-resolution image synthesis with latent diffusion models,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T05:58:02.331753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:58:02.331753Z digest=sha256:e6419a26792787bc2a47e0ad261db9055354a8a0616fdbacd4a9839e7090d8bf

Observation 406847a3-bac8-415e-b23d-bce3dce0f5e7 · outbound

This paper cites Vggsound: A large-scale audio-visual dataset,.

Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models Vggsound: A large-scale audio-visual dataset,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:58:02.526857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:58:02.335799Z digest=sha256:db2f7c86eb392f8b6b2869563b9a39e6e1601dd0a7dd3bf83381cc28081ce202

Observation 96a2c20c-d4bb-4263-b611-9c364e4e3414 · outbound

This paper cites Segment anything,.

Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models Segment anything,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:58:02.512415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:58:02.340053Z digest=sha256:b81deed212fed85a4f77b70b689236d3d04f8d9b8cd6134648dabb8212660bb3

Observation 3dec38ed-1032-4ad8-8887-aca36f518f87 · outbound

This paper cites Unraveling instance associations: A closer look for audio-visual segmentation,.

Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models Unraveling instance associations: A closer look for audio-visual segmentation,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:58:02.497818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:58:02.344260Z digest=sha256:6ac0a11b748981304c2e76f44cfa33dd2c6bcb8d75669548d4ce4769adc06ea8

Observation 005345c1-4e4b-48e5-b82e-acfb4c9df8b5 · outbound

This paper cites A closer look at weakly-supervised audio-visual source localization,.

Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models A closer look at weakly-supervised audio-visual source localization,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:58:02.483950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:58:02.348596Z digest=sha256:17a824bcd71c36acb0a0426928734c18c065c5cbe8e6e4eb3e2ed082368de178

Observation 9f67a6e9-214e-4beb-8c99-694d45256cc4 · outbound

This paper cites Bridg- ing vision and language encoders: Parameter-efficient tuning for referring image segmentation,.

Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models Bridg- ing vision and language encoders: Parameter-efficient tuning for referring image segmentation,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:58:02.469077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:58:02.352781Z digest=sha256:f45277e2c44d33bc9cfc36a6e8ee03ae9d8c01cbc4dc430bd7b79329c4303e71

Observation f16500da-de98-4eac-9740-46fc5134d510 · outbound

This paper cites Cris: Clip-driven referring image segmentation,.

Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models Cris: Clip-driven referring image segmentation,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:58:02.455112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:58:02.357429Z digest=sha256:4a02f2e094a00703fb371e757c08ad6734ea7096af89ad1cdcc4e7903e6b7ed9

Pith citing papers

Observation 5b467b46-f159-4823-b8ab-669a8d04bc28 · inbound

Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models cites this paper.

Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-08-07T05:58:02.440658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:58:02.208194Z digest=sha256:8701c8d796f3d4af63ab16e44bee00c80eba0bebb2946543be8905d192f7cb61