Pith. sign in

Paper Citation Record · LEDGER

Learning from Silence and Noise for Visual Sound Source Localization

As of 8 August 2026, this Paper Citation Record lists 69 of 69 outbound references and 0 inbound Pith citation observations for arXiv:2508.21761.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.21761 v1

Coverage vector

measured 69 of 69 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T14:03:22.242986Z

measured 69 of 69 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

69 of 69 outbound references displayed

  • verified exact6
  • verified fuzzy56
  • unresolved6
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3e4a7fab-55d5-457f-86a4-8f3593837185 · outbound

This paper cites Adobe audition sound effects, 2023.

Learning from Silence and Noise for Visual Sound Source Localization Adobe audition sound effects, 2023

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:34.971984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T14:03:15.918896Z digest=sha256:270f692f34857c5d0ee4afc135d464e4e82af854645db82a8f6b9b9ef39729bb

Observation 2487fb81-ca58-4a6e-9bbc-3c103ff6edd9 · outbound

This paper cites Self- supervised learning of audio-visual objects from video.

Learning from Silence and Noise for Visual Sound Source Localization Self- supervised learning of audio-visual objects from video

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:34.794728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T14:03:16.037998Z digest=sha256:b1f7d3235fde5fbf925ee2e281bf7eb6d789177fa659784e91e4dd1639345b9c

Observation 167a7fe3-becb-4103-88ac-23a71fd954ee · outbound

This paper cites Look, listen and learn.

Learning from Silence and Noise for Visual Sound Source Localization Look, listen and learn

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:34.604618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T14:03:16.212958Z digest=sha256:e61226f6d4653fcb6401688c7ec2108c4368f300df427286b01a41b412e29503

Observation 36d7af9e-35fc-4dd2-a2c6-11d63e7e929c · outbound

This paper cites Objects that sound.

Learning from Silence and Noise for Visual Sound Source Localization Objects that sound

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:34.384136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T14:03:16.317154Z digest=sha256:d91c1fcfdee3e876eac2b004b67f9494809409b07e75e8070ad0b35bf434448f

Observation 06da2763-a6f0-434b-915e-20afb4a4cf7f · outbound

This paper cites Vggsound: A large-scale audio-visual dataset.

Learning from Silence and Noise for Visual Sound Source Localization Vggsound: A large-scale audio-visual dataset

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:34.185269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T14:03:16.462096Z digest=sha256:4c67df28caf93a254d1d4d182d9fa832c323b53b3afc285053be9b4c4cb02042

Observation 8f8dccc0-9e7d-48aa-8820-e48257bbecff · outbound

This paper cites Localizing visual sounds the hard way.

Learning from Silence and Noise for Visual Sound Source Localization Localizing visual sounds the hard way

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:33.877814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T14:03:16.662025Z digest=sha256:a8215fab6ba21158eee311ee6e62dad7b2633942e13b930c64303db423c60e68

Observation ce4379dd-7053-4613-9993-a857258bcb61 · outbound

This paper cites A simple framework for contrastive learning of visual representations.

Learning from Silence and Noise for Visual Sound Source Localization A simple framework for contrastive learning of visual representations

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T14:03:16.836987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:03:16.836987Z digest=sha256:ffd75ef921646e8bd9ad4067c7dbb87088ac8331807273a463f39001802504df

Observation 9612ced4-e39e-4d8f-98e7-88a8bb288676 · outbound

This paper cites Exploring simple siamese representation learning.

Learning from Silence and Noise for Visual Sound Source Localization Exploring simple siamese representation learning

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:33.654480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T14:03:16.989821Z digest=sha256:7754f2f23a7ae7811fd979569d168865558462caf3267ce6873ea4c3d6ec4adf

Observation 18f1a1bf-66a1-46a6-84d2-e8af7158bfb4 · outbound

This paper cites Integrating Audio, Visual, and Semantic Information for Enhanced Multimodal Speaker Diarization.

Learning from Silence and Noise for Visual Sound Source Localization Integrating Audio, Visual, and Semantic Information for Enhanced Multimodal Speaker Diarization

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-08-05T14:03:23.762550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T14:03:17.100791Z digest=sha256:b29576ab3dfac361fd1af39b07492e17287d81489a15492217b2e57bf723327c

Observation 0a716c15-8ebd-4fd0-88bb-34b2ee5bbf71 · outbound

This paper cites Learning a similarity metric discrimi- natively, with application to face verification.

Learning from Silence and Noise for Visual Sound Source Localization Learning a similarity metric discrimi- natively, with application to face verification

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:33.487670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T14:03:17.193287Z digest=sha256:17b2406cf5f121dade47f00a548e9abdb530bb8a53431bbc33e3078803979980

Observation e14f1e04-fc7c-4f20-a678-de813f45fdb2 · outbound

This paper cites Imagenet: A large-scale hierarchical image database.

Learning from Silence and Noise for Visual Sound Source Localization Imagenet: A large-scale hierarchical image database

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:33.265198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T14:03:17.318838Z digest=sha256:86ea339321fa8c65fb422909a7e839acf772bfcafa34c91da54b4e2af6dd2d89

Observation d975b22f-4ca3-4589-a2e4-72ee818513cb · outbound

This paper cites Condi- tional generation of audio from video via foley analogies.

Learning from Silence and Noise for Visual Sound Source Localization Condi- tional generation of audio from video via foley analogies

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:33.105208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T14:03:17.370379Z digest=sha256:3d81c19d371dfd257d65d820a0ebc1db3bfda20591d85b2355b8ff851b308ad2

Observation 289e305d-b66e-4ba4-bce0-e02e77410ff9 · outbound

This paper cites Audio-Visual Approach For Multimodal Concurrent Speaker Detection.

Learning from Silence and Noise for Visual Sound Source Localization Audio-Visual Approach For Multimodal Concurrent Speaker Detection

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-08-05T14:03:23.586547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T14:03:17.449504Z digest=sha256:39b58482149bcd0ac7ec7aef0a2b3b5752a2812612d20bc1eeae356f3fc67d2d

Observation 4ab82633-9016-4ffb-81d5-3db50d0ec276 · outbound

This paper cites Effect of acoustic scene complexity and visual scene representation on auditory perception in virtual audio-visual environments.

Learning from Silence and Noise for Visual Sound Source Localization Effect of acoustic scene complexity and visual scene representation on auditory perception in virtual audio-visual environments

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:32.961414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T14:03:17.494864Z digest=sha256:a076cc9af3d3fc096d5a33509a3cb45b4f3b7f868a5ef50de5b70c951d3f0a11

Observation 7035228b-f75d-47ad-8bb6-095d71fefd7f · outbound

This paper cites Learning joint sta- tistical models for audio-visual fusion and segregation.

Learning from Silence and Noise for Visual Sound Source Localization Learning joint sta- tistical models for audio-visual fusion and segregation

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:32.769252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T14:03:17.608532Z digest=sha256:ffee1043827519347eb664a2347848f3f7aead24385730c9190590e09eecef77

Observation 21b16037-084b-48f8-9727-8994e6d63e22 · outbound

This paper cites Visualvoice: Audio-visual speech separation with cross-modal consistency.

Learning from Silence and Noise for Visual Sound Source Localization Visualvoice: Audio-visual speech separation with cross-modal consistency

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:32.569106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T14:03:17.717130Z digest=sha256:3aa2b91d97e6a309a8487b0230c63d1dc288be88ef2d1eaf1af8f6c7fd7d042b

Observation 7a8dccc9-a42b-4a45-9b98-5a25c45fe29f · outbound

This paper cites Cyclip: Cyclic contrastive language-image pretraining.

Learning from Silence and Noise for Visual Sound Source Localization Cyclip: Cyclic contrastive language-image pretraining

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:32.373662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T14:03:17.817540Z digest=sha256:03428679030cd248d378bb4b0d31680325ca294c3a7652638ab11608e1a58a73

Observation 555accc2-bd1e-4443-be62-406557953cf7 · outbound

This paper cites Bootstrap your own latent-a new approach to self-supervised learning.

Learning from Silence and Noise for Visual Sound Source Localization Bootstrap your own latent-a new approach to self-supervised learning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T14:03:17.872247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:03:17.872247Z digest=sha256:0f3af2e8e9642a589a847effa6d63dc6f9ae17d0a77d25e74e9d88b14217a1f6

Observation 9117824f-d757-446f-a13a-c5ca6b004c8a · outbound

This paper cites chirp" from the.

Learning from Silence and Noise for Visual Sound Source Localization chirp" from the

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:32.139912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T14:03:17.970801Z digest=sha256:1d89281d6c45c8e783931dbe2514c4a2f116928f567799dfce4921e2c9130258

Observation c1f857d6-a920-4f0c-8c65-9a2f577b3ec1 · outbound

This paper cites Canonical correlation analysis: An overview with application to learning methods.

Learning from Silence and Noise for Visual Sound Source Localization Canonical correlation analysis: An overview with application to learning methods

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:31.989916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T14:03:18.046880Z digest=sha256:ddf1c41446c711e568e10637609a19b1a7a59640efc87b1eabf23c93d1163203

Observation f9ac9d16-c864-4f9a-874c-84270023d2d5 · outbound

This paper cites Deep residual learning for image recognition.

Learning from Silence and Noise for Visual Sound Source Localization Deep residual learning for image recognition

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:31.716044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T14:03:18.151815Z digest=sha256:0ae91c8f1a22a9422749bf4976a6b03510a17b4175e67edadc09aa1042c710ab

Observation bdcaa123-6481-48ef-8d32-f6d0a08e3848 · outbound

This paper cites Momentum contrast for unsupervised visual representation learning.

Learning from Silence and Noise for Visual Sound Source Localization Momentum contrast for unsupervised visual representation learning

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:31.579954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T14:03:18.223211Z digest=sha256:3dba54db850c83582f6cbc66d79f6a9ee17bb4338368b5b2965736dc10d20abd

Observation 1531a6c8-67b1-4cd1-b540-fb97086735e4 · outbound

This paper cites Audio vision: Using audio-visual synchrony to locate sounds.

Learning from Silence and Noise for Visual Sound Source Localization Audio vision: Using audio-visual synchrony to locate sounds

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:31.401708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T14:03:18.309446Z digest=sha256:e4b9973383f60987cd62aa892b736dcdb8790f7994ba6b2a55c9c94ced903333

Observation 2f85e6e0-8b9c-4b19-b9d8-e3c329afbaa8 · outbound

This paper cites Discriminative sounding objects localization via self-supervised audiovisual matching.

Learning from Silence and Noise for Visual Sound Source Localization Discriminative sounding objects localization via self-supervised audiovisual matching

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:31.135238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T14:03:18.380386Z digest=sha256:8d3e4a38405969e360a7b20d829c0221012e8f38f9c37a566c3b0c7e62871fb2

Observation d41ebcae-edae-4131-a152-077352edea82 · outbound

This paper cites Mix and localize: Localizing sound sources in mixtures.

Learning from Silence and Noise for Visual Sound Source Localization Mix and localize: Localizing sound sources in mixtures

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:30.939933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T14:03:18.493766Z digest=sha256:b93afc3d2e7887fcb6eb6247ecd18ed64905dc0439389e8700a9567353097165

Observation f24628e5-4deb-4bd6-a161-2ec939372d7b · outbound

This paper cites You said that?: Synthesising talking faces from audio.

Learning from Silence and Noise for Visual Sound Source Localization You said that?: Synthesising talking faces from audio

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:30.689850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T14:03:18.591004Z digest=sha256:a25acba94d3fd28b0aaeed9f0ead7486a77eed33fdbfd42444e6d4f19398e2b4

Observation fe8c4062-3b87-4099-b58b-ab27916b5b14 · outbound

This paper cites A critical assessment of visual sound source localization models including negative audio.

Learning from Silence and Noise for Visual Sound Source Localization A critical assessment of visual sound source localization models including negative audio

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:30.496216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T14:03:18.664703Z digest=sha256:29485c408c050a0940fb819a73f815d26c70ed5f329235cae044cfa2eb31742f

Observation 3c20823b-4950-455a-9f2d-e625d405f647 · outbound

This paper cites Pixels that sound.

Learning from Silence and Noise for Visual Sound Source Localization Pixels that sound

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:30.279438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T14:03:18.762759Z digest=sha256:24317b62120b72928b8ed422ad3e5d8649947ecf47410e2d393940349ddd43f9

Observation 86a4e8a1-aca1-4a30-a092-b47e19d647af · outbound

This paper cites Learning to visually localize sound sources from mixtures without prior source knowledge.

Learning from Silence and Noise for Visual Sound Source Localization Learning to visually localize sound sources from mixtures without prior source knowledge

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:30.115536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T14:03:18.830545Z digest=sha256:ff53650003e1665c908f864ac15c9fd2f2cd8ee1120adf144ccc4cb441cbafea

Observation f74c8c59-b5b1-4b15-9869-a16ad8230e4a · outbound

This paper cites Adam: A Method for Stochastic Optimization.

Learning from Silence and Noise for Visual Sound Source Localization Adam: A Method for Stochastic Optimization

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-05T14:03:18.902349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:03:18.902349Z digest=sha256:9a96786b73cd8e3caff07180cf6095200b3a9cce13f08cc664c9006d569294aa

Observation 26e5a637-3c5a-4d02-83c3-de343d8cab76 · outbound

This paper cites Cooperative learning of audio and video models from self-supervised synchronization.

Learning from Silence and Noise for Visual Sound Source Localization Cooperative learning of audio and video models from self-supervised synchronization

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:29.932477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T14:03:19.044555Z digest=sha256:7b53315f98acd9719274a8093b094dfdd7b808a7a2a292142cef2a4428b40d36

Observation c3270f41-9b6a-4e61-a63d-3ae7c5b4fc9e · outbound

This paper cites Recent Advances in Multi-modal 3D Intelligence: A Comprehensive Survey and Evaluation.

Learning from Silence and Noise for Visual Sound Source Localization Recent Advances in Multi-modal 3D Intelligence: A Comprehensive Survey and Evaluation

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-08-05T14:03:23.396154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T14:03:19.177165Z digest=sha256:b6bc09e764e0e0fe77da4f4ff2399b494c0b0a3da7701327b04967959d998364

Observation 5ac2cb49-56c7-419a-97c4-702c5577b400 · outbound

This paper cites Do Audio-Visual Segmentation Models Truly Segment Sounding Objects?.

Learning from Silence and Noise for Visual Sound Source Localization Do Audio-Visual Segmentation Models Truly Segment Sounding Objects?

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-05T14:03:19.216337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:03:19.216337Z digest=sha256:89a5f4be0159cb32d486a82351427b43cc2ddb9d315c9bdc0c008f470ae28e6c

Observation 5726eb08-efbe-48d2-8b9a-57ff590efcdd · outbound

This paper cites Av-nerf: Learning neural fields for real-world audio-visual scene synthesis.

Learning from Silence and Noise for Visual Sound Source Localization Av-nerf: Learning neural fields for real-world audio-visual scene synthesis

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:29.749676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T14:03:19.287441Z digest=sha256:fb1930b81451e16e459aa573cf96026be40acab2fb85e67db6e05502c3b1d1f7

Observation 545b8489-112a-4a7c-acb4-3bf5445ac3b9 · outbound

This paper cites Exploiting transformation invariance and equivariance for self-supervised sound localisation.

Learning from Silence and Noise for Visual Sound Source Localization Exploiting transformation invariance and equivariance for self-supervised sound localisation

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:29.444007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T14:03:19.364018Z digest=sha256:0df4ebf31c7232f68dccccbd1851b9cc44db58d3165b07b136c6174979c49363

Observation d6efb336-eb53-4d2b-8eb8-ead267617474 · outbound

This paper cites Visual sound localization in the wild by cross-modal interference erasing.

Learning from Silence and Noise for Visual Sound Source Localization Visual sound localization in the wild by cross-modal interference erasing

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:29.268918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T14:03:19.430163Z digest=sha256:cae03444abf3f90dabbb9ed3a432dde1dfe6c7dc152112a7e5e054911734aaa1

Observation dfadd034-3502-4d21-8367-823a8361e709 · outbound

This paper cites Image segmentation using text and image prompts.

Learning from Silence and Noise for Visual Sound Source Localization Image segmentation using text and image prompts

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:29.062912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T14:03:19.524109Z digest=sha256:34a9d08dd3d290e9e4b580544b009bfcff9a48844ad591776d89ef8bddea95d4

Observation f3b6196e-4d12-4e59-9b43-ef4eb58d925e · outbound

This paper cites T-vsl: Text-guided visual sound source localization in mixtures.

Learning from Silence and Noise for Visual Sound Source Localization T-vsl: Text-guided visual sound source localization in mixtures

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:28.730076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T14:03:19.613935Z digest=sha256:838f8e67ff13dcd1d95629a37249f90037563ba484fcd0fd2e82a350c5bc8a70

Observation 1b55d64d-b16d-409d-bd20-36164e3dab4a · outbound

This paper cites Localizing visual sounds the easy way.

Learning from Silence and Noise for Visual Sound Source Localization Localizing visual sounds the easy way

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:28.457229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T14:03:19.694552Z digest=sha256:3c7476c9c67d5d262b79586219f9bf4b73db7c684145836260f9851acd87c28b

Observation 5198aa3c-400b-4ebf-a71e-8d6cee7cc734 · outbound

This paper cites A closer look at weakly-supervised audio-visual source localization.

Learning from Silence and Noise for Visual Sound Source Localization A closer look at weakly-supervised audio-visual source localization

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:28.241052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T14:03:19.728379Z digest=sha256:42ee7132a4e597ab08a459a7d453307f8b91de383feddc85c45c15b779305797

Observation 29abf2f1-28ad-4e4d-9c21-79903fa57789 · outbound

This paper cites Audio-visual grouping network for sound localization from mixtures.

Learning from Silence and Noise for Visual Sound Source Localization Audio-visual grouping network for sound localization from mixtures

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:28.027267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T14:03:19.819750Z digest=sha256:10b8a10ffbe95439a31bc9a2fd3192736ec1a11e289ffc3e0023fcabb60aeeaa

Observation f6d1524a-ddb1-46a0-8027-bc1a17cdaaee · outbound

This paper cites V ovit: Low latency graph-based audio-visual voice separation transformer.

Learning from Silence and Noise for Visual Sound Source Localization V ovit: Low latency graph-based audio-visual voice separation transformer

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:27.867430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T14:03:19.877524Z digest=sha256:c3759186d3e5becad5cc5f201600c5bcae22beaed376e934a652dab3fa4f7c56

Observation 4212176c-8216-4b62-9e64-87aa0313ae39 · outbound

This paper cites Speech inpainting: Context-based speech synthesis guided by video.

Learning from Silence and Noise for Visual Sound Source Localization Speech inpainting: Context-based speech synthesis guided by video

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:27.640497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T14:03:19.963920Z digest=sha256:bf6a4b56961a496a0acab9326d003c79829b66ff9de3715c572e309e9da1e87e

Observation 35c1b0bd-8031-4bfc-87ec-decfcf16a85b · outbound

This paper cites Representation Learning with Contrastive Predictive Coding.

Learning from Silence and Noise for Visual Sound Source Localization Representation Learning with Contrastive Predictive Coding

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-05T14:03:20.022923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:03:20.022923Z digest=sha256:045b25eb97bfb3689470f367d2b7d5539bd6d065a197a0418df297e1445d57db

Observation ccb7fd8a-a7cb-4ea2-8729-807d1485a669 · outbound

This paper cites Audio-visual scene analysis with self-supervised multisensory features.

Learning from Silence and Noise for Visual Sound Source Localization Audio-visual scene analysis with self-supervised multisensory features

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:27.447874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T14:03:20.134878Z digest=sha256:95c1d2cc214c53f35dc3ab502db9a4edcbdb87885c59a00073ab599e765c58c8

Observation 8bff9481-9821-4878-a193-e6c244a8ef4b · outbound

This paper cites Do we need sound for sound source localization? In Asian Conference on Computer Vision, 2020.

Learning from Silence and Noise for Visual Sound Source Localization Do we need sound for sound source localization? In Asian Conference on Computer Vision, 2020

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:27.224197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T14:03:20.231901Z digest=sha256:02c3fc6b420db3e80630a8c49cd10410a7711ab26db63f6f2189cc10dc156e42

Observation 40558691-05dd-42d0-b7bb-56127c08ea6e · outbound

This paper cites Marginnce: Robust sound localization with a negative margin.

Learning from Silence and Noise for Visual Sound Source Localization Marginnce: Robust sound localization with a negative margin

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:27.021849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T14:03:20.299684Z digest=sha256:528c2e05553a78afcecb600df6e790ee266341aee48e6e5a16cac1291d87fc61

Observation b267725d-7f77-42e1-88e6-f3f085c22a92 · outbound

This paper cites Can clip help sound source localization? In IEEE/CVF Winter Conference on Applications of Computer Vision, pages 5711–5720, 2024.

Learning from Silence and Noise for Visual Sound Source Localization Can clip help sound source localization? In IEEE/CVF Winter Conference on Applications of Computer Vision, pages 5711–5720, 2024

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:26.681627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T14:03:20.368803Z digest=sha256:2030d91c6aa1cecf229028775cd9f3dea82663d93f1702f59a999002cae0d5b0

Observation 41058908-543f-4c88-a706-d013f88c9a81 · outbound

This paper cites Multiple sound sources localization from coarse to fine.

Learning from Silence and Noise for Visual Sound Source Localization Multiple sound sources localization from coarse to fine

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:26.217672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T14:03:20.433210Z digest=sha256:3a91f30affe38c25c79d3a22a7a114c87a894576a776df28ef27f54ac02eea90

Observation b70c6e83-750c-47ff-bb69-fae5b234ca61 · outbound

This paper cites See the sound, hear the pixels.

Learning from Silence and Noise for Visual Sound Source Localization See the sound, hear the pixels

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:26.068140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T14:03:20.496314Z digest=sha256:7f3fc8da8cc0aa5b6b074ef670e7d34992ebfe148b854bb3625d74cb6bc83149

Observation c89f614e-64eb-4afe-9485-56ad4c7a6dc3 · outbound

This paper cites Sound source localization.

Learning from Silence and Noise for Visual Sound Source Localization Sound source localization

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:25.968208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T14:03:20.563495Z digest=sha256:cb5d8885c6b0cf998f258f5184475008bbdb3ec934acf0f459871c87ed941de4

Observation d1015605-cbf7-49cd-b70f-aeb0766ec065 · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

Learning from Silence and Noise for Visual Sound Source Localization High-resolution image synthesis with latent diffusion models

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:25.822302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T14:03:20.665585Z digest=sha256:50b659ca085ffe74802c0b5cc1cfc67405892cb5e984fa4939231bd14b114561

Observation 7bdcae6c-96c0-4e71-94d6-f754b9f90e0f · outbound

This paper cites Multimodal emotion recognition based on a fusion of audiovi- sual information with temporal dynamics.

Learning from Silence and Noise for Visual Sound Source Localization Multimodal emotion recognition based on a fusion of audiovi- sual information with temporal dynamics

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:25.641943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T14:03:20.753635Z digest=sha256:f0caee4f34338b5d4e4671d451fbf638d3b629ef36c6561134d82159a744a7fa

Observation 76cf4767-5cb6-4e57-95b5-ec764ceea4c5 · outbound

This paper cites Learn- ing to localize sound source in visual scenes.

Learning from Silence and Noise for Visual Sound Source Localization Learn- ing to localize sound source in visual scenes

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:25.491772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T14:03:20.867559Z digest=sha256:79ff31100157ac0235bb3924e6beef133cea4e32f49186e87e872837e45921f4

Observation fa0e209e-fb30-4268-a906-a1141e833c0f · outbound

This paper cites Learn- ing to localize sound sources in visual scenes: Analysis and applications.

Learning from Silence and Noise for Visual Sound Source Localization Learn- ing to localize sound sources in visual scenes: Analysis and applications

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:25.322500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T14:03:20.939327Z digest=sha256:b125039c44f991d09337d539ab3aa7562857071f63ca710d3d08658e4954fe42

Observation c0ed37e3-6dda-4b6b-845c-8fd6163ece78 · outbound

This paper cites Learning sound localization better from semantically similar samples.

Learning from Silence and Noise for Visual Sound Source Localization Learning sound localization better from semantically similar samples

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:25.181472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T14:03:21.045335Z digest=sha256:7b26c1a38f13fedd81f0c5622ba0725676f77ce313820f9e7f05b258c6f1084a

Observation c3ee7642-6288-4b1d-a719-e0eea20dadac · outbound

This paper cites Less can be more: Sound source localization with a classification model.

Learning from Silence and Noise for Visual Sound Source Localization Less can be more: Sound source localization with a classification model

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:25.058136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T14:03:21.119429Z digest=sha256:8f172197059d1bc2e91046d7a6bb42c3f9e5e0a0d500f5d41a06b60fc4c54359

Observation a86842d6-bc6f-4693-9f2b-08d3fc350837 · outbound

This paper cites Aligning Sight and Sound: Advanced Sound Source Localization Through Audio-Visual Alignment.

Learning from Silence and Noise for Visual Sound Source Localization Aligning Sight and Sound: Advanced Sound Source Localization Through Audio-Visual Alignment

Reference 58

Resolution
verified exact
local_arxiv, observed 2026-08-05T14:03:23.177348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T14:03:21.201865Z digest=sha256:8206afe6cf430b112b2efb5d0c29c87a29dfe706de008df9763783b57573bb2d

Observation 3c7d422b-5187-4053-8ea8-18ee6b274945 · outbound

This paper cites A Survey on Audio Synthesis and Audio-Visual Multimodal Processing.

Learning from Silence and Noise for Visual Sound Source Localization A Survey on Audio Synthesis and Audio-Visual Multimodal Processing

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-05T14:03:21.327224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:03:21.327224Z digest=sha256:1495de8dcd73f052bf39a36b98a76bcbbca75c66894f881dec2f300bbb79f847

Observation c6f32e40-f962-4dbf-a376-12ab955ba608 · outbound

This paper cites En- hancing sound source localization via false negative elimination.

Learning from Silence and Noise for Visual Sound Source Localization En- hancing sound source localization via false negative elimination

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:24.933449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T14:03:21.398564Z digest=sha256:b6bbbd97c5fe2b91ac2fa988ca21f3a7ec233ec9e797ddcdc66880b5f5e5f4ad

Observation b90020e4-a3a1-4aaf-a769-4fc824b1fef4 · outbound

This paper cites Learning audio-visual source localization via false negative aware contrastive learning.

Learning from Silence and Noise for Visual Sound Source Localization Learning audio-visual source localization via false negative aware contrastive learning

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:24.767706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T14:03:21.503132Z digest=sha256:f8dc3d7430cbe7726385186fc51012a947afdbdb1ef000206c2e60d27948574a

Observation 582a0c3a-1945-42fb-b321-9d4ed159de83 · outbound

This paper cites Sound to visual scene generation by audio-to-visual latent alignment.

Learning from Silence and Noise for Visual Sound Source Localization Sound to visual scene generation by audio-to-visual latent alignment

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:24.635351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T14:03:21.569517Z digest=sha256:815fba0d0736f73122f7c63d6d18c2ca64bf6d9aa60daa49f239524ddc99a22a

Observation fd6bf942-0d75-4daa-8780-f3e02d52b7b9 · outbound

This paper cites Sound2Vision: Generating Diverse Visuals from Audio through Cross-Modal Latent Alignment.

Learning from Silence and Noise for Visual Sound Source Localization Sound2Vision: Generating Diverse Visuals from Audio through Cross-Modal Latent Alignment

Reference 63

Resolution
verified exact
local_arxiv, observed 2026-08-05T14:03:23.027866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T14:03:21.674724Z digest=sha256:a049f0c1da5563ea3f4fa9febc5a15468e40e0936b1faa6e585a1c8909f7c947

Observation d030f199-b42e-4bd6-a4d8-bf6c0aef12c9 · outbound

This paper cites Audio-visual event localization in unconstrained videos.

Learning from Silence and Noise for Visual Sound Source Localization Audio-visual event localization in unconstrained videos

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:24.453977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T14:03:21.770318Z digest=sha256:79987b8e5dab8f37f9e6a3d00fa85a69369a8e21e158afb7e1c37aa2de21c724

Observation ff4b3a39-ebe0-402b-b3e6-63d44945da3b · outbound

This paper cites Phrasecut: Language-based image segmentation in the wild.

Learning from Silence and Noise for Visual Sound Source Localization Phrasecut: Language-based image segmentation in the wild

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:24.279956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T14:03:21.859711Z digest=sha256:80da95e7b763cf813eb6ad88f2d2849866b9ba9421f2b114952c1dd1c337d0f2

Observation 913503cf-93d5-4e6b-a859-7063cfbc141a · outbound

This paper cites How to listen? rethinking visual sound localization.

Learning from Silence and Noise for Visual Sound Source Localization How to listen? rethinking visual sound localization

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:24.077187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T14:03:21.931296Z digest=sha256:7ccba0309cdb46ad130ae42e4e7e06f89bb422cb7a10e302e760e76c4a03680e

Observation 36113071-3274-493b-80ff-fc83234719e8 · outbound

This paper cites Acoustic and visual knowledge distillation for contrastive audio-visual localization.

Learning from Silence and Noise for Visual Sound Source Localization Acoustic and visual knowledge distillation for contrastive audio-visual localization

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:23.900431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T14:03:22.024661Z digest=sha256:0331007fa1098627702ad6b7788f53d4d758e78e2ab7c0e287759ab22d042689

Observation 5472dad6-6e1f-401b-a811-31d910bf3686 · outbound

This paper cites Diagnosing and Rectifying Vision Models using Language.

Learning from Silence and Noise for Visual Sound Source Localization Diagnosing and Rectifying Vision Models using Language

Reference 68

Resolution
verified exact
local_arxiv, observed 2026-08-05T14:03:22.863662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T14:03:22.115862Z digest=sha256:40f7a329a564d79ea83d09c925ff09fc42f7546d86e08093700e9de33a9c167a

Observation 0e808f95-19fd-484e-813b-c4a6e281eee7 · outbound

This paper cites chicken clucking.

Learning from Silence and Noise for Visual Sound Source Localization chicken clucking

Reference 69

Resolution
malformed identifier
raw_fallback, observed 2026-08-05T14:03:22.624081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T14:03:22.242986Z digest=sha256:92b8f3c343ed71584d7681fdbe0e17020347af90f6b4b77ccf70d7e05ddcb6b7

Pith citing papers

No inbound Pith citation observations are available.