Pith. sign in

Paper Citation Record · LEDGER

Learning from Silence and Noise for Visual Sound Source Localization

As of 18 August 2026, this Paper Citation Record lists 69 of 69 outbound references and 0 inbound Pith citation observations for arXiv:2508.21761.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.21761 v1

Coverage vector

measured 69 of 69 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T14:03:22.242986Z

measured 69 of 69 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

69 of 69 outbound references displayed

  • verified exact6
  • verified fuzzy56
  • unresolved6
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3e4a7fab-55d5-457f-86a4-8f3593837185 · outbound

This paper cites Adobe audition sound effects, 2023.

Learning from Silence and Noise for Visual Sound Source Localization Adobe audition sound effects, 2023

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:34.971984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T14:03:15.918896Z digest=sha256:dfc67cc9d4567d909abdf12dee77b226dcf4289d50b676b697a00733ce95e67e

Observation 2487fb81-ca58-4a6e-9bbc-3c103ff6edd9 · outbound

This paper cites Self- supervised learning of audio-visual objects from video.

Learning from Silence and Noise for Visual Sound Source Localization Self- supervised learning of audio-visual objects from video

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:34.794728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T14:03:16.037998Z digest=sha256:c70a625aa126c19a750d4520508553c64fc55c0b55e4fcb37ee6db98a9dfbe21

Observation 167a7fe3-becb-4103-88ac-23a71fd954ee · outbound

This paper cites Look, listen and learn.

Learning from Silence and Noise for Visual Sound Source Localization Look, listen and learn

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:34.604618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T14:03:16.212958Z digest=sha256:5fe4ad467089aa9678c9431c255e086677a4709bdfe747d29572c05558904711

Observation 36d7af9e-35fc-4dd2-a2c6-11d63e7e929c · outbound

This paper cites Objects that sound.

Learning from Silence and Noise for Visual Sound Source Localization Objects that sound

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:34.384136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T14:03:16.317154Z digest=sha256:ccfb310023bfda6c058f0fd71c8f05ac88baab9cbd6ffcbe6c2ec463bb45c96b

Observation 06da2763-a6f0-434b-915e-20afb4a4cf7f · outbound

This paper cites Vggsound: A large-scale audio-visual dataset.

Learning from Silence and Noise for Visual Sound Source Localization Vggsound: A large-scale audio-visual dataset

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:34.185269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T14:03:16.462096Z digest=sha256:48107576d820973c5cdce11f7cec86961c3ad31792caecb39d3a38baa50cf11b

Observation 8f8dccc0-9e7d-48aa-8820-e48257bbecff · outbound

This paper cites Localizing visual sounds the hard way.

Learning from Silence and Noise for Visual Sound Source Localization Localizing visual sounds the hard way

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:33.877814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T14:03:16.662025Z digest=sha256:d35f28234b7c38617781c87d9bac9999792bf5b6e862e13d9c545c29ad731095

Observation ce4379dd-7053-4613-9993-a857258bcb61 · outbound

This paper cites A simple framework for contrastive learning of visual representations.

Learning from Silence and Noise for Visual Sound Source Localization A simple framework for contrastive learning of visual representations

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T14:03:16.836987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:03:16.836987Z digest=sha256:9020c152dadfc50e0824b4d8e7c5b697ceb013c1839d20c4aa3025449a2b1ab5

Observation 9612ced4-e39e-4d8f-98e7-88a8bb288676 · outbound

This paper cites Exploring simple siamese representation learning.

Learning from Silence and Noise for Visual Sound Source Localization Exploring simple siamese representation learning

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:33.654480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T14:03:16.989821Z digest=sha256:832de799228284d242c2d0cf62c5bbb7260484e8ed8afa386f3405e2a6066d61

Observation 18f1a1bf-66a1-46a6-84d2-e8af7158bfb4 · outbound

This paper cites Integrating Audio, Visual, and Semantic Information for Enhanced Multimodal Speaker Diarization.

Learning from Silence and Noise for Visual Sound Source Localization Integrating Audio, Visual, and Semantic Information for Enhanced Multimodal Speaker Diarization

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-08-05T14:03:23.762550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T14:03:17.100791Z digest=sha256:26c51e563839aff2fcc91910ce72cf3bc520df27121fb802cc38e835187b777a

Observation 0a716c15-8ebd-4fd0-88bb-34b2ee5bbf71 · outbound

This paper cites Learning a similarity metric discrimi- natively, with application to face verification.

Learning from Silence and Noise for Visual Sound Source Localization Learning a similarity metric discrimi- natively, with application to face verification

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:33.487670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T14:03:17.193287Z digest=sha256:8e678d6226017b78ec8ff94581d87535ea78d2309ef50bb18d0a90be283cb535

Observation e14f1e04-fc7c-4f20-a678-de813f45fdb2 · outbound

This paper cites Imagenet: A large-scale hierarchical image database.

Learning from Silence and Noise for Visual Sound Source Localization Imagenet: A large-scale hierarchical image database

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:33.265198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T14:03:17.318838Z digest=sha256:ccb65b744897b0387a55b0d732754f8f9388a8e1d14ea089570309669ff7d949

Observation d975b22f-4ca3-4589-a2e4-72ee818513cb · outbound

This paper cites Condi- tional generation of audio from video via foley analogies.

Learning from Silence and Noise for Visual Sound Source Localization Condi- tional generation of audio from video via foley analogies

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:33.105208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T14:03:17.370379Z digest=sha256:f471dfec74946fab6ec74e1e3b57c8c85c497bffa424547741ed7d1ac9343ac3

Observation 289e305d-b66e-4ba4-bce0-e02e77410ff9 · outbound

This paper cites Audio-Visual Approach For Multimodal Concurrent Speaker Detection.

Learning from Silence and Noise for Visual Sound Source Localization Audio-Visual Approach For Multimodal Concurrent Speaker Detection

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-08-05T14:03:23.586547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T14:03:17.449504Z digest=sha256:798da52e03cc9f7426b153ac19736dfff6963fb3d0619f1224e7d0b8551eb5c4

Observation 4ab82633-9016-4ffb-81d5-3db50d0ec276 · outbound

This paper cites Effect of acoustic scene complexity and visual scene representation on auditory perception in virtual audio-visual environments.

Learning from Silence and Noise for Visual Sound Source Localization Effect of acoustic scene complexity and visual scene representation on auditory perception in virtual audio-visual environments

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:32.961414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T14:03:17.494864Z digest=sha256:9217d0b6eb833a4224115707738926d9c2f48655ee70ba34c401f87a503e6eae

Observation 7035228b-f75d-47ad-8bb6-095d71fefd7f · outbound

This paper cites Learning joint sta- tistical models for audio-visual fusion and segregation.

Learning from Silence and Noise for Visual Sound Source Localization Learning joint sta- tistical models for audio-visual fusion and segregation

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:32.769252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T14:03:17.608532Z digest=sha256:4b0034c1be70a155d8f679a40f45b9d1cbcf91caeb94eb48f71d3dc96e8c6456

Observation 21b16037-084b-48f8-9727-8994e6d63e22 · outbound

This paper cites Visualvoice: Audio-visual speech separation with cross-modal consistency.

Learning from Silence and Noise for Visual Sound Source Localization Visualvoice: Audio-visual speech separation with cross-modal consistency

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:32.569106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T14:03:17.717130Z digest=sha256:d0bb000e01f039d931151fb59389f32dfa5843fb9f155820f8f39733acb406b8

Observation 7a8dccc9-a42b-4a45-9b98-5a25c45fe29f · outbound

This paper cites Cyclip: Cyclic contrastive language-image pretraining.

Learning from Silence and Noise for Visual Sound Source Localization Cyclip: Cyclic contrastive language-image pretraining

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:32.373662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T14:03:17.817540Z digest=sha256:c78cfbf675dc58fdc98dcf3f5f701abc7423d7750c76d66f7cbebf8cf3bb5038

Observation 555accc2-bd1e-4443-be62-406557953cf7 · outbound

This paper cites Bootstrap your own latent-a new approach to self-supervised learning.

Learning from Silence and Noise for Visual Sound Source Localization Bootstrap your own latent-a new approach to self-supervised learning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T14:03:17.872247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:03:17.872247Z digest=sha256:b8b9949cf757ec0c51e84b521ca6aef4ed6fb3891ac2a6626db09ce201abe386

Observation 9117824f-d757-446f-a13a-c5ca6b004c8a · outbound

This paper cites chirp" from the.

Learning from Silence and Noise for Visual Sound Source Localization chirp" from the

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:32.139912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T14:03:17.970801Z digest=sha256:fc3a7d576c9c173b2b246472b667fe3571deb71021334e582ab98bfced2d405b

Observation c1f857d6-a920-4f0c-8c65-9a2f577b3ec1 · outbound

This paper cites Canonical correlation analysis: An overview with application to learning methods.

Learning from Silence and Noise for Visual Sound Source Localization Canonical correlation analysis: An overview with application to learning methods

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:31.989916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T14:03:18.046880Z digest=sha256:65ad66aa65361d5b4e14d81af4a0a29c49cd9cbc4556bec0adbabca19e2e2314

Observation f9ac9d16-c864-4f9a-874c-84270023d2d5 · outbound

This paper cites Deep residual learning for image recognition.

Learning from Silence and Noise for Visual Sound Source Localization Deep residual learning for image recognition

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:31.716044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T14:03:18.151815Z digest=sha256:934e1af85898d584c73d4ce8bcca0e647f50d4c52ecb4fae93f9cc886bb4bbde

Observation bdcaa123-6481-48ef-8d32-f6d0a08e3848 · outbound

This paper cites Momentum contrast for unsupervised visual representation learning.

Learning from Silence and Noise for Visual Sound Source Localization Momentum contrast for unsupervised visual representation learning

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:31.579954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T14:03:18.223211Z digest=sha256:513e3098636059c90bfa428301fc4e6d787d153736b8f88ed766d608644b2724

Observation 1531a6c8-67b1-4cd1-b540-fb97086735e4 · outbound

This paper cites Audio vision: Using audio-visual synchrony to locate sounds.

Learning from Silence and Noise for Visual Sound Source Localization Audio vision: Using audio-visual synchrony to locate sounds

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:31.401708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T14:03:18.309446Z digest=sha256:73570db20d7ae1e08fec63a9c035c14a02d151c2560acda57021f0f5d342db84

Observation 2f85e6e0-8b9c-4b19-b9d8-e3c329afbaa8 · outbound

This paper cites Discriminative sounding objects localization via self-supervised audiovisual matching.

Learning from Silence and Noise for Visual Sound Source Localization Discriminative sounding objects localization via self-supervised audiovisual matching

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:31.135238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T14:03:18.380386Z digest=sha256:45aa367710a6eb10e7b2ec6747b6d609fe2bc2135d604e5c4b59adc24b3cb43b

Observation d41ebcae-edae-4131-a152-077352edea82 · outbound

This paper cites Mix and localize: Localizing sound sources in mixtures.

Learning from Silence and Noise for Visual Sound Source Localization Mix and localize: Localizing sound sources in mixtures

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:30.939933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T14:03:18.493766Z digest=sha256:8bff542e2e8b97c09ca53a9e36ab44ce4a0ba98704c23b52830d5b51246b3ae7

Observation f24628e5-4deb-4bd6-a161-2ec939372d7b · outbound

This paper cites You said that?: Synthesising talking faces from audio.

Learning from Silence and Noise for Visual Sound Source Localization You said that?: Synthesising talking faces from audio

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:30.689850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T14:03:18.591004Z digest=sha256:eb8aff1d02f2f140d09bf9c201922207133b1bea40ee0661dd825a5d2ee4274f

Observation fe8c4062-3b87-4099-b58b-ab27916b5b14 · outbound

This paper cites A critical assessment of visual sound source localization models including negative audio.

Learning from Silence and Noise for Visual Sound Source Localization A critical assessment of visual sound source localization models including negative audio

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:30.496216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T14:03:18.664703Z digest=sha256:cd53c682e8d82842fe3b76232732ea46fee436fd8f12d524f20f6492c499884e

Observation 3c20823b-4950-455a-9f2d-e625d405f647 · outbound

This paper cites Pixels that sound.

Learning from Silence and Noise for Visual Sound Source Localization Pixels that sound

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:30.279438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T14:03:18.762759Z digest=sha256:e722d1a3e1a8bd226c4023f984e1ca8b30165fa4ddf73436647dc33d66214066

Observation 86a4e8a1-aca1-4a30-a092-b47e19d647af · outbound

This paper cites Learning to visually localize sound sources from mixtures without prior source knowledge.

Learning from Silence and Noise for Visual Sound Source Localization Learning to visually localize sound sources from mixtures without prior source knowledge

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:30.115536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T14:03:18.830545Z digest=sha256:2f11df9cede6e739038e3fd5392ecc18dd3fce496d44043d15300dcde2d05b70

Observation f74c8c59-b5b1-4b15-9869-a16ad8230e4a · outbound

This paper cites Adam: A Method for Stochastic Optimization.

Learning from Silence and Noise for Visual Sound Source Localization Adam: A Method for Stochastic Optimization

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-05T14:03:18.902349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:03:18.902349Z digest=sha256:bca2488061151e12da14d44fea15108af2c881f21f7a2efe4c6b32a5477a8e74

Observation 26e5a637-3c5a-4d02-83c3-de343d8cab76 · outbound

This paper cites Cooperative learning of audio and video models from self-supervised synchronization.

Learning from Silence and Noise for Visual Sound Source Localization Cooperative learning of audio and video models from self-supervised synchronization

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:29.932477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T14:03:19.044555Z digest=sha256:06c1d1a04e876599d958eef75999a38c3819bf359e14683814ad1c98b9066b9b

Observation c3270f41-9b6a-4e61-a63d-3ae7c5b4fc9e · outbound

This paper cites Recent Advances in Multi-modal 3D Intelligence: A Comprehensive Survey and Evaluation.

Learning from Silence and Noise for Visual Sound Source Localization Recent Advances in Multi-modal 3D Intelligence: A Comprehensive Survey and Evaluation

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-08-05T14:03:23.396154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T14:03:19.177165Z digest=sha256:321124dbcbbaa12adb222bfc4d39520566bf26a3977c353a3e09a7a2d8147d78

Observation 5ac2cb49-56c7-419a-97c4-702c5577b400 · outbound

This paper cites Do Audio-Visual Segmentation Models Truly Segment Sounding Objects?.

Learning from Silence and Noise for Visual Sound Source Localization Do Audio-Visual Segmentation Models Truly Segment Sounding Objects?

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-05T14:03:19.216337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:03:19.216337Z digest=sha256:1a43723cd1c34bed53d659baf737f4bffefffc2eb26b082ec803b910ec6c454c

Observation 5726eb08-efbe-48d2-8b9a-57ff590efcdd · outbound

This paper cites Av-nerf: Learning neural fields for real-world audio-visual scene synthesis.

Learning from Silence and Noise for Visual Sound Source Localization Av-nerf: Learning neural fields for real-world audio-visual scene synthesis

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:29.749676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T14:03:19.287441Z digest=sha256:61e25c9c7e45d7108f7e131742e8acbf57a3c055db9760882ca4a79acf13b324

Observation 545b8489-112a-4a7c-acb4-3bf5445ac3b9 · outbound

This paper cites Exploiting transformation invariance and equivariance for self-supervised sound localisation.

Learning from Silence and Noise for Visual Sound Source Localization Exploiting transformation invariance and equivariance for self-supervised sound localisation

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:29.444007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T14:03:19.364018Z digest=sha256:c64f1780ebcd2676f7f05ad711ee335b6a4e0dee14120995dd8f7bd564c4c487

Observation d6efb336-eb53-4d2b-8eb8-ead267617474 · outbound

This paper cites Visual sound localization in the wild by cross-modal interference erasing.

Learning from Silence and Noise for Visual Sound Source Localization Visual sound localization in the wild by cross-modal interference erasing

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:29.268918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T14:03:19.430163Z digest=sha256:681f148c14fe5ab954f9427458194890db595ae494e53bb33944c333d7f5a284

Observation dfadd034-3502-4d21-8367-823a8361e709 · outbound

This paper cites Image segmentation using text and image prompts.

Learning from Silence and Noise for Visual Sound Source Localization Image segmentation using text and image prompts

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:29.062912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T14:03:19.524109Z digest=sha256:3de189bc582a56644337caeec7bd40dd0a0104d148ebc4471e6c3bf24586b736

Observation f3b6196e-4d12-4e59-9b43-ef4eb58d925e · outbound

This paper cites T-vsl: Text-guided visual sound source localization in mixtures.

Learning from Silence and Noise for Visual Sound Source Localization T-vsl: Text-guided visual sound source localization in mixtures

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:28.730076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T14:03:19.613935Z digest=sha256:6a6e002e91e5a205fb8b8fd1553e874d28b42832c86f7c757fbbe97dc1f420a6

Observation 1b55d64d-b16d-409d-bd20-36164e3dab4a · outbound

This paper cites Localizing visual sounds the easy way.

Learning from Silence and Noise for Visual Sound Source Localization Localizing visual sounds the easy way

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:28.457229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T14:03:19.694552Z digest=sha256:4238cb387ced5594f382a7b67a4b655a5eaef18032cf8fefd4bb770dd1951d8c

Observation 5198aa3c-400b-4ebf-a71e-8d6cee7cc734 · outbound

This paper cites A closer look at weakly-supervised audio-visual source localization.

Learning from Silence and Noise for Visual Sound Source Localization A closer look at weakly-supervised audio-visual source localization

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:28.241052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T14:03:19.728379Z digest=sha256:bb78f147c4048646ec686f60a54ddb367fdc75f9e4e9f999b19ffd2ff2bce448

Observation 29abf2f1-28ad-4e4d-9c21-79903fa57789 · outbound

This paper cites Audio-visual grouping network for sound localization from mixtures.

Learning from Silence and Noise for Visual Sound Source Localization Audio-visual grouping network for sound localization from mixtures

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:28.027267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T14:03:19.819750Z digest=sha256:cb45c27e227a064667dd6f54c3f54711c1adb2ad0810997e78e0cfae64e5f996

Observation f6d1524a-ddb1-46a0-8027-bc1a17cdaaee · outbound

This paper cites V ovit: Low latency graph-based audio-visual voice separation transformer.

Learning from Silence and Noise for Visual Sound Source Localization V ovit: Low latency graph-based audio-visual voice separation transformer

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:27.867430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T14:03:19.877524Z digest=sha256:d489f9b05d82dbaed21e9a872bef5a517391ae217790207b406ae2cd5b002460

Observation 4212176c-8216-4b62-9e64-87aa0313ae39 · outbound

This paper cites Speech inpainting: Context-based speech synthesis guided by video.

Learning from Silence and Noise for Visual Sound Source Localization Speech inpainting: Context-based speech synthesis guided by video

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:27.640497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T14:03:19.963920Z digest=sha256:2fd667577cfc562be281c766f1490fb66a8e1e73c530f24698c588bf7d44508e

Observation 35c1b0bd-8031-4bfc-87ec-decfcf16a85b · outbound

This paper cites Representation Learning with Contrastive Predictive Coding.

Learning from Silence and Noise for Visual Sound Source Localization Representation Learning with Contrastive Predictive Coding

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-05T14:03:20.022923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:03:20.022923Z digest=sha256:113df317d1bf4ab9ecc0dde332156b27bd89c23e07142ceca372a7c2329aaed2

Observation ccb7fd8a-a7cb-4ea2-8729-807d1485a669 · outbound

This paper cites Audio-visual scene analysis with self-supervised multisensory features.

Learning from Silence and Noise for Visual Sound Source Localization Audio-visual scene analysis with self-supervised multisensory features

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:27.447874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T14:03:20.134878Z digest=sha256:8ad838e71b5caeaed8a20dabbdabc21e9ab2204883625dcc03711230138f3010

Observation 8bff9481-9821-4878-a193-e6c244a8ef4b · outbound

This paper cites Do we need sound for sound source localization? In Asian Conference on Computer Vision, 2020.

Learning from Silence and Noise for Visual Sound Source Localization Do we need sound for sound source localization? In Asian Conference on Computer Vision, 2020

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:27.224197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T14:03:20.231901Z digest=sha256:e1151bbcbd342b97fb1d613b87094eedfe431b474075d951f357636f808c8a75

Observation 40558691-05dd-42d0-b7bb-56127c08ea6e · outbound

This paper cites Marginnce: Robust sound localization with a negative margin.

Learning from Silence and Noise for Visual Sound Source Localization Marginnce: Robust sound localization with a negative margin

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:27.021849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T14:03:20.299684Z digest=sha256:0b7c97711db7cd4fecf3f864804489fa649d01deb9ffe839dd5b6339afe4964b

Observation b267725d-7f77-42e1-88e6-f3f085c22a92 · outbound

This paper cites Can clip help sound source localization? In IEEE/CVF Winter Conference on Applications of Computer Vision, pages 5711–5720, 2024.

Learning from Silence and Noise for Visual Sound Source Localization Can clip help sound source localization? In IEEE/CVF Winter Conference on Applications of Computer Vision, pages 5711–5720, 2024

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:26.681627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T14:03:20.368803Z digest=sha256:a59f815097156edce0d4754706f987ac7732aba5a5586ef471387f86039a21d2

Observation 41058908-543f-4c88-a706-d013f88c9a81 · outbound

This paper cites Multiple sound sources localization from coarse to fine.

Learning from Silence and Noise for Visual Sound Source Localization Multiple sound sources localization from coarse to fine

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:26.217672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T14:03:20.433210Z digest=sha256:6cc081596084c0871ac3692c3f4ff1d6cdaa5b051ec56bfd131227b445fad55b

Observation b70c6e83-750c-47ff-bb69-fae5b234ca61 · outbound

This paper cites See the sound, hear the pixels.

Learning from Silence and Noise for Visual Sound Source Localization See the sound, hear the pixels

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:26.068140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T14:03:20.496314Z digest=sha256:29734fe063c1047b1088b16305d55ddb6966df061e148f2c4a552315310c65ef

Observation c89f614e-64eb-4afe-9485-56ad4c7a6dc3 · outbound

This paper cites Sound source localization.

Learning from Silence and Noise for Visual Sound Source Localization Sound source localization

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:25.968208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T14:03:20.563495Z digest=sha256:e06d4779e9b007185217a21f9885f8b7547c7e62815cf415d7a8187b0cdb0fc8

Observation d1015605-cbf7-49cd-b70f-aeb0766ec065 · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

Learning from Silence and Noise for Visual Sound Source Localization High-resolution image synthesis with latent diffusion models

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:25.822302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T14:03:20.665585Z digest=sha256:29ab0064f963265b20973e0d1a84f59438130adf503ad108e4df0786f1d9f881

Observation 7bdcae6c-96c0-4e71-94d6-f754b9f90e0f · outbound

This paper cites Multimodal emotion recognition based on a fusion of audiovi- sual information with temporal dynamics.

Learning from Silence and Noise for Visual Sound Source Localization Multimodal emotion recognition based on a fusion of audiovi- sual information with temporal dynamics

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:25.641943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T14:03:20.753635Z digest=sha256:6e3bf895226d8c50a7f73f1e8a5d202c24ff89e1828c28e9880d2b58407a8c86

Observation 76cf4767-5cb6-4e57-95b5-ec764ceea4c5 · outbound

This paper cites Learn- ing to localize sound source in visual scenes.

Learning from Silence and Noise for Visual Sound Source Localization Learn- ing to localize sound source in visual scenes

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:25.491772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T14:03:20.867559Z digest=sha256:55e0f06804806c141a2dd6d97b2fcdde64f4492d666e53ae9361d73758c88544

Observation fa0e209e-fb30-4268-a906-a1141e833c0f · outbound

This paper cites Learn- ing to localize sound sources in visual scenes: Analysis and applications.

Learning from Silence and Noise for Visual Sound Source Localization Learn- ing to localize sound sources in visual scenes: Analysis and applications

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:25.322500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T14:03:20.939327Z digest=sha256:91fda260e08eaa830b7ca81e7ca2eb438f3dc992ff16f036225ae34934dadb4d

Observation c0ed37e3-6dda-4b6b-845c-8fd6163ece78 · outbound

This paper cites Learning sound localization better from semantically similar samples.

Learning from Silence and Noise for Visual Sound Source Localization Learning sound localization better from semantically similar samples

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:25.181472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T14:03:21.045335Z digest=sha256:87dfde465d8f7758babf9b4f0a5e5a8b4a18b151732062a92220e75beca45cd0

Observation c3ee7642-6288-4b1d-a719-e0eea20dadac · outbound

This paper cites Less can be more: Sound source localization with a classification model.

Learning from Silence and Noise for Visual Sound Source Localization Less can be more: Sound source localization with a classification model

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:25.058136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T14:03:21.119429Z digest=sha256:66419af46038639860eb1b2b95348f230d2c5180ef3ba1f979518f86ac058196

Observation a86842d6-bc6f-4693-9f2b-08d3fc350837 · outbound

This paper cites Aligning Sight and Sound: Advanced Sound Source Localization Through Audio-Visual Alignment.

Learning from Silence and Noise for Visual Sound Source Localization Aligning Sight and Sound: Advanced Sound Source Localization Through Audio-Visual Alignment

Reference 58

Resolution
verified exact
local_arxiv, observed 2026-08-05T14:03:23.177348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T14:03:21.201865Z digest=sha256:f5737156bc398c9ca0278a97c20fbab87bc679381f6fb4f5ed1516f5f4950d11

Observation 3c7d422b-5187-4053-8ea8-18ee6b274945 · outbound

This paper cites A Survey on Audio Synthesis and Audio-Visual Multimodal Processing.

Learning from Silence and Noise for Visual Sound Source Localization A Survey on Audio Synthesis and Audio-Visual Multimodal Processing

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-05T14:03:21.327224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:03:21.327224Z digest=sha256:e9956d23af8e7f1c005af2512ee4eac158cdc48044d81425020624f89030264b

Observation c6f32e40-f962-4dbf-a376-12ab955ba608 · outbound

This paper cites En- hancing sound source localization via false negative elimination.

Learning from Silence and Noise for Visual Sound Source Localization En- hancing sound source localization via false negative elimination

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:24.933449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T14:03:21.398564Z digest=sha256:2bd7498385ceaf07e526fc5444cf5fad004a63ec34fd51acf406d28d5c0f0e9b

Observation b90020e4-a3a1-4aaf-a769-4fc824b1fef4 · outbound

This paper cites Learning audio-visual source localization via false negative aware contrastive learning.

Learning from Silence and Noise for Visual Sound Source Localization Learning audio-visual source localization via false negative aware contrastive learning

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:24.767706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T14:03:21.503132Z digest=sha256:6dd301a1c64057cc03a8c960c60bf364e9c05cf841f7dfecc2a745a43708eac2

Observation 582a0c3a-1945-42fb-b321-9d4ed159de83 · outbound

This paper cites Sound to visual scene generation by audio-to-visual latent alignment.

Learning from Silence and Noise for Visual Sound Source Localization Sound to visual scene generation by audio-to-visual latent alignment

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:24.635351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T14:03:21.569517Z digest=sha256:30ba70a8b3313d44e1fde3e695e187e6806db696c4c63186b63a7aec54569e20

Observation fd6bf942-0d75-4daa-8780-f3e02d52b7b9 · outbound

This paper cites Sound2Vision: Generating Diverse Visuals from Audio through Cross-Modal Latent Alignment.

Learning from Silence and Noise for Visual Sound Source Localization Sound2Vision: Generating Diverse Visuals from Audio through Cross-Modal Latent Alignment

Reference 63

Resolution
verified exact
local_arxiv, observed 2026-08-05T14:03:23.027866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T14:03:21.674724Z digest=sha256:53cc46ec8a9f2cd6f2be7fcb5430d3f1501a8b118464e03ee778a8c7658f33b0

Observation d030f199-b42e-4bd6-a4d8-bf6c0aef12c9 · outbound

This paper cites Audio-visual event localization in unconstrained videos.

Learning from Silence and Noise for Visual Sound Source Localization Audio-visual event localization in unconstrained videos

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:24.453977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T14:03:21.770318Z digest=sha256:3aacec6562c345ae1fdd8155e0586dc6af516f2e323b132c6a8517ff746d37c3

Observation ff4b3a39-ebe0-402b-b3e6-63d44945da3b · outbound

This paper cites Phrasecut: Language-based image segmentation in the wild.

Learning from Silence and Noise for Visual Sound Source Localization Phrasecut: Language-based image segmentation in the wild

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:24.279956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T14:03:21.859711Z digest=sha256:495d0a2c7403b88f41107ef8eef722fdf146ffe016d43ed64c0399bf8f696c04

Observation 913503cf-93d5-4e6b-a859-7063cfbc141a · outbound

This paper cites How to listen? rethinking visual sound localization.

Learning from Silence and Noise for Visual Sound Source Localization How to listen? rethinking visual sound localization

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:24.077187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T14:03:21.931296Z digest=sha256:3cc374c4bc9a4d4f0f3d03df8e5db3d7017f4696f4561849cd08e27335f1b2e2

Observation 36113071-3274-493b-80ff-fc83234719e8 · outbound

This paper cites Acoustic and visual knowledge distillation for contrastive audio-visual localization.

Learning from Silence and Noise for Visual Sound Source Localization Acoustic and visual knowledge distillation for contrastive audio-visual localization

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:03:23.900431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T14:03:22.024661Z digest=sha256:08fac42597eb49a63a32e22a731614a38cae3c60f913f23456dd520b28aa40b0

Observation 5472dad6-6e1f-401b-a811-31d910bf3686 · outbound

This paper cites Diagnosing and Rectifying Vision Models using Language.

Learning from Silence and Noise for Visual Sound Source Localization Diagnosing and Rectifying Vision Models using Language

Reference 68

Resolution
verified exact
local_arxiv, observed 2026-08-05T14:03:22.863662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T14:03:22.115862Z digest=sha256:2a0ad2da7a7608d5ad340ab665f698cdbe063141785f47e2e2193b407f3a80cf

Observation 0e808f95-19fd-484e-813b-c4a6e281eee7 · outbound

This paper cites chicken clucking.

Learning from Silence and Noise for Visual Sound Source Localization chicken clucking

Reference 69

Resolution
malformed identifier
raw_fallback, observed 2026-08-05T14:03:22.624081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T14:03:22.242986Z digest=sha256:c391da359bd1fd6cb0d89fda4bde12c218fdaf7201eac53c4944f92a668d649c

Pith citing papers

No inbound Pith citation observations are available.