Pith. sign in

Paper Citation Record · LEDGER

Object-aware Sound Source Localization via Audio-Visual Scene Understanding

As of 18 August 2026, this Paper Citation Record lists 49 of 49 outbound references and 0 inbound Pith citation observations for arXiv:2506.18557.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.18557 v2

Coverage vector

measured 49 of 49 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T23:20:49.374967Z

measured 49 of 49 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

49 of 49 outbound references displayed

  • verified exact1
  • verified fuzzy40
  • unresolved8
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation cc21691c-a3a3-4237-830a-168b45d996dc · outbound

This paper cites write newline.

Object-aware Sound Source Localization via Audio-Visual Scene Understanding write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:44.623191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:20:44.623191Z digest=sha256:a1f899f80ec7c98aafe10e967ca71b2e254414ddaad4a1e23a22a96c8d49fd14

Observation e009438d-2da3-44a7-937e-2cab930c2b7a · outbound

This paper cites Objects that sound.

Object-aware Sound Source Localization via Audio-Visual Scene Understanding Objects that sound

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:59.016524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T23:20:44.739030Z digest=sha256:00879c6dd03f41efb3e0eda0b9c2ba09afcf5fb4242b57c749da53cd53868867

Observation dcfd5eab-c51b-4ab0-9a80-ba0fed71de2a · outbound

This paper cites Vggsound: A large-scale audio-visual dataset.

Object-aware Sound Source Localization via Audio-Visual Scene Understanding Vggsound: A large-scale audio-visual dataset

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:58.812481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T23:20:44.827505Z digest=sha256:74dd1c38544fb4a4b23e07f17a0beacc64878b3c9e032da35fcca34c140baf81

Observation f152976f-0ba2-4acb-954e-ec303d1b96a0 · outbound

This paper cites Localizing visual sounds the hard way.

Object-aware Sound Source Localization via Audio-Visual Scene Understanding Localizing visual sounds the hard way

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:58.588619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T23:20:44.963367Z digest=sha256:40b17aa5706691f9ec748083c3b5ca5f3282eb0cf32b0779fe516c6f68b9b2a7

Observation dd4af06e-17c5-4d08-8ab4-0589e05ecb75 · outbound

This paper cites Exploring simple siamese representation learning.

Object-aware Sound Source Localization via Audio-Visual Scene Understanding Exploring simple siamese representation learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:45.067294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:20:45.067294Z digest=sha256:45fba23b97cc310cdc8310562813c80f9eb4910c76c07d4f38e156233c4af4cb

Observation 9d7937df-f7d3-4751-b852-a6663328b554 · outbound

This paper cites Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks.

Object-aware Sound Source Localization via Audio-Visual Scene Understanding Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:58.416200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T23:20:45.124745Z digest=sha256:3d040c05373c2d7597dda0f7e5e8d0a47a72a80386ce5c08f282be176c557d49

Observation 71495029-6d41-4882-920a-4fd2ee8da924 · outbound

This paper cites Sinkhorn distances: Lightspeed computation of optimal transport.

Object-aware Sound Source Localization via Audio-Visual Scene Understanding Sinkhorn distances: Lightspeed computation of optimal transport

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:45.241334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:20:45.241334Z digest=sha256:dcb09762ba8dae92721fd24e3a7e750a9e6ff0669ff7330fee77a8b110d8eaeb

Observation cb5fff76-df51-49c0-8092-58bc5216a9f1 · outbound

This paper cites Imagenet: A large-scale hierarchical image database.

Object-aware Sound Source Localization via Audio-Visual Scene Understanding Imagenet: A large-scale hierarchical image database

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:45.352669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:20:45.352669Z digest=sha256:2b5d88f317cfe90317fc2f7d24e15046adcf1cd6f539cd6c9112197f5067a84f

Observation 100fedaa-026b-4114-907f-c8f5cbcad8dd · outbound

This paper cites Cross-modal prompts: Adapting large pre-trained models for audio-visual downstream tasks.

Object-aware Sound Source Localization via Audio-Visual Scene Understanding Cross-modal prompts: Adapting large pre-trained models for audio-visual downstream tasks

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:58.265975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T23:20:45.494963Z digest=sha256:3cdff66b2bf442070cdda5d26bc0c0b5079b5ccdf8d3fd9dbe06420a604e27c3

Observation 12d58a27-c3a6-42b7-879b-6a065254e04c · outbound

This paper cites With a little help from my friends: Nearest-neighbor contrastive learning of visual representations.

Object-aware Sound Source Localization via Audio-Visual Scene Understanding With a little help from my friends: Nearest-neighbor contrastive learning of visual representations

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:58.116336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T23:20:45.594337Z digest=sha256:d5974ea7e173dc18ac08230ef2d6fc1f7fb3b9b5a60530771ca9d96ba5b9fb74

Observation ffcba95a-21df-47a5-8d80-a8e8ed04626e · outbound

This paper cites Hear the flow: Optical flow-based self-supervised visual sound source localization.

Object-aware Sound Source Localization via Audio-Visual Scene Understanding Hear the flow: Optical flow-based self-supervised visual sound source localization

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:57.831456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T23:20:45.678394Z digest=sha256:53d31418cdf294c4316ae89bf750bf352830d9401d055eae63027312c9164669

Observation 686ef128-6d45-4c11-a4da-ae432dc4a78a · outbound

This paper cites Audioclip: Extending clip to image, text and audio.

Object-aware Sound Source Localization via Audio-Visual Scene Understanding Audioclip: Extending clip to image, text and audio

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:57.654611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T23:20:45.810588Z digest=sha256:f6d1b76782d4251dd9ba435ed566cfd6352fcb6b13383fab196272307a70f74f

Observation 314c11cc-f68e-4daa-97a6-0aad424811eb · outbound

This paper cites Deep residual learning for image recognition.

Object-aware Sound Source Localization via Audio-Visual Scene Understanding Deep residual learning for image recognition

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:45.882461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:20:45.882461Z digest=sha256:ca84487ad4f4457c88e14a1c0f90f7688b35e3dd0681ee4e4c541d514f4cda3d

Observation 45025512-b2a6-4da2-b930-10160292978c · outbound

This paper cites Deep multimodal clustering for unsupervised audiovisual learning.

Object-aware Sound Source Localization via Audio-Visual Scene Understanding Deep multimodal clustering for unsupervised audiovisual learning

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:57.394754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T23:20:45.999412Z digest=sha256:65a6538a9440213f266cf52635aa518df96fc97ba383b7833f6bcdd1e7304a74

Observation 64deb004-f7ae-4e9b-abbe-a48c365b3762 · outbound

This paper cites Discriminative sounding objects localization via self-supervised audiovisual matching.

Object-aware Sound Source Localization via Audio-Visual Scene Understanding Discriminative sounding objects localization via self-supervised audiovisual matching

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:57.153525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T23:20:46.075610Z digest=sha256:e3119f5b241c8cc67f0fdb1280e90fbaa458443fe77fd9e48cebc7fc0a288162

Observation cc069196-7209-4cd1-bb72-3dc76df8ebdf · outbound

This paper cites Mix and localize: Localizing sound sources in mixtures.

Object-aware Sound Source Localization via Audio-Visual Scene Understanding Mix and localize: Localizing sound sources in mixtures

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:56.914740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T23:20:46.152099Z digest=sha256:c13b1d6a14bb4397bd1d2e2027d015da1c0ab6ad6318722f6d2b0e75cc05ae27

Observation 9faa0021-b426-4b9a-85af-cd5d45bafa24 · outbound

This paper cites Boosting contrastive self-supervised learning with false negative cancellation.

Object-aware Sound Source Localization via Audio-Visual Scene Understanding Boosting contrastive self-supervised learning with false negative cancellation

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:56.674748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T23:20:46.260804Z digest=sha256:6265253316fd1bcdf7bebd06ada8d009b45071fa7eae64f860e98168d52cdcb0

Observation f11684f8-189a-4455-bcfd-dfee9d4c5ca1 · outbound

This paper cites A review of recent advances on deep learning methods for audio-visual speech recognition.

Object-aware Sound Source Localization via Audio-Visual Scene Understanding A review of recent advances on deep learning methods for audio-visual speech recognition

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:56.385786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T23:20:46.326449Z digest=sha256:f93cab2a5337ed43e92d01ee5a25fb1fe1fe97286f08df8b642c88d913b5e126

Observation ac83f528-b373-4ce5-94d9-41338915fbbb · outbound

This paper cites Learning to visually localize sound sources from mixtures without prior source knowledge.

Object-aware Sound Source Localization via Audio-Visual Scene Understanding Learning to visually localize sound sources from mixtures without prior source knowledge

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:56.134749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T23:20:46.404760Z digest=sha256:f6e35afe80b8f42a86d66dbc77d83fd48d952cedc989af237e5ca3c1fc62c94e

Observation 085df064-f62c-4d87-8d26-60ae9757c918 · outbound

This paper cites Adam: A Method for Stochastic Optimization.

Object-aware Sound Source Localization via Audio-Visual Scene Understanding Adam: A Method for Stochastic Optimization

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:46.466395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:20:46.466395Z digest=sha256:c707e392765cdc666b673cdd927a232fd599b546a5014852ffd2ee41c0459769

Observation 572067d7-d2fc-498e-968e-d1122aeb18d6 · outbound

This paper cites Unsupervised sound localization via iterative contrastive learning.

Object-aware Sound Source Localization via Audio-Visual Scene Understanding Unsupervised sound localization via iterative contrastive learning

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:55.942335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T23:20:46.543786Z digest=sha256:790c0e2a845f270b6496e67806a2a6f2bbc8a2e4be32c45a8b1a9da4f78e2667

Observation e9005eba-beb0-4a48-98b5-bd5048d15ffb · outbound

This paper cites Exploiting transformation invariance and equivariance for self-supervised sound localisation.

Object-aware Sound Source Localization via Audio-Visual Scene Understanding Exploiting transformation invariance and equivariance for self-supervised sound localisation

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:55.749982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T23:20:46.716945Z digest=sha256:49d64e6ae390a67226daafb759fbb8ab18a6dee4293d91d5a26677a655b20405

Observation 5ee4ac45-00ea-492c-aee1-293e2cc937d9 · outbound

This paper cites Generalized video anomaly event detection: Systematic taxonomy and comparison of deep models.

Object-aware Sound Source Localization via Audio-Visual Scene Understanding Generalized video anomaly event detection: Systematic taxonomy and comparison of deep models

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:55.434742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T23:20:46.825427Z digest=sha256:f5db1f25f2ebf5f6c9d58d7f1d32124a01aa6a4a7e0a69e95f876cd7a463f8a8

Observation dab2a57e-d48f-462e-ae71-3874f91a96fa · outbound

This paper cites T-vsl: Text-guided visual sound source localization in mixtures.

Object-aware Sound Source Localization via Audio-Visual Scene Understanding T-vsl: Text-guided visual sound source localization in mixtures

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:55.207310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T23:20:47.179413Z digest=sha256:defe74a5a95ff7482cb9027dceefd045b9d742300c92196707f6e01d718884ba

Observation d6076cb8-c9fa-45a4-a79f-0e9f64d5713c · outbound

This paper cites Localizing visual sounds the easy way.

Object-aware Sound Source Localization via Audio-Visual Scene Understanding Localizing visual sounds the easy way

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:54.989857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T23:20:47.319708Z digest=sha256:f61dbf1f896b30aec272b017c7aa9e4b910f6547c0fb44b4537a4316c56ada1d

Observation 510a2c9a-c178-450c-9083-5d6387ce1c88 · outbound

This paper cites A closer look at weakly-supervised audio-visual source localization.

Object-aware Sound Source Localization via Audio-Visual Scene Understanding A closer look at weakly-supervised audio-visual source localization

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:54.793986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T23:20:47.399677Z digest=sha256:25d1ade8ea64f1fd64e02940b8d9877c038aa26c9f453a3a2f74c7ed3d31e79b

Observation 7ead8897-a468-4d84-9723-bbdd899358ea · outbound

This paper cites Audio-visual grouping network for sound localization from mixtures.

Object-aware Sound Source Localization via Audio-Visual Scene Understanding Audio-visual grouping network for sound localization from mixtures

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:54.602287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T23:20:47.445298Z digest=sha256:fc2211d0b02ef1274e1909f8d3428d5d4d3eefb65bb72c84ff19a389a9d3e62f

Observation a7f2f281-70a1-44d5-a856-c0cdf82229f3 · outbound

This paper cites Audio-visual scene analysis with self-supervised multisensory features.

Object-aware Sound Source Localization via Audio-Visual Scene Understanding Audio-visual scene analysis with self-supervised multisensory features

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:54.380654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T23:20:47.474313Z digest=sha256:1a638080adbad84b106235390cf21ec1709dfb686e72d5a62fecb1be7aaf9551

Observation 78868aa1-7c8c-4746-98c3-374d3e091bc9 · outbound

This paper cites Multiple sound sources localization from coarse to fine.

Object-aware Sound Source Localization via Audio-Visual Scene Understanding Multiple sound sources localization from coarse to fine

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:54.184834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T23:20:47.545828Z digest=sha256:f722c46491637464d086952ebed99c338875d82e75a22217cbb512d880a6f299

Observation a8698035-bfd8-4a4e-993b-51396e7e26a4 · outbound

This paper cites Multimodal Open-Vocabulary Video Classification via Pre-Trained Vision and Language Models.

Object-aware Sound Source Localization via Audio-Visual Scene Understanding Multimodal Open-Vocabulary Video Classification via Pre-Trained Vision and Language Models

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-08-06T23:20:49.591703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T23:20:47.607798Z digest=sha256:d074760890f9934c2b4e79bcf81b85fa321128f935b847ee1a671ace0fd36c96

Observation 9b0ae4b5-47ac-48ad-9615-77b480b67d4a · outbound

This paper cites Mm-diffusion: Learning multi-modal diffusion models for joint audio and video generation.

Object-aware Sound Source Localization via Audio-Visual Scene Understanding Mm-diffusion: Learning multi-modal diffusion models for joint audio and video generation

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:53.964747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T23:20:47.662872Z digest=sha256:de46010444d8376eb391e4702a84f6499fb6898161a41ca2f4f605b5af6f9825

Observation d51aa6db-5f85-46d9-ad3c-9af470c889bc · outbound

This paper cites Learning to localize sound source in visual scenes.

Object-aware Sound Source Localization via Audio-Visual Scene Understanding Learning to localize sound source in visual scenes

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:53.797386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T23:20:47.742209Z digest=sha256:940c0f4b4bbf91d651bcbd3b45ccfe50fc5f270562ffd38e428eb6aa4ab80183

Observation fc4440a1-9511-4a3e-a515-fb31081f8a81 · outbound

This paper cites Learning sound localization better from semantically similar samples.

Object-aware Sound Source Localization via Audio-Visual Scene Understanding Learning sound localization better from semantically similar samples

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:53.536440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T23:20:47.780015Z digest=sha256:76c289176d60da77f6ff183e30b8c3d9d8a6f94f575023540edef398f76cef08

Observation c3c94bfc-8ce8-4f7c-99ff-cfaa9bf03581 · outbound

This paper cites Sound source localization is all about cross-modal alignment.

Object-aware Sound Source Localization via Audio-Visual Scene Understanding Sound source localization is all about cross-modal alignment

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:53.104744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T23:20:47.952858Z digest=sha256:a458231658689c43a7801d9c50b484ea5682e7fd8a002e3552e33c17124e0f8b

Observation edc0cf83-fb89-4bc0-8973-1cce97dd04a5 · outbound

This paper cites Unsupervised sounding object localization with bottom-up and top-down attention.

Object-aware Sound Source Localization via Audio-Visual Scene Understanding Unsupervised sounding object localization with bottom-up and top-down attention

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:52.876261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T23:20:47.985284Z digest=sha256:ca4448c96d1ccfa02cab9849f2e83dbb8ca93f5f960cd65798179299df8d9ed1

Observation e821573b-01b8-42d4-b2ee-4af3b572fda0 · outbound

This paper cites Flowgrad: Using motion for visual sound source localization.

Object-aware Sound Source Localization via Audio-Visual Scene Understanding Flowgrad: Using motion for visual sound source localization

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:52.534945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T23:20:48.061595Z digest=sha256:85f20835d2550caa57bc9c95d7e4da3c4abdde662afae39d137c3f077c8826d4

Observation 9b7a432f-dd00-4e2e-880b-98c6df771e1f · outbound

This paper cites Self-supervised predictive learning: A negative-free method for sound source localization in visual scenes.

Object-aware Sound Source Localization via Audio-Visual Scene Understanding Self-supervised predictive learning: A negative-free method for sound source localization in visual scenes

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:52.215295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T23:20:48.196924Z digest=sha256:8dc5d24456389620af381c072b7889d4870263bdf19d082dc2cd9d46536876d8

Observation 0099b1eb-197c-4982-bbdd-470db32f7dbe · outbound

This paper cites Learning audio-visual source localization via false negative aware contrastive learning.

Object-aware Sound Source Localization via Audio-Visual Scene Understanding Learning audio-visual source localization via false negative aware contrastive learning

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:51.998121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T23:20:48.316561Z digest=sha256:c64236b15f79c22b44a5af2f33674130066376e10f562f1dc83a7358328f10ff

Observation 04793bae-3c14-47f6-b257-7b79e6ed984b · outbound

This paper cites Audio-visual spatial integration and recursive attention for robust sound source localization.

Object-aware Sound Source Localization via Audio-Visual Scene Understanding Audio-visual spatial integration and recursive attention for robust sound source localization

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:51.739907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T23:20:48.413073Z digest=sha256:792f24bf1971da24b6f85f49fd96881def9cb4f9dfcb66aa6196ac9c16cee9e1

Observation a4793b17-fe85-4b43-8de0-53d48099bcae · outbound

This paper cites Watch video, catch keyword: Context-aware keyword attention for moment retrieval and highlight detection.

Object-aware Sound Source Localization via Audio-Visual Scene Understanding Watch video, catch keyword: Context-aware keyword attention for moment retrieval and highlight detection

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:51.514758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T23:20:48.481956Z digest=sha256:057495b5b05c2caac97b7fb2d4d78120acf167b4f163c75bcfcd7807108c1050

Observation de45388e-df01-4cc3-aafa-63a089495525 · outbound

This paper cites V2a-mapper: A lightweight solution for vision-to-audio generation by connecting foundation models.

Object-aware Sound Source Localization via Audio-Visual Scene Understanding V2a-mapper: A lightweight solution for vision-to-audio generation by connecting foundation models

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:51.283695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T23:20:48.538040Z digest=sha256:8480f40328832b9dfa114a5f2a003cb877ec0baf0e0597a05519961812b28e31

Observation 02f9acb8-8e93-4110-856a-999f7b07683d · outbound

This paper cites Multimodal large language models: A survey.

Object-aware Sound Source Localization via Audio-Visual Scene Understanding Multimodal large language models: A survey

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:51.046232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T23:20:48.669587Z digest=sha256:067998ac01200d39d5e9f0915b71496308e1f722d0a19b27255c288277dcdcf7

Observation caf9bdfc-63e5-467e-a050-368ba2d6297b · outbound

This paper cites Sonicvisionlm: Playing sound with vision language models.

Object-aware Sound Source Localization via Audio-Visual Scene Understanding Sonicvisionlm: Playing sound with vision language models

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:50.814775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T23:20:48.722137Z digest=sha256:5b98cd1d967dc9ca9ada893934b70cc493d02e0e48dfd25cc1d3e5fce54e26ed

Observation e09ecbac-9ff0-495b-9354-2b1ed2b00458 · outbound

This paper cites A proposal-based paradigm for self-supervised sound source localization in videos.

Object-aware Sound Source Localization via Audio-Visual Scene Understanding A proposal-based paradigm for self-supervised sound source localization in videos

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:50.588081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T23:20:48.915324Z digest=sha256:d66136ac496171a0cd8d8685de4a808cbc9f63f66f660aa34626979da5949f8a

Observation 0ae44f15-eda2-4597-b370-81865df6e06d · outbound

This paper cites A Survey on Multimodal Large Language Models.

Object-aware Sound Source Localization via Audio-Visual Scene Understanding A Survey on Multimodal Large Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:48.985249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:20:48.985249Z digest=sha256:2002a5aaff38ac594a53142d4162aecc707cebdb4f745beb7c81e712fd36a190

Observation 52a526d2-17b8-4015-8ead-64026d200fcc · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

Object-aware Sound Source Localization via Audio-Visual Scene Understanding Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:49.075435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:20:49.075435Z digest=sha256:ef3157ef3f0a016ef005a84b3b3a6c6a5b1e899bc4a5a5251efc7c18f2f0cd62

Observation 7d765c18-5d41-445c-a565-addc16c9924e · outbound

This paper cites The sound of pixels.

Object-aware Sound Source Localization via Audio-Visual Scene Understanding The sound of pixels

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:50.348537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T23:20:49.160956Z digest=sha256:dea6149cb35cfe3fb6ac59c2fd363b99aed867ec2ace73ca62263d03bd0fdd80

Observation 37e29985-6257-4e8f-af2e-e41c0dfa0027 · outbound

This paper cites Weakly supervised contrastive learning.

Object-aware Sound Source Localization via Audio-Visual Scene Understanding Weakly supervised contrastive learning

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:50.052232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T23:20:49.253720Z digest=sha256:08cd22f87eacd64f592893ce2b427f0bb508c5dd9db01fc635de0ef33c5e24bc

Observation 69480215-0ca9-4260-b091-d46be833e310 · outbound

This paper cites Exploiting visual context semantics for sound source localization.

Object-aware Sound Source Localization via Audio-Visual Scene Understanding Exploiting visual context semantics for sound source localization

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:49.795672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T23:20:49.374967Z digest=sha256:44bc60c87f677d479f19df1d178620ac3f691f7305dbd3e0757e647ab2c4d471

Pith citing papers

No inbound Pith citation observations are available.