Pith. sign in

Paper Citation Record · LEDGER

Object-aware Sound Source Localization via Audio-Visual Scene Understanding

As of 14 August 2026, this Paper Citation Record lists 49 of 49 outbound references and 0 inbound Pith citation observations for arXiv:2506.18557.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.18557 v2

Coverage vector

measured 49 of 49 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T23:20:49.374967Z

measured 49 of 49 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

49 of 49 outbound references displayed

  • verified exact1
  • verified fuzzy40
  • unresolved8
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation cc21691c-a3a3-4237-830a-168b45d996dc · outbound

This paper cites write newline.

Object-aware Sound Source Localization via Audio-Visual Scene Understanding write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:44.623191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:20:44.623191Z digest=sha256:a1f899f80ec7c98aafe10e967ca71b2e254414ddaad4a1e23a22a96c8d49fd14

Observation e009438d-2da3-44a7-937e-2cab930c2b7a · outbound

This paper cites Objects that sound.

Object-aware Sound Source Localization via Audio-Visual Scene Understanding Objects that sound

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:59.016524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-06T23:20:44.739030Z digest=sha256:72150943b4a0f4bf0d17159a8f8e2a3ad61865c4d8b860ce129debf97509ea3d

Observation dcfd5eab-c51b-4ab0-9a80-ba0fed71de2a · outbound

This paper cites Vggsound: A large-scale audio-visual dataset.

Object-aware Sound Source Localization via Audio-Visual Scene Understanding Vggsound: A large-scale audio-visual dataset

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:58.812481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-06T23:20:44.827505Z digest=sha256:c24885dd564ded0ea7bfc72dcab8c1139cd762ac6ade1880e233326aaf695364

Observation f152976f-0ba2-4acb-954e-ec303d1b96a0 · outbound

This paper cites Localizing visual sounds the hard way.

Object-aware Sound Source Localization via Audio-Visual Scene Understanding Localizing visual sounds the hard way

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:58.588619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-06T23:20:44.963367Z digest=sha256:78716b57aca69ac350c0608894bbd82a3164401a346c4b1a628c67cd49a7d1a5

Observation dd4af06e-17c5-4d08-8ab4-0589e05ecb75 · outbound

This paper cites Exploring simple siamese representation learning.

Object-aware Sound Source Localization via Audio-Visual Scene Understanding Exploring simple siamese representation learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:45.067294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:20:45.067294Z digest=sha256:45fba23b97cc310cdc8310562813c80f9eb4910c76c07d4f38e156233c4af4cb

Observation 9d7937df-f7d3-4751-b852-a6663328b554 · outbound

This paper cites Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks.

Object-aware Sound Source Localization via Audio-Visual Scene Understanding Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:58.416200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-06T23:20:45.124745Z digest=sha256:17fd348a5fd83ad4ddbff4890cb4e13a00f2c7d9ec6fa8a95de2a247cc09d5a1

Observation 71495029-6d41-4882-920a-4fd2ee8da924 · outbound

This paper cites Sinkhorn distances: Lightspeed computation of optimal transport.

Object-aware Sound Source Localization via Audio-Visual Scene Understanding Sinkhorn distances: Lightspeed computation of optimal transport

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:45.241334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:20:45.241334Z digest=sha256:dcb09762ba8dae92721fd24e3a7e750a9e6ff0669ff7330fee77a8b110d8eaeb

Observation cb5fff76-df51-49c0-8092-58bc5216a9f1 · outbound

This paper cites Imagenet: A large-scale hierarchical image database.

Object-aware Sound Source Localization via Audio-Visual Scene Understanding Imagenet: A large-scale hierarchical image database

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:45.352669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:20:45.352669Z digest=sha256:2b5d88f317cfe90317fc2f7d24e15046adcf1cd6f539cd6c9112197f5067a84f

Observation 100fedaa-026b-4114-907f-c8f5cbcad8dd · outbound

This paper cites Cross-modal prompts: Adapting large pre-trained models for audio-visual downstream tasks.

Object-aware Sound Source Localization via Audio-Visual Scene Understanding Cross-modal prompts: Adapting large pre-trained models for audio-visual downstream tasks

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:58.265975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-06T23:20:45.494963Z digest=sha256:4c92a3c951f1a783b78e3e46eec31beb1bae741029d4620d13f9663edd725ad8

Observation 12d58a27-c3a6-42b7-879b-6a065254e04c · outbound

This paper cites With a little help from my friends: Nearest-neighbor contrastive learning of visual representations.

Object-aware Sound Source Localization via Audio-Visual Scene Understanding With a little help from my friends: Nearest-neighbor contrastive learning of visual representations

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:58.116336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-06T23:20:45.594337Z digest=sha256:36e1269124acdada63d96563ba2147971588be22e650dd27539b1c7bda01fc14

Observation ffcba95a-21df-47a5-8d80-a8e8ed04626e · outbound

This paper cites Hear the flow: Optical flow-based self-supervised visual sound source localization.

Object-aware Sound Source Localization via Audio-Visual Scene Understanding Hear the flow: Optical flow-based self-supervised visual sound source localization

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:57.831456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-06T23:20:45.678394Z digest=sha256:28be81b52a2e93eee5e9011f26e21d8dd4df1b6d7935c3b477c761b67597ad1b

Observation 686ef128-6d45-4c11-a4da-ae432dc4a78a · outbound

This paper cites Audioclip: Extending clip to image, text and audio.

Object-aware Sound Source Localization via Audio-Visual Scene Understanding Audioclip: Extending clip to image, text and audio

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:57.654611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-06T23:20:45.810588Z digest=sha256:02916fa172c499eb94e1c4bbf82cdcec4920fbd80b91ce9a74bdea2fac910089

Observation 314c11cc-f68e-4daa-97a6-0aad424811eb · outbound

This paper cites Deep residual learning for image recognition.

Object-aware Sound Source Localization via Audio-Visual Scene Understanding Deep residual learning for image recognition

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:45.882461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:20:45.882461Z digest=sha256:ca84487ad4f4457c88e14a1c0f90f7688b35e3dd0681ee4e4c541d514f4cda3d

Observation 45025512-b2a6-4da2-b930-10160292978c · outbound

This paper cites Deep multimodal clustering for unsupervised audiovisual learning.

Object-aware Sound Source Localization via Audio-Visual Scene Understanding Deep multimodal clustering for unsupervised audiovisual learning

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:57.394754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-06T23:20:45.999412Z digest=sha256:4dbd338ced85b0b0485ec87c42b16b2fa767097a319f7658faff87ff82aae728

Observation 64deb004-f7ae-4e9b-abbe-a48c365b3762 · outbound

This paper cites Discriminative sounding objects localization via self-supervised audiovisual matching.

Object-aware Sound Source Localization via Audio-Visual Scene Understanding Discriminative sounding objects localization via self-supervised audiovisual matching

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:57.153525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-06T23:20:46.075610Z digest=sha256:98c8aa0c425e744554d4b326a001f745cec2388a969b6a36854335a8a1bd0d3a

Observation cc069196-7209-4cd1-bb72-3dc76df8ebdf · outbound

This paper cites Mix and localize: Localizing sound sources in mixtures.

Object-aware Sound Source Localization via Audio-Visual Scene Understanding Mix and localize: Localizing sound sources in mixtures

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:56.914740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-06T23:20:46.152099Z digest=sha256:8564e939712da8262e97ce28e79b199981889d0b0f049ac40180f06c16e6bea0

Observation 9faa0021-b426-4b9a-85af-cd5d45bafa24 · outbound

This paper cites Boosting contrastive self-supervised learning with false negative cancellation.

Object-aware Sound Source Localization via Audio-Visual Scene Understanding Boosting contrastive self-supervised learning with false negative cancellation

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:56.674748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-06T23:20:46.260804Z digest=sha256:f2e9e846dc98428a5c48819a2534c0ec3fd274a30f6e13c9ce274f79627a601e

Observation f11684f8-189a-4455-bcfd-dfee9d4c5ca1 · outbound

This paper cites A review of recent advances on deep learning methods for audio-visual speech recognition.

Object-aware Sound Source Localization via Audio-Visual Scene Understanding A review of recent advances on deep learning methods for audio-visual speech recognition

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:56.385786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-06T23:20:46.326449Z digest=sha256:026ec62f223de2ead3fb82ab8d201aa1f7820783de3037767c4d265135f9bfc2

Observation ac83f528-b373-4ce5-94d9-41338915fbbb · outbound

This paper cites Learning to visually localize sound sources from mixtures without prior source knowledge.

Object-aware Sound Source Localization via Audio-Visual Scene Understanding Learning to visually localize sound sources from mixtures without prior source knowledge

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:56.134749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-06T23:20:46.404760Z digest=sha256:90932d9f8355d3711638240ddda9324516c5ac4227555170160e49ba8c39c92b

Observation 085df064-f62c-4d87-8d26-60ae9757c918 · outbound

This paper cites Adam: A Method for Stochastic Optimization.

Object-aware Sound Source Localization via Audio-Visual Scene Understanding Adam: A Method for Stochastic Optimization

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:46.466395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:20:46.466395Z digest=sha256:0de9a6bfab9b9b7e1eaef91907af99e328428d253962c820b0809f822d34af2e

Observation 572067d7-d2fc-498e-968e-d1122aeb18d6 · outbound

This paper cites Unsupervised sound localization via iterative contrastive learning.

Object-aware Sound Source Localization via Audio-Visual Scene Understanding Unsupervised sound localization via iterative contrastive learning

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:55.942335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-06T23:20:46.543786Z digest=sha256:32ff92eb70739360ed73c2c0bffc85773c8b88fc7b1d3edb2cd9137b1d18e819

Observation e9005eba-beb0-4a48-98b5-bd5048d15ffb · outbound

This paper cites Exploiting transformation invariance and equivariance for self-supervised sound localisation.

Object-aware Sound Source Localization via Audio-Visual Scene Understanding Exploiting transformation invariance and equivariance for self-supervised sound localisation

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:55.749982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-06T23:20:46.716945Z digest=sha256:7f170f3ed18aef0544b7848c366319b684472dd73473c1edd1ac66b25a743368

Observation 5ee4ac45-00ea-492c-aee1-293e2cc937d9 · outbound

This paper cites Generalized video anomaly event detection: Systematic taxonomy and comparison of deep models.

Object-aware Sound Source Localization via Audio-Visual Scene Understanding Generalized video anomaly event detection: Systematic taxonomy and comparison of deep models

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:55.434742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-06T23:20:46.825427Z digest=sha256:9032c1fa1fbc21c339320111cca36d30dc9479a9b2afc4b453d4f0d56bdaa87a

Observation dab2a57e-d48f-462e-ae71-3874f91a96fa · outbound

This paper cites T-vsl: Text-guided visual sound source localization in mixtures.

Object-aware Sound Source Localization via Audio-Visual Scene Understanding T-vsl: Text-guided visual sound source localization in mixtures

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:55.207310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-06T23:20:47.179413Z digest=sha256:487f66bbc1ddb2c1bd36a0ac78e56bb3ec7800904c2e033da3d5fab7df3359d7

Observation d6076cb8-c9fa-45a4-a79f-0e9f64d5713c · outbound

This paper cites Localizing visual sounds the easy way.

Object-aware Sound Source Localization via Audio-Visual Scene Understanding Localizing visual sounds the easy way

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:54.989857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-06T23:20:47.319708Z digest=sha256:dc2dc97da96285d3c3cdf26fe5179b747c1722527dbcb390c2beba59e7ded8b5

Observation 510a2c9a-c178-450c-9083-5d6387ce1c88 · outbound

This paper cites A closer look at weakly-supervised audio-visual source localization.

Object-aware Sound Source Localization via Audio-Visual Scene Understanding A closer look at weakly-supervised audio-visual source localization

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:54.793986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-06T23:20:47.399677Z digest=sha256:591fb8b0fe4a9bf133c04589a420f214439d15f00fd1f87ae331548ee94e2dbb

Observation 7ead8897-a468-4d84-9723-bbdd899358ea · outbound

This paper cites Audio-visual grouping network for sound localization from mixtures.

Object-aware Sound Source Localization via Audio-Visual Scene Understanding Audio-visual grouping network for sound localization from mixtures

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:54.602287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-06T23:20:47.445298Z digest=sha256:812620173ff665c8043fbabad078ce78579d1fd68f83a545d33b6dd88e85d32c

Observation a7f2f281-70a1-44d5-a856-c0cdf82229f3 · outbound

This paper cites Audio-visual scene analysis with self-supervised multisensory features.

Object-aware Sound Source Localization via Audio-Visual Scene Understanding Audio-visual scene analysis with self-supervised multisensory features

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:54.380654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-06T23:20:47.474313Z digest=sha256:b19af7c3defe66f80a69db751d4e322f198e5fac1223d86b56af8016b5108c91

Observation 78868aa1-7c8c-4746-98c3-374d3e091bc9 · outbound

This paper cites Multiple sound sources localization from coarse to fine.

Object-aware Sound Source Localization via Audio-Visual Scene Understanding Multiple sound sources localization from coarse to fine

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:54.184834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-06T23:20:47.545828Z digest=sha256:2137413de19b1d5b4a44cdbb1e54a60bebf2d450b88e9e45f5193a57bc1367de

Observation a8698035-bfd8-4a4e-993b-51396e7e26a4 · outbound

This paper cites Multimodal Open-Vocabulary Video Classification via Pre-Trained Vision and Language Models.

Object-aware Sound Source Localization via Audio-Visual Scene Understanding Multimodal Open-Vocabulary Video Classification via Pre-Trained Vision and Language Models

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-08-06T23:20:49.591703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-06T23:20:47.607798Z digest=sha256:413d1441c0f22f3de7495c7ee47f85afe54661d24f5782a2a1d052fa6f3f328b

Observation 9b0ae4b5-47ac-48ad-9615-77b480b67d4a · outbound

This paper cites Mm-diffusion: Learning multi-modal diffusion models for joint audio and video generation.

Object-aware Sound Source Localization via Audio-Visual Scene Understanding Mm-diffusion: Learning multi-modal diffusion models for joint audio and video generation

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:53.964747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-06T23:20:47.662872Z digest=sha256:d121264c330cd98ef9c7a29771fa1604c67bebb22c633125237a4ded340b4d6c

Observation d51aa6db-5f85-46d9-ad3c-9af470c889bc · outbound

This paper cites Learning to localize sound source in visual scenes.

Object-aware Sound Source Localization via Audio-Visual Scene Understanding Learning to localize sound source in visual scenes

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:53.797386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-06T23:20:47.742209Z digest=sha256:658b5893bf7de44adb7ed7bf0481525493d65b4df3159302477211903145a074

Observation fc4440a1-9511-4a3e-a515-fb31081f8a81 · outbound

This paper cites Learning sound localization better from semantically similar samples.

Object-aware Sound Source Localization via Audio-Visual Scene Understanding Learning sound localization better from semantically similar samples

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:53.536440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-06T23:20:47.780015Z digest=sha256:9f6a604b03fc09a8202a860571e8847b7df8c99c595fbb1a64bd7d9f448a6798

Observation c3c94bfc-8ce8-4f7c-99ff-cfaa9bf03581 · outbound

This paper cites Sound source localization is all about cross-modal alignment.

Object-aware Sound Source Localization via Audio-Visual Scene Understanding Sound source localization is all about cross-modal alignment

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:53.104744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-06T23:20:47.952858Z digest=sha256:28e49b87c4afd09fd3d56749084b9086550bbd6e3f85df24603e0482851ed1de

Observation edc0cf83-fb89-4bc0-8973-1cce97dd04a5 · outbound

This paper cites Unsupervised sounding object localization with bottom-up and top-down attention.

Object-aware Sound Source Localization via Audio-Visual Scene Understanding Unsupervised sounding object localization with bottom-up and top-down attention

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:52.876261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-06T23:20:47.985284Z digest=sha256:d44321077629a948df2b87db24b417a5e6ed8d34de2049576b44f477af33e64d

Observation e821573b-01b8-42d4-b2ee-4af3b572fda0 · outbound

This paper cites Flowgrad: Using motion for visual sound source localization.

Object-aware Sound Source Localization via Audio-Visual Scene Understanding Flowgrad: Using motion for visual sound source localization

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:52.534945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-06T23:20:48.061595Z digest=sha256:a8465da87f8f7c8cd0b1b8622911f0184ddb18c1e80834e68897a1fb8abbb0da

Observation 9b7a432f-dd00-4e2e-880b-98c6df771e1f · outbound

This paper cites Self-supervised predictive learning: A negative-free method for sound source localization in visual scenes.

Object-aware Sound Source Localization via Audio-Visual Scene Understanding Self-supervised predictive learning: A negative-free method for sound source localization in visual scenes

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:52.215295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-06T23:20:48.196924Z digest=sha256:6abbcad82febfa72332b3aa44c767380d3fc312b61b96a333ad1246fc8bb86a3

Observation 0099b1eb-197c-4982-bbdd-470db32f7dbe · outbound

This paper cites Learning audio-visual source localization via false negative aware contrastive learning.

Object-aware Sound Source Localization via Audio-Visual Scene Understanding Learning audio-visual source localization via false negative aware contrastive learning

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:51.998121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-06T23:20:48.316561Z digest=sha256:2c278a169a7d604c80908f0d68fbe001cc881d736c7226bb572bba4bacc82f9f

Observation 04793bae-3c14-47f6-b257-7b79e6ed984b · outbound

This paper cites Audio-visual spatial integration and recursive attention for robust sound source localization.

Object-aware Sound Source Localization via Audio-Visual Scene Understanding Audio-visual spatial integration and recursive attention for robust sound source localization

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:51.739907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-06T23:20:48.413073Z digest=sha256:4edfc8cbbcc4f719bc54134641799883bcb35883abd54a262719970df9af7b98

Observation a4793b17-fe85-4b43-8de0-53d48099bcae · outbound

This paper cites Watch video, catch keyword: Context-aware keyword attention for moment retrieval and highlight detection.

Object-aware Sound Source Localization via Audio-Visual Scene Understanding Watch video, catch keyword: Context-aware keyword attention for moment retrieval and highlight detection

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:51.514758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-06T23:20:48.481956Z digest=sha256:2dc9bf3cece8e31cdeec1ea48fafbed9008518101e72550dbd21bcd362d7cb5a

Observation de45388e-df01-4cc3-aafa-63a089495525 · outbound

This paper cites V2a-mapper: A lightweight solution for vision-to-audio generation by connecting foundation models.

Object-aware Sound Source Localization via Audio-Visual Scene Understanding V2a-mapper: A lightweight solution for vision-to-audio generation by connecting foundation models

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:51.283695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-06T23:20:48.538040Z digest=sha256:269de87c33ef47c12823fc62f8ead797e8541fe56cf23d921db2a23b73ccd921

Observation 02f9acb8-8e93-4110-856a-999f7b07683d · outbound

This paper cites Multimodal large language models: A survey.

Object-aware Sound Source Localization via Audio-Visual Scene Understanding Multimodal large language models: A survey

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:51.046232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-06T23:20:48.669587Z digest=sha256:b8a8c50caabc4514795c011e5f5f316f6d4c64536bce21cc110dd8c911b48b67

Observation caf9bdfc-63e5-467e-a050-368ba2d6297b · outbound

This paper cites Sonicvisionlm: Playing sound with vision language models.

Object-aware Sound Source Localization via Audio-Visual Scene Understanding Sonicvisionlm: Playing sound with vision language models

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:50.814775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-06T23:20:48.722137Z digest=sha256:f776f8ba24bb4b9ae0cef94fd60c50e1cf757ceb8f1cc4fac72e14c344eebe4a

Observation e09ecbac-9ff0-495b-9354-2b1ed2b00458 · outbound

This paper cites A proposal-based paradigm for self-supervised sound source localization in videos.

Object-aware Sound Source Localization via Audio-Visual Scene Understanding A proposal-based paradigm for self-supervised sound source localization in videos

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:50.588081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-06T23:20:48.915324Z digest=sha256:d2bf05f09254a8e27e0a9854cc9ddeb9bb277f64347dc1bf67073de8c75e6c18

Observation 0ae44f15-eda2-4597-b370-81865df6e06d · outbound

This paper cites A Survey on Multimodal Large Language Models.

Object-aware Sound Source Localization via Audio-Visual Scene Understanding A Survey on Multimodal Large Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:48.985249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:20:48.985249Z digest=sha256:2002a5aaff38ac594a53142d4162aecc707cebdb4f745beb7c81e712fd36a190

Observation 52a526d2-17b8-4015-8ead-64026d200fcc · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

Object-aware Sound Source Localization via Audio-Visual Scene Understanding Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:49.075435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:20:49.075435Z digest=sha256:ef3157ef3f0a016ef005a84b3b3a6c6a5b1e899bc4a5a5251efc7c18f2f0cd62

Observation 7d765c18-5d41-445c-a565-addc16c9924e · outbound

This paper cites The sound of pixels.

Object-aware Sound Source Localization via Audio-Visual Scene Understanding The sound of pixels

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:50.348537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-06T23:20:49.160956Z digest=sha256:0ce421da876813cbfccda712ff7531ed5a20d5ab7d2c8319214e0b41a0a3ce42

Observation 37e29985-6257-4e8f-af2e-e41c0dfa0027 · outbound

This paper cites Weakly supervised contrastive learning.

Object-aware Sound Source Localization via Audio-Visual Scene Understanding Weakly supervised contrastive learning

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:50.052232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-06T23:20:49.253720Z digest=sha256:aecf9e422b1bfb1f3f11421ae6bacc24fa68fdf9ce78d89543387789dc290283

Observation 69480215-0ca9-4260-b091-d46be833e310 · outbound

This paper cites Exploiting visual context semantics for sound source localization.

Object-aware Sound Source Localization via Audio-Visual Scene Understanding Exploiting visual context semantics for sound source localization

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:49.795672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-06T23:20:49.374967Z digest=sha256:1dcb0e530374ad5b824be82557c9834bd5a83a64dc6b2d1887c3b6daffb36149

Pith citing papers

No inbound Pith citation observations are available.