Pith. sign in

Paper Citation Record · LEDGER

What's Making That Sound Right Now? Video-centric Audio-Visual Localization

As of 8 August 2026, this Paper Citation Record lists 37 of 37 outbound references and 0 inbound Pith citation observations for arXiv:2507.04667.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.04667 v2

Coverage vector

measured 37 of 37 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T19:47:28.634151Z

measured 37 of 37 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

37 of 37 outbound references displayed

  • verified exact0
  • verified fuzzy28
  • unresolved9
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5d25d9bc-8170-472e-af95-cf5808dfbd8f · outbound

This paper cites Self-supervised learning of audio-visual objects from video.

What's Making That Sound Right Now? Video-centric Audio-Visual Localization Self-supervised learning of audio-visual objects from video

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:29.045291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:47:28.120371Z digest=sha256:32a7b87a845ee75522db90db4434f39dd19adaa0ec22a403e9df63ae6933e07a

Observation 1aa1ec69-56e9-4da3-90c5-fc901871e907 · outbound

This paper cites Vivit: A video vision transformer.

What's Making That Sound Right Now? Video-centric Audio-Visual Localization Vivit: A video vision transformer

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T19:47:28.262239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:47:28.262239Z digest=sha256:417600c5a03462569e4f3cb627c3b586b9c42780e34bc5fe444a81f5dd1acd2c

Observation 3320ce3b-6bf2-4ec7-b40e-b04bbcb804e0 · outbound

This paper cites Soundnet: Learning sound representations from unlabeled video.

What's Making That Sound Right Now? Video-centric Audio-Visual Localization Soundnet: Learning sound representations from unlabeled video

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:29.025881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:47:28.379125Z digest=sha256:378cd4e7ba2f5e0e6ffb5f419c67f45569d42e13f5937ec33a77906688fe789b

Observation 6077323f-fcd4-43c7-ae78-702acd364b30 · outbound

This paper cites Is space-time attention all you need for video understanding? In ICML, page 4, 2021.

What's Making That Sound Right Now? Video-centric Audio-Visual Localization Is space-time attention all you need for video understanding? In ICML, page 4, 2021

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:29.013865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:47:28.442925Z digest=sha256:b18db91962c3d3804fcb45fb2a7c8e400aec10247730dde374ecdad9e6fda486

Observation 4b2c3113-78f9-4010-b99d-9ab5c5d32334 · outbound

This paper cites Vggsound: A large-scale audio-visual dataset.

What's Making That Sound Right Now? Video-centric Audio-Visual Localization Vggsound: A large-scale audio-visual dataset

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:29.001333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:47:28.516909Z digest=sha256:c086a283409dae5cb7e658f2b9bdc2d261cbaa3cd0285b90e61d04ac5f658999

Observation 6db0cbf5-aea7-41de-a057-ff38cf9d5ef5 · outbound

This paper cites Localizing visual sounds the hard way.

What's Making That Sound Right Now? Video-centric Audio-Visual Localization Localizing visual sounds the hard way

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:28.989452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:47:28.520441Z digest=sha256:673c830929107deab53f1000dc2362c16a3d356dafe07f53d18752d23abf271c

Observation dbff7220-d553-448d-a284-ba15e0d503bc · outbound

This paper cites an unresolved cited work.

What's Making That Sound Right Now? Video-centric Audio-Visual Localization Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-06T19:47:28.978673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:47:28.523621Z digest=sha256:6faae954bb41fbfaceb43c31313c6a5d15131c9eab55c9a634a5f50a5a27851d

Observation 4bc77b1f-e344-41a9-a1ee-cecfed6faeef · outbound

This paper cites Gemmeke, Daniel P.

What's Making That Sound Right Now? Video-centric Audio-Visual Localization Gemmeke, Daniel P

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:28.966826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:47:28.526956Z digest=sha256:87f76de04839d65692aaecd77433d1e5a437e54dc8b3b0e8b1e3683e6d3bf611

Observation 3cf20378-63ed-4945-8d0a-db61d337984e · outbound

This paper cites Liu, David Harwath, Leonid Karlinsky, Hilde Kuehne, and James R.

What's Making That Sound Right Now? Video-centric Audio-Visual Localization Liu, David Harwath, Leonid Karlinsky, Hilde Kuehne, and James R

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:28.955716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:47:28.530439Z digest=sha256:361563edb5a6172b853341aebc346fead1341d9a2beaaf4fb2e986782e475cb3

Observation 2d615601-776e-4ec3-9ab3-5d54eb5462bb · outbound

This paper cites Dual mean-teacher: An unbiased semi-supervised framework for audio-visual source localization.

What's Making That Sound Right Now? Video-centric Audio-Visual Localization Dual mean-teacher: An unbiased semi-supervised framework for audio-visual source localization

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:28.945502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:47:28.533809Z digest=sha256:92ec0a5c572f76ca95bf3bfc9eab88c83d2e4f0fae7e204f8b1bcf887c6c7663

Observation ace7c10c-b03b-4f78-b764-a66dcc45fe87 · outbound

This paper cites Deep residual learning for image recognition.

What's Making That Sound Right Now? Video-centric Audio-Visual Localization Deep residual learning for image recognition

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T19:47:28.536881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:47:28.536881Z digest=sha256:5e9919559d079dbe9097407921cea62583b211879bb9d482291e33b160c4872c

Observation ccd609c0-720a-4ce9-b980-5a02641883a7 · outbound

This paper cites Deep multimodal clus- tering for unsupervised audiovisual learning.

What's Making That Sound Right Now? Video-centric Audio-Visual Localization Deep multimodal clus- tering for unsupervised audiovisual learning

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:28.928999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:47:28.540100Z digest=sha256:b01e7f22c7394745996415953b406d05ad114e984cb6df7865556416ed9b6182

Observation 729f0c5a-698e-4cbc-9517-57dffab53d7d · outbound

This paper cites Mix and local- ize: Localizing sound sources in mixtures.

What's Making That Sound Right Now? Video-centric Audio-Visual Localization Mix and local- ize: Localizing sound sources in mixtures

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T19:47:28.543459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:47:28.543459Z digest=sha256:6985195f9ee88649656b5b2c2c013aeec8010ec48434f93ae9150673f80210e1

Observation 6523c2d7-6bed-4c0b-9da2-d0bda742bb2b · outbound

This paper cites Egocentric audio-visual object localization.

What's Making That Sound Right Now? Video-centric Audio-Visual Localization Egocentric audio-visual object localization

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:28.910711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:47:28.546983Z digest=sha256:dd81b77f5971c9f1f41a3985da87b87e72dd08681d9605643c5fb956af5fb14f

Observation b556ca4e-d0e4-4e44-8542-91f5fe1c0163 · outbound

This paper cites Learning to visually localize sound sources from mix- tures without prior source knowledge.

What's Making That Sound Right Now? Video-centric Audio-Visual Localization Learning to visually localize sound sources from mix- tures without prior source knowledge

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:28.899326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:47:28.550853Z digest=sha256:07330123f183eb0c6c586ae2074002b58bda2cda639e2067db50725d78efe251

Observation 7a83fb99-9282-440f-aa18-6bb13256a02c · outbound

This paper cites Segment any- thing.

What's Making That Sound Right Now? Video-centric Audio-Visual Localization Segment any- thing

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T19:47:28.555315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:47:28.555315Z digest=sha256:5af2dff9536e0a3c8e1ff7fcf82455b532a56bdfc8e72b78355fa083d49feecf

Observation dd8d6a7a-de6a-4da8-aa46-360b7106627c · outbound

This paper cites Openimages: A public dataset for large-scale multi-label and multi-class image classification.

What's Making That Sound Right Now? Video-centric Audio-Visual Localization Openimages: A public dataset for large-scale multi-label and multi-class image classification

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:28.882386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:47:28.558796Z digest=sha256:e6be4b139eb68c688eb29d7b867fddc092c50727ec5d913f652b63a3a1582140

Observation 756d42fb-3d7f-48ff-805f-e39fab60c70a · outbound

This paper cites Ex- ploiting transformation invariance and equivariance for self- supervised sound localisation.

What's Making That Sound Right Now? Video-centric Audio-Visual Localization Ex- ploiting transformation invariance and equivariance for self- supervised sound localisation

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:28.872036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:47:28.562746Z digest=sha256:12000b54fd54922a2994cd66722e6a9aab2c80db87225ac32fc185715260e9e0

Observation 64222c37-fb0c-4454-8815-d7f71efe56a5 · outbound

This paper cites A framework for multiple-instance learning.

What's Making That Sound Right Now? Video-centric Audio-Visual Localization A framework for multiple-instance learning

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:28.861737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:47:28.566583Z digest=sha256:2f397147cbbd59495a50d48965409c4f144a070c03ad8448ac72a64a3edd907c

Observation 4b09be43-d2d4-4ca5-9ca1-702280eeaded · outbound

This paper cites A closer look at weakly- supervised audio-visual source localization.

What's Making That Sound Right Now? Video-centric Audio-Visual Localization A closer look at weakly- supervised audio-visual source localization

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:28.850639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:47:28.570189Z digest=sha256:aa07174921ab82dce196a8a0a0ea219e913539ff8ca8f703e7beada7cf62dff2

Observation 1f88f78c-6ba2-4fe3-8716-cefef78857d1 · outbound

This paper cites Localizing visual sounds the easy way.

What's Making That Sound Right Now? Video-centric Audio-Visual Localization Localizing visual sounds the easy way

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:28.839103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:47:28.574111Z digest=sha256:688a82d69b6aeb02c8345113537ee1b8a86754a83f39215bac4e1b3a2fc4e8eb

Observation 4f6053e6-4114-40b8-a6a6-35d4f6652d1e · outbound

This paper cites Audio-visual grouping net- work for sound localization from mixtures.

What's Making That Sound Right Now? Video-centric Audio-Visual Localization Audio-visual grouping net- work for sound localization from mixtures

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:28.827331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:47:28.577704Z digest=sha256:329bd865acb9cf6bfcbf85ee3d4ba64807e6b6a96eacf5a415919622a3dd34cc

Observation bae5192a-2750-4212-9aaf-42fa629ceb0d · outbound

This paper cites Learn- ing representations from audio-visual spatial alignment.

What's Making That Sound Right Now? Video-centric Audio-Visual Localization Learn- ing representations from audio-visual spatial alignment

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:28.816423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:47:28.581769Z digest=sha256:3bfc5f2a66fc5eb2c80be886f64405329311632a0720e49a5e04cef2c42e1474

Observation 834fcbfb-1546-4393-b36b-9a6975cd24c2 · outbound

This paper cites Can clip help sound source localization? In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pages 5711–5720, 2024.

What's Making That Sound Right Now? Video-centric Audio-Visual Localization Can clip help sound source localization? In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pages 5711–5720, 2024

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:28.805247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:47:28.585329Z digest=sha256:26024be0ae9c513c9b94dd1b56e8194f0db7dfe61b937e36ff83b2bf9ec32c26

Observation b500f5a6-4eff-41d1-bf5b-fb9e77c5dedd · outbound

This paper cites Audio-visual object localization and separation using low- rank and sparsity.

What's Making That Sound Right Now? Video-centric Audio-Visual Localization Audio-visual object localization and separation using low- rank and sparsity

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:28.794982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:47:28.588633Z digest=sha256:8a3e43d0e18fcf8a620c9a19f70bf97b0fcceb1c64bdc33034222cfeef199093

Observation 99621e61-5379-41c4-8f43-73d3fa554fcc · outbound

This paper cites Multiple sound sources localization from coarse to fine.

What's Making That Sound Right Now? Video-centric Audio-Visual Localization Multiple sound sources localization from coarse to fine

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:28.784403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:47:28.592354Z digest=sha256:838dc49a73a3067fbc57b551306b5e90d0ec268b13f8c99108a413ef148f152b

Observation 73338dac-fbe3-46c6-bbfb-0bdb26c5117e · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

What's Making That Sound Right Now? Video-centric Audio-Visual Localization Learning transferable visual models from natural language supervi- sion

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T19:47:28.595754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:47:28.595754Z digest=sha256:77a527dc4da6aca288bd65b90bf01524b35e83679ffc365524eacd4497520d42

Observation 1bed4dee-3972-4944-a7e4-65461141bebf · outbound

This paper cites Real-Time Flying Object Detection with YOLOv8.

What's Making That Sound Right Now? Video-centric Audio-Visual Localization Real-Time Flying Object Detection with YOLOv8

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T19:47:28.600083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:47:28.600083Z digest=sha256:72616a382422117bef75dd462fb115be7506a3cbae1a9d69f4d394df828535b9

Observation 0bc6ec16-b7bf-4a6e-bba4-366b0b56ffc8 · outbound

This paper cites Learning to localize sound source in visual scenes.

What's Making That Sound Right Now? Video-centric Audio-Visual Localization Learning to localize sound source in visual scenes

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:28.767085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:47:28.604552Z digest=sha256:b93fffd27c5e743f3cc047ebfcf14233c4f5d90dd43d0caeb0db4c17662e8094

Observation 46019f8d-bf6f-498f-bbea-5173633a346d · outbound

This paper cites Sound source local- ization is all about cross-modal alignment.

What's Making That Sound Right Now? Video-centric Audio-Visual Localization Sound source local- ization is all about cross-modal alignment

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:28.756066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:47:28.608218Z digest=sha256:135be0240a5c890b87c9bd8a77ae6a782db9ead5e859ec7648b8f357007ee3b1

Observation 1a1ba15e-929a-455c-a454-bce4a1f66423 · outbound

This paper cites Aligning Sight and Sound: Advanced Sound Source Localization Through Audio-Visual Alignment.

What's Making That Sound Right Now? Video-centric Audio-Visual Localization Aligning Sight and Sound: Advanced Sound Source Localization Through Audio-Visual Alignment

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T19:47:28.611835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:47:28.611835Z digest=sha256:44c400683701a75e682d06782f59f44ee3e511b23ce7f59f2201c7fa9ea8a63b

Observation 1108011b-256e-48b9-bb8f-624531372e34 · outbound

This paper cites Learning audio-visual source localization via false negative aware contrastive learning.

What's Making That Sound Right Now? Video-centric Audio-Visual Localization Learning audio-visual source localization via false negative aware contrastive learning

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:28.743921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:47:28.615704Z digest=sha256:5b870ed55bd1012c114b8ab9c965235a9bcb60a38118b9ae5fba9098e07b699b

Observation 8da0e691-4b13-4b6b-b5b5-c3adc29ce8f7 · outbound

This paper cites Audio-visual event localization in unconstrained videos.

What's Making That Sound Right Now? Video-centric Audio-Visual Localization Audio-visual event localization in unconstrained videos

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:28.732696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:47:28.619505Z digest=sha256:0b01a82377c618a1b4cdf919113344d5015fd437b1b011b2dba18b5dab24563e

Observation d421adb9-2239-4eb8-b456-106d4c5313dc · outbound

This paper cites Scaling autoregressive video models.

What's Making That Sound Right Now? Video-centric Audio-Visual Localization Scaling autoregressive video models

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:28.722139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:47:28.622770Z digest=sha256:75ecefcb2ceae4731625c517def0cce1d421adf8d3943169b770c64ccca67381

Observation ebbd4606-7474-40bd-ada4-3083f496b229 · outbound

This paper cites The sound of pixels.

What's Making That Sound Right Now? Video-centric Audio-Visual Localization The sound of pixels

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:28.710135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:47:28.626178Z digest=sha256:0d7144ddd478502bca84ecd5ad20a68c86bcae14a92b03aaa6a09d9b462cde16

Observation b9dcf0b8-4232-4fb7-b00c-f6ec9afee9f1 · outbound

This paper cites Audio-visual segmentation.

What's Making That Sound Right Now? Video-centric Audio-Visual Localization Audio-visual segmentation

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:28.698811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:47:28.630375Z digest=sha256:4f840e8996ae5b596b057d84bba2cc336a3d3bc58c67ee61e16255967cf18c8f

Observation 4f421f1c-d8a9-4b1f-b122-40798004a8a2 · outbound

This paper cites Audio-Visual Segmentation with Semantics.

What's Making That Sound Right Now? Video-centric Audio-Visual Localization Audio-Visual Segmentation with Semantics

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T19:47:28.634151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:47:28.634151Z digest=sha256:22abd1e590ba00f67d7fea70254e889ec8ec69558893fffdce220f914950ac36

Pith citing papers

No inbound Pith citation observations are available.