Pith. sign in

Paper Citation Record · LEDGER

What's Making That Sound Right Now? Video-centric Audio-Visual Localization

As of 13 August 2026, this Paper Citation Record lists 37 of 37 outbound references and 0 inbound Pith citation observations for arXiv:2507.04667.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.04667 v2

Coverage vector

measured 37 of 37 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T19:47:28.634151Z

measured 37 of 37 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

37 of 37 outbound references displayed

  • verified exact0
  • verified fuzzy28
  • unresolved9
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5d25d9bc-8170-472e-af95-cf5808dfbd8f · outbound

This paper cites Self-supervised learning of audio-visual objects from video.

What's Making That Sound Right Now? Video-centric Audio-Visual Localization Self-supervised learning of audio-visual objects from video

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:29.045291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T19:47:28.120371Z digest=sha256:893a551dcc8dc6909c1fa680cc8033296b1a9b26beb4a745ab281ef5fb14fcde

Observation 1aa1ec69-56e9-4da3-90c5-fc901871e907 · outbound

This paper cites Vivit: A video vision transformer.

What's Making That Sound Right Now? Video-centric Audio-Visual Localization Vivit: A video vision transformer

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T19:47:28.262239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:47:28.262239Z digest=sha256:ce1b4e8c3bd013e1e05a57d1c1f74d093be67918396b9e793a30a50d30e44baa

Observation 3320ce3b-6bf2-4ec7-b40e-b04bbcb804e0 · outbound

This paper cites Soundnet: Learning sound representations from unlabeled video.

What's Making That Sound Right Now? Video-centric Audio-Visual Localization Soundnet: Learning sound representations from unlabeled video

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:29.025881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T19:47:28.379125Z digest=sha256:81bd612b6da338b8b346ce0b2867db12d5e7c9057e702ce38adf27ad99358821

Observation 6077323f-fcd4-43c7-ae78-702acd364b30 · outbound

This paper cites Is space-time attention all you need for video understanding? In ICML, page 4, 2021.

What's Making That Sound Right Now? Video-centric Audio-Visual Localization Is space-time attention all you need for video understanding? In ICML, page 4, 2021

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:29.013865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T19:47:28.442925Z digest=sha256:71bd354228b1fffb064ded4cd7558b29038105984af3e9e4ccd3ff89213379bf

Observation 4b2c3113-78f9-4010-b99d-9ab5c5d32334 · outbound

This paper cites Vggsound: A large-scale audio-visual dataset.

What's Making That Sound Right Now? Video-centric Audio-Visual Localization Vggsound: A large-scale audio-visual dataset

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:29.001333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T19:47:28.516909Z digest=sha256:355130dbf28fed46ea754e2523afab19c69514d82affbcc3dbe6bd522fe2be24

Observation 6db0cbf5-aea7-41de-a057-ff38cf9d5ef5 · outbound

This paper cites Localizing visual sounds the hard way.

What's Making That Sound Right Now? Video-centric Audio-Visual Localization Localizing visual sounds the hard way

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:28.989452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T19:47:28.520441Z digest=sha256:5a0f32c7a1a818928682f67504db04eb11cd06794606e6d7986272065e9552c0

Observation dbff7220-d553-448d-a284-ba15e0d503bc · outbound

This paper cites an unresolved cited work.

What's Making That Sound Right Now? Video-centric Audio-Visual Localization Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-06T19:47:28.978673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T19:47:28.523621Z digest=sha256:ea0f3591ab05a0f9af2bbee082de496175cac6f024940f75e1a2e6d495045d88

Observation 4bc77b1f-e344-41a9-a1ee-cecfed6faeef · outbound

This paper cites Gemmeke, Daniel P.

What's Making That Sound Right Now? Video-centric Audio-Visual Localization Gemmeke, Daniel P

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:28.966826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T19:47:28.526956Z digest=sha256:2c473467215e44f03f8c88a65d66da997f837fe7f743eefd389a9b7c5b902a54

Observation 3cf20378-63ed-4945-8d0a-db61d337984e · outbound

This paper cites Liu, David Harwath, Leonid Karlinsky, Hilde Kuehne, and James R.

What's Making That Sound Right Now? Video-centric Audio-Visual Localization Liu, David Harwath, Leonid Karlinsky, Hilde Kuehne, and James R

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:28.955716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T19:47:28.530439Z digest=sha256:67dac062bd23a81b21ea0a35c127434aca718480ba3aca99717b5fc9e5532207

Observation 2d615601-776e-4ec3-9ab3-5d54eb5462bb · outbound

This paper cites Dual mean-teacher: An unbiased semi-supervised framework for audio-visual source localization.

What's Making That Sound Right Now? Video-centric Audio-Visual Localization Dual mean-teacher: An unbiased semi-supervised framework for audio-visual source localization

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:28.945502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T19:47:28.533809Z digest=sha256:d7accfd6edf881fdb585974e281a54817d63f3779e0a083bea6cb6d2f2d25638

Observation ace7c10c-b03b-4f78-b764-a66dcc45fe87 · outbound

This paper cites Deep residual learning for image recognition.

What's Making That Sound Right Now? Video-centric Audio-Visual Localization Deep residual learning for image recognition

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T19:47:28.536881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:47:28.536881Z digest=sha256:24d45073fa5f50ca47c14847dc49fcd1e66cbbe080ae947f403be9bce811113d

Observation ccd609c0-720a-4ce9-b980-5a02641883a7 · outbound

This paper cites Deep multimodal clus- tering for unsupervised audiovisual learning.

What's Making That Sound Right Now? Video-centric Audio-Visual Localization Deep multimodal clus- tering for unsupervised audiovisual learning

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:28.928999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T19:47:28.540100Z digest=sha256:56930b7c3f76a850de06a7fe10053dc5c9dbad56c7604af6fb8081035bef1420

Observation 729f0c5a-698e-4cbc-9517-57dffab53d7d · outbound

This paper cites Mix and local- ize: Localizing sound sources in mixtures.

What's Making That Sound Right Now? Video-centric Audio-Visual Localization Mix and local- ize: Localizing sound sources in mixtures

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T19:47:28.543459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:47:28.543459Z digest=sha256:6787f4139e6890c7e5c5eb845a81f9ab2e4a2db2ce7d89695bb1dc8ddae1583d

Observation 6523c2d7-6bed-4c0b-9da2-d0bda742bb2b · outbound

This paper cites Egocentric audio-visual object localization.

What's Making That Sound Right Now? Video-centric Audio-Visual Localization Egocentric audio-visual object localization

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:28.910711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T19:47:28.546983Z digest=sha256:6f7e89960c3625326dbd846ccf3fe964ceddbb1108fcad2b019e64a23b122891

Observation b556ca4e-d0e4-4e44-8542-91f5fe1c0163 · outbound

This paper cites Learning to visually localize sound sources from mix- tures without prior source knowledge.

What's Making That Sound Right Now? Video-centric Audio-Visual Localization Learning to visually localize sound sources from mix- tures without prior source knowledge

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:28.899326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T19:47:28.550853Z digest=sha256:a89be018e4be19d5b31c87300e2cf67228e955a8a4e3c4b983073e1a7c967c60

Observation 7a83fb99-9282-440f-aa18-6bb13256a02c · outbound

This paper cites Segment any- thing.

What's Making That Sound Right Now? Video-centric Audio-Visual Localization Segment any- thing

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T19:47:28.555315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:47:28.555315Z digest=sha256:3193e8de4cb33ebaf5ea28d3dc9a9a276658cd24d747bc24db8b3e730b1c8a7e

Observation dd8d6a7a-de6a-4da8-aa46-360b7106627c · outbound

This paper cites Openimages: A public dataset for large-scale multi-label and multi-class image classification.

What's Making That Sound Right Now? Video-centric Audio-Visual Localization Openimages: A public dataset for large-scale multi-label and multi-class image classification

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:28.882386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T19:47:28.558796Z digest=sha256:7a0ee637d647d145d6dd118bd0b3472abe28518a83aa63a6e7bb4352341b3285

Observation 756d42fb-3d7f-48ff-805f-e39fab60c70a · outbound

This paper cites Ex- ploiting transformation invariance and equivariance for self- supervised sound localisation.

What's Making That Sound Right Now? Video-centric Audio-Visual Localization Ex- ploiting transformation invariance and equivariance for self- supervised sound localisation

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:28.872036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T19:47:28.562746Z digest=sha256:7b1728c9d63a401f5d4451a77ddf8af67cafe3ed8d16bf8e39dddd5b29d77827

Observation 64222c37-fb0c-4454-8815-d7f71efe56a5 · outbound

This paper cites A framework for multiple-instance learning.

What's Making That Sound Right Now? Video-centric Audio-Visual Localization A framework for multiple-instance learning

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:28.861737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T19:47:28.566583Z digest=sha256:bf78f053379d13614d2e09e9323ea3215f2e296cb69ecbd788da4407423bac28

Observation 4b09be43-d2d4-4ca5-9ca1-702280eeaded · outbound

This paper cites A closer look at weakly- supervised audio-visual source localization.

What's Making That Sound Right Now? Video-centric Audio-Visual Localization A closer look at weakly- supervised audio-visual source localization

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:28.850639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T19:47:28.570189Z digest=sha256:a71a772c69500e6475d7186aadecc9732f5a26e17f7a0d08a3ad0dc3d919d854

Observation 1f88f78c-6ba2-4fe3-8716-cefef78857d1 · outbound

This paper cites Localizing visual sounds the easy way.

What's Making That Sound Right Now? Video-centric Audio-Visual Localization Localizing visual sounds the easy way

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:28.839103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T19:47:28.574111Z digest=sha256:56ed11b43bc4d6eded82e710fc8e40cf3e169d97f401fd0052025026f7f8c793

Observation 4f6053e6-4114-40b8-a6a6-35d4f6652d1e · outbound

This paper cites Audio-visual grouping net- work for sound localization from mixtures.

What's Making That Sound Right Now? Video-centric Audio-Visual Localization Audio-visual grouping net- work for sound localization from mixtures

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:28.827331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T19:47:28.577704Z digest=sha256:0abaa120dc42aea9e5a7efca129bcb05f914b48160714935802ca40150e81f36

Observation bae5192a-2750-4212-9aaf-42fa629ceb0d · outbound

This paper cites Learn- ing representations from audio-visual spatial alignment.

What's Making That Sound Right Now? Video-centric Audio-Visual Localization Learn- ing representations from audio-visual spatial alignment

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:28.816423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T19:47:28.581769Z digest=sha256:59c6f3a73169265a687042e0a56dc00d309fa53d025db0030b02290cd7711b57

Observation 834fcbfb-1546-4393-b36b-9a6975cd24c2 · outbound

This paper cites Can clip help sound source localization? In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pages 5711–5720, 2024.

What's Making That Sound Right Now? Video-centric Audio-Visual Localization Can clip help sound source localization? In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pages 5711–5720, 2024

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:28.805247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T19:47:28.585329Z digest=sha256:5ecf03acb427716eef8f3476bb69093ad88299e4c4099b1d3b382bcbe6bc81ed

Observation b500f5a6-4eff-41d1-bf5b-fb9e77c5dedd · outbound

This paper cites Audio-visual object localization and separation using low- rank and sparsity.

What's Making That Sound Right Now? Video-centric Audio-Visual Localization Audio-visual object localization and separation using low- rank and sparsity

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:28.794982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T19:47:28.588633Z digest=sha256:45259988f970aa7bf5ba26daddcdd9df7f52b1ae035feeb197a684714130072e

Observation 99621e61-5379-41c4-8f43-73d3fa554fcc · outbound

This paper cites Multiple sound sources localization from coarse to fine.

What's Making That Sound Right Now? Video-centric Audio-Visual Localization Multiple sound sources localization from coarse to fine

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:28.784403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T19:47:28.592354Z digest=sha256:6f8526cac309b264348121b471b0a66a50f07b8eb879fc74b4f37639b2b55325

Observation 73338dac-fbe3-46c6-bbfb-0bdb26c5117e · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

What's Making That Sound Right Now? Video-centric Audio-Visual Localization Learning transferable visual models from natural language supervi- sion

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T19:47:28.595754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:47:28.595754Z digest=sha256:e90aa41f7118c60ab4b24a6a63dca5abfc4f56d1309a8657ed034559e8b016f0

Observation 1bed4dee-3972-4944-a7e4-65461141bebf · outbound

This paper cites Real-Time Flying Object Detection with YOLOv8.

What's Making That Sound Right Now? Video-centric Audio-Visual Localization Real-Time Flying Object Detection with YOLOv8

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T19:47:28.600083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:47:28.600083Z digest=sha256:84a279233a896841bd28e5c743fdad66f1a4c6cf6577d69aacf5140a3272b776

Observation 0bc6ec16-b7bf-4a6e-bba4-366b0b56ffc8 · outbound

This paper cites Learning to localize sound source in visual scenes.

What's Making That Sound Right Now? Video-centric Audio-Visual Localization Learning to localize sound source in visual scenes

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:28.767085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T19:47:28.604552Z digest=sha256:87473b51462853672432d641ccdd92cd34b3c64b1dc0563443a9434ae8cddd5c

Observation 46019f8d-bf6f-498f-bbea-5173633a346d · outbound

This paper cites Sound source local- ization is all about cross-modal alignment.

What's Making That Sound Right Now? Video-centric Audio-Visual Localization Sound source local- ization is all about cross-modal alignment

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:28.756066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T19:47:28.608218Z digest=sha256:d6ad8af380949cbdfa320019b7ebc64ecc9577d397198db36f17b96f3bba87c1

Observation 1a1ba15e-929a-455c-a454-bce4a1f66423 · outbound

This paper cites Aligning Sight and Sound: Advanced Sound Source Localization Through Audio-Visual Alignment.

What's Making That Sound Right Now? Video-centric Audio-Visual Localization Aligning Sight and Sound: Advanced Sound Source Localization Through Audio-Visual Alignment

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T19:47:28.611835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:47:28.611835Z digest=sha256:bcaa59c4760edfc15795410253e189a5d05edd924c4ad4c48e50736d112cbe7c

Observation 1108011b-256e-48b9-bb8f-624531372e34 · outbound

This paper cites Learning audio-visual source localization via false negative aware contrastive learning.

What's Making That Sound Right Now? Video-centric Audio-Visual Localization Learning audio-visual source localization via false negative aware contrastive learning

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:28.743921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T19:47:28.615704Z digest=sha256:eef5c508605c1c61343b5c35b06b837aad25855ce7326097e29fae2a658f932f

Observation 8da0e691-4b13-4b6b-b5b5-c3adc29ce8f7 · outbound

This paper cites Audio-visual event localization in unconstrained videos.

What's Making That Sound Right Now? Video-centric Audio-Visual Localization Audio-visual event localization in unconstrained videos

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:28.732696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T19:47:28.619505Z digest=sha256:c8602499c9d1d713187bcbb243d5e8d859549af6eec1ecfae414a86da2f2d55c

Observation d421adb9-2239-4eb8-b456-106d4c5313dc · outbound

This paper cites Scaling autoregressive video models.

What's Making That Sound Right Now? Video-centric Audio-Visual Localization Scaling autoregressive video models

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:28.722139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T19:47:28.622770Z digest=sha256:293fab8ae592c338023de79f02afe26334e3f8f7541b893c408eef0d97545c1e

Observation ebbd4606-7474-40bd-ada4-3083f496b229 · outbound

This paper cites The sound of pixels.

What's Making That Sound Right Now? Video-centric Audio-Visual Localization The sound of pixels

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:28.710135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T19:47:28.626178Z digest=sha256:d3827257e5dbdc833bb4dff60e5f417790d3ca367b53e121dbe95b8f325680ca

Observation b9dcf0b8-4232-4fb7-b00c-f6ec9afee9f1 · outbound

This paper cites Audio-visual segmentation.

What's Making That Sound Right Now? Video-centric Audio-Visual Localization Audio-visual segmentation

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:47:28.698811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T19:47:28.630375Z digest=sha256:07f196022c9c12bc8e6c6f8e1e1c87afca4d1b6b74ecaa74ac50fc20f874dcf5

Observation 4f421f1c-d8a9-4b1f-b122-40798004a8a2 · outbound

This paper cites Audio-Visual Segmentation with Semantics.

What's Making That Sound Right Now? Video-centric Audio-Visual Localization Audio-Visual Segmentation with Semantics

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T19:47:28.634151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:47:28.634151Z digest=sha256:0187af62416d858df2686ffca6495cbd29a8747ab51d61bedf414d6f7b195d90

Pith citing papers

No inbound Pith citation observations are available.