Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-09T19:24:20.464475Z
Paper Citation Record · LEDGER
As of 10 August 2026, this Paper Citation Record lists 54 of 54 outbound references and 2 inbound Pith citation observations for arXiv:2502.00358.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-09T19:24:20.464475Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-05T14:03:19.216337Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-01T23:06:21.321184Z
54 of 54 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 7590ae97-f740-46c7-9be8-1ea125844558 · outbound
Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? Look, listen and learn
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation f289fa68-93f3-4d34-bcf9-ab473b252dda · outbound
Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? Objects that sound
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation df9469b7-b5f7-4678-9a25-de7ab25123dc · outbound
Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? Multimodal syn- chronization in musical ensembles: Investigating audio and visual cues
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 1e6d503c-2281-4486-97d5-1353591a9210 · outbound
Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? Unresolved cited work
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 069c42ab-714b-4596-92a4-4fdb5ee3e258 · outbound
Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? Vggsound: A large-scale audio-visual dataset
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b0faeb93-dd04-482f-bb6f-8cea463edf83 · outbound
Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? Localizing visual sounds the hard way
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 250e5ccd-80b4-49fa-87e0-7ad5cea08e18 · outbound
Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? Deeplab: Semantic image segmentation with deep convolutional nets, atrous convolu- tion, and fully connected crfs
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 403ffee5-05a0-49ee-b7e8-5077c94e0489 · outbound
Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? Unraveling in- stance associations: A closer look for audio-visual segmenta- tion
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dc9b2cb6-1412-48c9-a838-bfe33c019e85 · outbound
Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? Masked-attention mask transformer for universal image segmentation
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6ad4a3f6-4f28-485a-808c-cc706cdd7143 · outbound
Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? Self-awareness for autonomous systems
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 78b6a000-990a-419b-a67a-47308e4220f5 · outbound
Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? Design of intelligent human-computer inter- action system for hard of hearing and non-disabled people
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation bf269e69-6846-44f7-9c99-a8ad2ed65804 · outbound
Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? Avsegformer: Audio-visual segmentation with trans- former
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 6fc51821-cfa7-497a-99d6-a60706b9127c · outbound
Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? Audio set: An ontology and human- labeled dataset for audio events
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation f63d4afd-c2b0-454e-94a5-27e305641c1c · outbound
Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? Open-Vocabulary Audio-Visual Semantic Segmentation
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 5d722fc3-def5-40ea-9c2c-83845017fc62 · outbound
Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? Multi-modal instruction tuned llms with fine-grained visual perception
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation e19bd29a-cf44-4012-8854-4a31d3ab6e4a · outbound
Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? Cnn archi- tectures for large-scale audio classification
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation ce6648f2-77f2-4db1-bc78-df2fd710fecb · outbound
Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? Discriminative sounding objects localization via self-supervised audiovisual match- ing
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 88c8a958-b0eb-496f-9f24-16dd9d0e50b6 · outbound
Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? Mix and local- ize: Localizing sound sources in mixtures
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 65436e9e-d69b-4047-b0cb-9ca97860d307 · outbound
Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? Char- acterising soundscape research in human-computer interac- tion
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 1afa3391-9c82-4eff-b4af-bfeab31b4a5f · outbound
Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? A Critical Assessment of Visual Sound Source Localization Models Including Negative Audio
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation ce65fb0d-0146-4d7f-bcbb-b6981be59b81 · outbound
Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? Adam: A Method for Stochastic Optimization
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c6ab8375-0aec-4086-9608-c6667db94dee · outbound
Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? Panoptic feature pyramid networks
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 202a8954-99e9-420a-878d-3bae15edb23d · outbound
Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? Segment any- thing
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d0c00e84-ce2b-47f1-9ab0-9ab5363234cf · outbound
Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? Catr: Combinatorial-dependence audio-queried transformer for audio-visual video segmentation
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation ea1d684b-ae15-4ee4-9100-e6f73176d1f8 · outbound
Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? Qdformer: Towards ro- bust audiovisual segmentation in complex environments with quantization-based semantic decomposition
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 5e9f3fc5-d5cb-4f4c-9f29-88eaf79d2a51 · outbound
Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? Audio-aware Query-enhanced Transformer for Audio-Visual Segmentation
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6650e041-06a1-4230-9ba1-b5e8094d9522 · outbound
Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? Annotation-free audio-visual segmentation
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 22b3d3ef-30b4-4719-b152-91e9622b1339 · outbound
Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? Decoupled Weight Decay Regularization
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 23c13131-eb53-4202-8298-3281ddda682d · outbound
Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? Deep learning for intelligent human–computer in- teraction
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 8aec4654-2427-443b-8e48-7723b5de656a · outbound
Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? Stepping Stones: A Progressive Training Strategy for Audio-Visual Semantic Segmentation
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation e1f277c6-4b03-44bb-b872-4c440f75739e · outbound
Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? T-vsl: Text-guided visual sound source localization in mixtures
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a4d91c5-2570-482e-a667-a3d74f5351d0 · outbound
Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? A closer look at weakly- supervised audio-visual source localization
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 5e6e0b2f-196d-471b-8e61-9aa5d2eea53b · outbound
Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? AV-SAM: Segment Anything Model Meets Audio-Visual Localization and Segmentation
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 70a85fd3-07f4-43e8-8e23-f392a5eeadad · outbound
Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? Multi-scale Multi-instance Visual Sound Localization and Segmentation
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 59d920ce-cba4-45de-93d0-df03dcf96615 · outbound
Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? Balanced multimodal learning via on-the-fly gradient modulation
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 72e4aa2f-0220-42b7-9ed8-d2ad24c307f5 · outbound
Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? Multiple sound sources localization from coarse to fine
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 9e795381-6bd6-47f5-87ad-53956781f52e · outbound
Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? Imagenet large scale visual recognition challenge
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6242c99e-24e2-457f-b4df-af9dc4d793a9 · outbound
Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? Acoustic self-awareness of autonomous sys- tems in a world of sounds
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 7c6af101-8642-46d1-96d2-f8c22e40c9d9 · outbound
Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? Learning to localize sound source in visual scenes
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 05e60b53-f984-4953-aa14-636976fbe689 · outbound
Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? Odor/taste integration and the perception of flavor
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 624ff189-17f9-4269-a392-e5f4fcf6ac16 · outbound
Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? Unveiling and Mitigating Bias in Audio Visual Segmentation
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation c664f050-f538-472c-a8fd-9484947e3cea · outbound
Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? Unified mul- tisensory perception: Weakly-supervised audio-visual video parsing
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 6011ac71-824c-45f6-a37b-5b3ddd781939 · outbound
Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? Assured Autonomy: Path Toward Living With Autonomous Systems We Can Trust
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6d0e6881-6005-4f12-aecb-dac4844dd551 · outbound
Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? AudioBench: A Universal Benchmark for Audio Large Language Models
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 001c15d9-596c-4bfa-a0c0-462ff0717e40 · outbound
Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? What makes train- ing multi-modal classification networks hard? In Proceed- ings of the IEEE/CVF conference on computer vision and pattern recognition, pages 12695–12705, 2020
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation e47230cf-78ed-4edf-8fbd-e72631751744 · outbound
Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? Pvt v2: Improved baselines with pyramid vision transformer
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 8a03ae3e-458a-4034-9434-c9856878e4ac · outbound
Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? Prompting segmentation with sound is gen- eralizable audio-visual source localizer
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation bd44ef52-0ab8-49ce-8156-768eb8918bc1 · outbound
Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? Can Textual Semantics Mitigate Sounding Object Segmentation Preference?
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation fc4dd78f-a199-4c1d-8114-ca097ae91fe3 · outbound
Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? Ref-AVS: Refer and Segment Objects in Audio-Visual Scenes
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 5f904dfc-2706-4170-9a55-646f8d6fc03e · outbound
Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? MMPareto: Boosting Multimodal Learning with Innocent Unimodal Assistance
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f9efe18-145c-4fdc-aa98-12120b0235c0 · outbound
Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? Analyzing audiovisual data for understanding user’s emotion in human- computer interaction environment
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation a4c5a844-0fa7-4fae-ba4a-a4ce09f3492d · outbound
Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? Cooperation does matter: Exploring multi-order bilateral relations for audio- visual segmentation
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation f416126d-e8c2-43e7-9803-d90922dd7128 · outbound
Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? Audio–visual segmentation
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6628f8ac-7e2f-4c0e-9575-3549b0176a7c · outbound
Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? 1, 2, 3, 4, 6, 7, 8, 14 Supplementary Material A
Reference 403
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 5ac2cb49-56c7-419a-97c4-702c5577b400 · inbound
Learning from Silence and Noise for Visual Sound Source Localization Do Audio-Visual Segmentation Models Truly Segment Sounding Objects?
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a7d7f17a-5f35-478f-9cd9-388ff19e134c · inbound
AVI-Bench: Toward Human-like Audio-Visual Intelligence of Omni-MLLMs Do Audio-Visual Segmentation Models Truly Segment Sounding Objects?
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.