Pith. sign in

Paper Citation Record · LEDGER

Do Audio-Visual Segmentation Models Truly Segment Sounding Objects?

As of 10 August 2026, this Paper Citation Record lists 54 of 54 outbound references and 2 inbound Pith citation observations for arXiv:2502.00358.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.00358 v2

Coverage vector

measured 54 of 54 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-09T19:24:20.464475Z

measured 56 of 56 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T14:03:19.216337Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T23:06:21.321184Z

Reference resolution

54 of 54 outbound references displayed

  • verified exact6
  • verified fuzzy31
  • unresolved16
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 7590ae97-f740-46c7-9be8-1ea125844558 · outbound

This paper cites Look, listen and learn.

Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? Look, listen and learn

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:24:21.108505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T19:24:20.277575Z digest=sha256:0810dad29257f52b1e9a641741bc09d742de73464ed518ebb42199cccbcf5482

Observation f289fa68-93f3-4d34-bcf9-ab473b252dda · outbound

This paper cites Objects that sound.

Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? Objects that sound

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:24:21.098943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T19:24:20.282196Z digest=sha256:3ee09e45319bb1ab6735dd19299b8aed43ce00e93b2b485abdf73e065089cf29

Observation df9469b7-b5f7-4678-9a25-de7ab25123dc · outbound

This paper cites Multimodal syn- chronization in musical ensembles: Investigating audio and visual cues.

Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? Multimodal syn- chronization in musical ensembles: Investigating audio and visual cues

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:24:21.088660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T19:24:20.286016Z digest=sha256:21d6f0afd600bbeeba10fc5c7cccac6cac3e4000fc09ce9ff8925ac8d8f62c7c

Observation 1e6d503c-2281-4486-97d5-1353591a9210 · outbound

This paper cites an unresolved cited work.

Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-09T19:24:21.077744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T19:24:20.290327Z digest=sha256:ffe749cde2f20aa3ac3f02fceb6a7e012797378a7617fb494ab11bb6df2e17c5

Observation 069c42ab-714b-4596-92a4-4fdb5ee3e258 · outbound

This paper cites Vggsound: A large-scale audio-visual dataset.

Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? Vggsound: A large-scale audio-visual dataset

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-09T19:24:20.294182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:24:20.294182Z digest=sha256:4ceade88cb0e1ba93afec7999dcec537f62a396e23ee7696a6a93555fdfcb0fa

Observation b0faeb93-dd04-482f-bb6f-8cea463edf83 · outbound

This paper cites Localizing visual sounds the hard way.

Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? Localizing visual sounds the hard way

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:24:21.059677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T19:24:20.298792Z digest=sha256:b881f64587559d7886dc2bbb10d6aed4c2783a29337d7397e58bab248d245c1e

Observation 250e5ccd-80b4-49fa-87e0-7ad5cea08e18 · outbound

This paper cites Deeplab: Semantic image segmentation with deep convolutional nets, atrous convolu- tion, and fully connected crfs.

Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? Deeplab: Semantic image segmentation with deep convolutional nets, atrous convolu- tion, and fully connected crfs

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:24:21.049165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T19:24:20.302763Z digest=sha256:6b30223441d802c570b397b4b8555753dc63d8b9bf0af7b02dadee645b7b4545

Observation 403ffee5-05a0-49ee-b7e8-5077c94e0489 · outbound

This paper cites Unraveling in- stance associations: A closer look for audio-visual segmenta- tion.

Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? Unraveling in- stance associations: A closer look for audio-visual segmenta- tion

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-09T19:24:20.306658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:24:20.306658Z digest=sha256:edceb9c4a3f8806c9c8533789ee705a97185f3366bf55a11eb6e7c7db95941a4

Observation dc9b2cb6-1412-48c9-a838-bfe33c019e85 · outbound

This paper cites Masked-attention mask transformer for universal image segmentation.

Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? Masked-attention mask transformer for universal image segmentation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-09T19:24:20.310828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:24:20.310828Z digest=sha256:05f95e7a256d399b05366fd8d8d7d51dc16a233a5587d2f513d4b69ac244b034

Observation 6ad4a3f6-4f28-485a-808c-cc706cdd7143 · outbound

This paper cites Self-awareness for autonomous systems.

Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? Self-awareness for autonomous systems

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:24:21.025332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T19:24:20.314450Z digest=sha256:f0477483fc05e7c83cf484e68bd1aa8492a3ca73d206125c53d7fd5322cac566

Observation 78b6a000-990a-419b-a67a-47308e4220f5 · outbound

This paper cites Design of intelligent human-computer inter- action system for hard of hearing and non-disabled people.

Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? Design of intelligent human-computer inter- action system for hard of hearing and non-disabled people

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:24:21.014811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T19:24:20.317908Z digest=sha256:92c26cc854b1c14addbd656247eefc01d6ee124d212ba060e3c722033cc399fe

Observation bf269e69-6846-44f7-9c99-a8ad2ed65804 · outbound

This paper cites Avsegformer: Audio-visual segmentation with trans- former.

Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? Avsegformer: Audio-visual segmentation with trans- former

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:24:21.004497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T19:24:20.321551Z digest=sha256:41127414e1fb2412a0d97f1669c84388601fdac0e59409b95b18ca60d73ae2ce

Observation 6fc51821-cfa7-497a-99d6-a60706b9127c · outbound

This paper cites Audio set: An ontology and human- labeled dataset for audio events.

Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? Audio set: An ontology and human- labeled dataset for audio events

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:24:20.994065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T19:24:20.324911Z digest=sha256:60262ebb53c42ec663e62234b607226aaa930ef1d44abd6427d55d3da8851b6b

Observation f63d4afd-c2b0-454e-94a5-27e305641c1c · outbound

This paper cites Open-Vocabulary Audio-Visual Semantic Segmentation.

Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? Open-Vocabulary Audio-Visual Semantic Segmentation

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-08-09T19:24:20.671139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T19:24:20.328347Z digest=sha256:49bb1a83c1e050da1e8d3c443914527ce20c1e45a8e752eef16aacb1b8be43ba

Observation 5d722fc3-def5-40ea-9c2c-83845017fc62 · outbound

This paper cites Multi-modal instruction tuned llms with fine-grained visual perception.

Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? Multi-modal instruction tuned llms with fine-grained visual perception

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:24:20.983016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T19:24:20.332139Z digest=sha256:455df7d36198b48579375500ae54ae2ab78d7acc65d817b1769ddf41749ec3ae

Observation e19bd29a-cf44-4012-8854-4a31d3ab6e4a · outbound

This paper cites Cnn archi- tectures for large-scale audio classification.

Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? Cnn archi- tectures for large-scale audio classification

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:24:20.971597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T19:24:20.335980Z digest=sha256:07ce9104b9291b459d69ef102bee26d9513b3c1989e771522289aed365cb460f

Observation ce6648f2-77f2-4db1-bc78-df2fd710fecb · outbound

This paper cites Discriminative sounding objects localization via self-supervised audiovisual match- ing.

Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? Discriminative sounding objects localization via self-supervised audiovisual match- ing

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:24:20.960336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T19:24:20.339591Z digest=sha256:f37d95d01e56f4edcd9048d33a0f20394dfe114b93dfcad29212485dc3f21ebf

Observation 88c8a958-b0eb-496f-9f24-16dd9d0e50b6 · outbound

This paper cites Mix and local- ize: Localizing sound sources in mixtures.

Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? Mix and local- ize: Localizing sound sources in mixtures

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:24:20.950095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T19:24:20.343114Z digest=sha256:8e2ae6bc8e72cbcc481f91ef27b5b73bf498283655cb910cce79c2cdb63652f7

Observation 65436e9e-d69b-4047-b0cb-9ca97860d307 · outbound

This paper cites Char- acterising soundscape research in human-computer interac- tion.

Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? Char- acterising soundscape research in human-computer interac- tion

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:24:20.941001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T19:24:20.346602Z digest=sha256:5148869fbfe9878698939147cccd79775a94d76356a689d0843e4ef306f1bee7

Observation 1afa3391-9c82-4eff-b4af-bfeab31b4a5f · outbound

This paper cites A Critical Assessment of Visual Sound Source Localization Models Including Negative Audio.

Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? A Critical Assessment of Visual Sound Source Localization Models Including Negative Audio

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-08-09T19:24:20.645939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T19:24:20.350125Z digest=sha256:b57cc7835971dab1494933dddb1dbc3026642215dbee3c4ba7f7067462d745d4

Observation ce65fb0d-0146-4d7f-bcbb-b6981be59b81 · outbound

This paper cites Adam: A Method for Stochastic Optimization.

Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? Adam: A Method for Stochastic Optimization

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-09T19:24:20.353719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:24:20.353719Z digest=sha256:85cba615bcf5b8730da2266ad59d25a9dae7d697c85efef0f69417e78885917c

Observation c6ab8375-0aec-4086-9608-c6667db94dee · outbound

This paper cites Panoptic feature pyramid networks.

Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? Panoptic feature pyramid networks

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:24:20.931140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T19:24:20.357125Z digest=sha256:25c2126c9dbf18ecef2fa52cc72f9f8115d8c2c5bfaa502d79b5c4f789ff908e

Observation 202a8954-99e9-420a-878d-3bae15edb23d · outbound

This paper cites Segment any- thing.

Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? Segment any- thing

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-09T19:24:20.360276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:24:20.360276Z digest=sha256:2c6137f054305fd90599118af23af4f3609310f28a0514ca58bbe7e0ea76118a

Observation d0c00e84-ce2b-47f1-9ab0-9ab5363234cf · outbound

This paper cites Catr: Combinatorial-dependence audio-queried transformer for audio-visual video segmentation.

Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? Catr: Combinatorial-dependence audio-queried transformer for audio-visual video segmentation

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:24:20.915529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T19:24:20.363453Z digest=sha256:40c14e4bebae3679358c09826d185574f686b26538c4bf5c0f5f7539b097a95e

Observation ea1d684b-ae15-4ee4-9100-e6f73176d1f8 · outbound

This paper cites Qdformer: Towards ro- bust audiovisual segmentation in complex environments with quantization-based semantic decomposition.

Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? Qdformer: Towards ro- bust audiovisual segmentation in complex environments with quantization-based semantic decomposition

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:24:20.905024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T19:24:20.366364Z digest=sha256:ac46307d9856fcc200cd63dc73529b50a9fe42c2ee4e2abe69a4572517cbf618

Observation 5e9f3fc5-d5cb-4f4c-9f29-88eaf79d2a51 · outbound

This paper cites Audio-aware Query-enhanced Transformer for Audio-Visual Segmentation.

Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? Audio-aware Query-enhanced Transformer for Audio-Visual Segmentation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-09T19:24:20.369849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:24:20.369849Z digest=sha256:94433e0cdadfbf00defc8bbb9485c2cb3a27cf4aa98445e243b25552be092a1f

Observation 6650e041-06a1-4230-9ba1-b5e8094d9522 · outbound

This paper cites Annotation-free audio-visual segmentation.

Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? Annotation-free audio-visual segmentation

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:24:20.894454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T19:24:20.374215Z digest=sha256:d6f7d01c8c07faae4161a1060282a134496f47fd01076cd37685457f1dd73da2

Observation 22b3d3ef-30b4-4719-b152-91e9622b1339 · outbound

This paper cites Decoupled Weight Decay Regularization.

Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? Decoupled Weight Decay Regularization

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-09T19:24:20.377738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:24:20.377738Z digest=sha256:0a141ded677d9c916aefcf2da6108b9053df085b74b4535bd5b01ab31811ab03

Observation 23c13131-eb53-4202-8298-3281ddda682d · outbound

This paper cites Deep learning for intelligent human–computer in- teraction.

Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? Deep learning for intelligent human–computer in- teraction

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:24:20.883756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T19:24:20.381448Z digest=sha256:b078b44125904c566d847ef221cc38f6891efd619cbf35eb70c2ae808bce6192

Observation 8aec4654-2427-443b-8e48-7723b5de656a · outbound

This paper cites Stepping Stones: A Progressive Training Strategy for Audio-Visual Semantic Segmentation.

Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? Stepping Stones: A Progressive Training Strategy for Audio-Visual Semantic Segmentation

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-08-09T19:24:20.600464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T19:24:20.384721Z digest=sha256:203e6895fc31a276743d8e7dc7485033d520f97c065083b1ad8c9ad97f3c5ec0

Observation e1f277c6-4b03-44bb-b872-4c440f75739e · outbound

This paper cites T-vsl: Text-guided visual sound source localization in mixtures.

Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? T-vsl: Text-guided visual sound source localization in mixtures

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-09T19:24:20.388365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:24:20.388365Z digest=sha256:339acd0ae4a5a5ce0108ce213e8a124ab2a208a052ffbf885998a2bb7854ce78

Observation 5a4d91c5-2570-482e-a667-a3d74f5351d0 · outbound

This paper cites A closer look at weakly- supervised audio-visual source localization.

Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? A closer look at weakly- supervised audio-visual source localization

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:24:20.863671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T19:24:20.391561Z digest=sha256:ca061a342a1b50b2c829a040c7da3adfa44e80543fd456171c88f0fd74a234ac

Observation 5e6e0b2f-196d-471b-8e61-9aa5d2eea53b · outbound

This paper cites AV-SAM: Segment Anything Model Meets Audio-Visual Localization and Segmentation.

Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? AV-SAM: Segment Anything Model Meets Audio-Visual Localization and Segmentation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-09T19:24:20.394989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:24:20.394989Z digest=sha256:372e292d276e672ff6140ac07cbc950bb61ef2089ed71887721323f420488bb2

Observation 70a85fd3-07f4-43e8-8e23-f392a5eeadad · outbound

This paper cites Multi-scale Multi-instance Visual Sound Localization and Segmentation.

Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? Multi-scale Multi-instance Visual Sound Localization and Segmentation

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-09T19:24:20.398584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:24:20.398584Z digest=sha256:918cf27953432db43d869c9054f8ad4dce83af18933d7c40d64bb58b6730d7a1

Observation 59d920ce-cba4-45de-93d0-df03dcf96615 · outbound

This paper cites Balanced multimodal learning via on-the-fly gradient modulation.

Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? Balanced multimodal learning via on-the-fly gradient modulation

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:24:20.850493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T19:24:20.402033Z digest=sha256:13863a47d8dd1925a21c8348b5419d086312d3b1914ef704bcd3af69fe1ad012

Observation 72e4aa2f-0220-42b7-9ed8-d2ad24c307f5 · outbound

This paper cites Multiple sound sources localization from coarse to fine.

Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? Multiple sound sources localization from coarse to fine

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:24:20.837026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T19:24:20.406084Z digest=sha256:048b07c123989e0666e5164170fda717243f03dcbf1e18529ab72b29f7bf80d7

Observation 9e795381-6bd6-47f5-87ad-53956781f52e · outbound

This paper cites Imagenet large scale visual recognition challenge.

Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? Imagenet large scale visual recognition challenge

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-09T19:24:20.409475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:24:20.409475Z digest=sha256:9861c2302fd2916de162acac4cf39c846116feb31320b85e04d5ab703b883da5

Observation 6242c99e-24e2-457f-b4df-af9dc4d793a9 · outbound

This paper cites Acoustic self-awareness of autonomous sys- tems in a world of sounds.

Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? Acoustic self-awareness of autonomous sys- tems in a world of sounds

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:24:20.808735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T19:24:20.412783Z digest=sha256:51d746b33755e640e2102c8f3ad8bec405bae8f0d8f6cb50cfe590d01f56831e

Observation 7c6af101-8642-46d1-96d2-f8c22e40c9d9 · outbound

This paper cites Learning to localize sound source in visual scenes.

Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? Learning to localize sound source in visual scenes

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:24:20.796559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T19:24:20.416214Z digest=sha256:c2e598428d8fbfb5eee4d47eb510ae985cdafe6b39f755bb2321bc8827db4820

Observation 05e60b53-f984-4953-aa14-636976fbe689 · outbound

This paper cites Odor/taste integration and the perception of flavor.

Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? Odor/taste integration and the perception of flavor

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:24:20.783592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T19:24:20.419549Z digest=sha256:38c81c1290e01a59eb83630e97793c6d357252c57971ea6f905911cca06ad2e8

Observation 624ff189-17f9-4269-a392-e5f4fcf6ac16 · outbound

This paper cites Unveiling and Mitigating Bias in Audio Visual Segmentation.

Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? Unveiling and Mitigating Bias in Audio Visual Segmentation

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-08-09T19:24:20.564339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T19:24:20.422869Z digest=sha256:c1bbb6e991c06099d6fc4a8609cc44e35ff355c7247c70aabbf1963566e65d51

Observation c664f050-f538-472c-a8fd-9484947e3cea · outbound

This paper cites Unified mul- tisensory perception: Weakly-supervised audio-visual video parsing.

Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? Unified mul- tisensory perception: Weakly-supervised audio-visual video parsing

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:24:20.770675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T19:24:20.426645Z digest=sha256:94c0c368cd90a34807c68243c9b2241b6041f93947eb877c21027007ef434bb5

Observation 6011ac71-824c-45f6-a37b-5b3ddd781939 · outbound

This paper cites Assured Autonomy: Path Toward Living With Autonomous Systems We Can Trust.

Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? Assured Autonomy: Path Toward Living With Autonomous Systems We Can Trust

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-09T19:24:20.430094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:24:20.430094Z digest=sha256:f19ee3d7f9c2311652555c6be7c9e1c323a0c6e6203974ca53b3bca5abdc1600

Observation 6d0e6881-6005-4f12-aecb-dac4844dd551 · outbound

This paper cites AudioBench: A Universal Benchmark for Audio Large Language Models.

Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? AudioBench: A Universal Benchmark for Audio Large Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-09T19:24:20.433964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:24:20.433964Z digest=sha256:875c3f77e89116413cbe0b828e924d50f1a60e6ed31dad25b864b18224d4f93e

Observation 001c15d9-596c-4bfa-a0c0-462ff0717e40 · outbound

This paper cites What makes train- ing multi-modal classification networks hard? In Proceed- ings of the IEEE/CVF conference on computer vision and pattern recognition, pages 12695–12705, 2020.

Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? What makes train- ing multi-modal classification networks hard? In Proceed- ings of the IEEE/CVF conference on computer vision and pattern recognition, pages 12695–12705, 2020

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:24:20.757431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T19:24:20.437710Z digest=sha256:b9378dc2c4070d2028bc09d684e56403168e836d18b325dddbc71f64faa06d67

Observation e47230cf-78ed-4edf-8fbd-e72631751744 · outbound

This paper cites Pvt v2: Improved baselines with pyramid vision transformer.

Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? Pvt v2: Improved baselines with pyramid vision transformer

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:24:20.745520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T19:24:20.441068Z digest=sha256:298e8550b8a7444c962c5d1929a9f60677bdc2827e12bbac7f3b2dac0e53ad4e

Observation 8a03ae3e-458a-4034-9434-c9856878e4ac · outbound

This paper cites Prompting segmentation with sound is gen- eralizable audio-visual source localizer.

Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? Prompting segmentation with sound is gen- eralizable audio-visual source localizer

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:24:20.733937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T19:24:20.444378Z digest=sha256:786802b1ed7e93e5ef78a567f9e64d244819dff93b9d79d9bdf29067e1c36ede

Observation bd44ef52-0ab8-49ce-8156-768eb8918bc1 · outbound

This paper cites Can Textual Semantics Mitigate Sounding Object Segmentation Preference?.

Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? Can Textual Semantics Mitigate Sounding Object Segmentation Preference?

Reference 48

Resolution
verified exact
local_arxiv, observed 2026-08-09T19:24:20.529588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T19:24:20.447630Z digest=sha256:24b9775306c18fd57259747cf191bf5aa8b82d81e54f2cf289dc26acf0e4f4d9

Observation fc4dd78f-a199-4c1d-8114-ca097ae91fe3 · outbound

This paper cites Ref-AVS: Refer and Segment Objects in Audio-Visual Scenes.

Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? Ref-AVS: Refer and Segment Objects in Audio-Visual Scenes

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-08-09T19:24:20.513951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T19:24:20.450615Z digest=sha256:07bc48132722d68da1e3adb2a1c127c396e3cc20dd9ff382d0559a69b6ace2fd

Observation 5f904dfc-2706-4170-9a55-646f8d6fc03e · outbound

This paper cites MMPareto: Boosting Multimodal Learning with Innocent Unimodal Assistance.

Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? MMPareto: Boosting Multimodal Learning with Innocent Unimodal Assistance

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-09T19:24:20.453394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:24:20.453394Z digest=sha256:c21c91e36f4a71c1e22b17ea319303907d573978a65e6f9f216627dc58dadf4d

Observation 9f9efe18-145c-4fdc-aa98-12120b0235c0 · outbound

This paper cites Analyzing audiovisual data for understanding user’s emotion in human- computer interaction environment.

Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? Analyzing audiovisual data for understanding user’s emotion in human- computer interaction environment

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:24:20.720184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T19:24:20.456185Z digest=sha256:ca4ab62d42d186dee83ee592a6f973f791d6edb5fa87b8e447d5137cc5d6ac49

Observation a4c5a844-0fa7-4fae-ba4a-a4ce09f3492d · outbound

This paper cites Cooperation does matter: Exploring multi-order bilateral relations for audio- visual segmentation.

Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? Cooperation does matter: Exploring multi-order bilateral relations for audio- visual segmentation

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:24:20.707155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T19:24:20.459070Z digest=sha256:1be7a421068462fd12ae2d5110b2129e7a3f88136839ed30b911619457fb5d97

Observation f416126d-e8c2-43e7-9803-d90922dd7128 · outbound

This paper cites Audio–visual segmentation.

Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? Audio–visual segmentation

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-09T19:24:20.461698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:24:20.461698Z digest=sha256:3593e5a7b916d028a2bc0451f903754838914cde855ccf7a9d9059099f037bfe

Observation 6628f8ac-7e2f-4c0e-9575-3549b0176a7c · outbound

This paper cites 1, 2, 3, 4, 6, 7, 8, 14 Supplementary Material A.

Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? 1, 2, 3, 4, 6, 7, 8, 14 Supplementary Material A

Reference 403

Resolution
malformed identifier
raw_fallback, observed 2026-08-09T19:24:20.684763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T19:24:20.464475Z digest=sha256:dc0030285c3f61934fe9c0a8f7143663cebf82d7dd325ac66a9f99940e4a5979

Pith citing papers

Observation 5ac2cb49-56c7-419a-97c4-702c5577b400 · inbound

Learning from Silence and Noise for Visual Sound Source Localization cites this paper.

Learning from Silence and Noise for Visual Sound Source Localization Do Audio-Visual Segmentation Models Truly Segment Sounding Objects?

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-05T14:03:19.216337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:03:19.216337Z digest=sha256:3c7c06dfff2577136eb0aaf7ca232463af51455c2504d7d33a630f4548b2e0d7

Observation a7d7f17a-5f35-478f-9cd9-388ff19e134c · inbound

AVI-Bench: Toward Human-like Audio-Visual Intelligence of Omni-MLLMs cites this paper.

AVI-Bench: Toward Human-like Audio-Visual Intelligence of Omni-MLLMs Do Audio-Visual Segmentation Models Truly Segment Sounding Objects?

Reference 79

Resolution
verified exact
arxiv_id, observed 2026-07-01T23:06:21.323631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-28T14:36:53.295540Z digest=sha256:b303fb641de0fd7bd531fa51df05309e93e5e747b04b6f31d92c9ebfbaf84a9d