Pith. sign in

Paper Citation Record · LEDGER

UWAV: Uncertainty-weighted Weakly-supervised Audio-Visual Video Parsing

As of 23 August 2026, this Paper Citation Record lists 63 of 63 outbound references and 0 inbound Pith citation observations for arXiv:2505.09615.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.09615 v1

Coverage vector

measured 63 of 63 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T21:31:59.807964Z

measured 63 of 63 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

63 of 63 outbound references displayed

  • verified exact0
  • verified fuzzy53
  • unresolved10
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9ec27250-d849-4af4-a668-d5a767e58f35 · outbound

This paper cites Layer Normalization.

UWAV: Uncertainty-weighted Weakly-supervised Audio-Visual Video Parsing Layer Normalization

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T21:31:59.133256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:31:59.133256Z digest=sha256:c2ff1f3cd684482af91135ab35048b298151de7cbec4cb874e058141f0d9c7e4

Observation 122a23f9-a735-48bd-8eac-98d5c7a863a1 · outbound

This paper cites Visual scene graphs for audio source separation.

UWAV: Uncertainty-weighted Weakly-supervised Audio-Visual Video Parsing Visual scene graphs for audio source separation

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:32:02.173760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T21:31:59.143601Z digest=sha256:75e31ce9d1844af2f82c2234b262989bd4a8a5afa608c7c3b0031b859f9f2c77

Observation 3a04554b-f6b1-4db9-a412-a73dafb7a1c2 · outbound

This paper cites Learning audio-visual dynamics using scene graphs for au- dio source separation.

UWAV: Uncertainty-weighted Weakly-supervised Audio-Visual Video Parsing Learning audio-visual dynamics using scene graphs for au- dio source separation

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:32:02.132139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T21:31:59.152983Z digest=sha256:2298b5cae143dcdf506d416a103778665b0b4bb7aa4cc83d632530255b40f4c4

Observation 01bcbeb9-a2ad-4f3f-b95a-0c16ff7ea937 · outbound

This paper cites Soundspaces: Audio-visual navigation in 3d environments.

UWAV: Uncertainty-weighted Weakly-supervised Audio-Visual Video Parsing Soundspaces: Audio-visual navigation in 3d environments

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:32:02.069914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T21:31:59.172888Z digest=sha256:e591579f5434db867e64412e4f8542ef5f233249ebd7a2cd0f63947838a49e51

Observation 75d4e832-1bc9-45b7-a475-ef6cc9f86031 · outbound

This paper cites Se- mantic audio-visual navigation.

UWAV: Uncertainty-weighted Weakly-supervised Audio-Visual Video Parsing Se- mantic audio-visual navigation

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:32:02.035096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T21:31:59.183218Z digest=sha256:1b5c8b44907f061bd6d196d79e599bb17bc6a928b8d1aee92f9da88042e35abf

Observation 7e00efec-21d4-4918-abb5-569a616b0f46 · outbound

This paper cites Learning to set waypoints for audio-visual navigation.

UWAV: Uncertainty-weighted Weakly-supervised Audio-Visual Video Parsing Learning to set waypoints for audio-visual navigation

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:32:02.008540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T21:31:59.195133Z digest=sha256:ae02bdfe2a7da401f4e0cb1b81a87dd01135a8231b6fa44393a3f51945f4457d

Observation 3bcb7c00-8c83-4b96-875d-2ff7baa9eb2c · outbound

This paper cites iquery: Instruments as queries for audio-visual sound separation.

UWAV: Uncertainty-weighted Weakly-supervised Audio-Visual Video Parsing iquery: Instruments as queries for audio-visual sound separation

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:32:01.964867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T21:31:59.200164Z digest=sha256:967b26e95eb48635d1373976dbc1d2ec477f877744cb0605fbc1de54bc7bb791

Observation 9af12bc2-4b01-4410-9bb8-425b9aa44a7e · outbound

This paper cites Joint-modal label denoising for weakly-supervised audio-visual video parsing.

UWAV: Uncertainty-weighted Weakly-supervised Audio-Visual Video Parsing Joint-modal label denoising for weakly-supervised audio-visual video parsing

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:32:01.932493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T21:31:59.206004Z digest=sha256:cdabf37e32d4ce3a9d0190a9f0724ba1613ce77fdf0d42eb47d52858fb75a304

Observation 49e3a0b9-753b-4b06-aa0b-60647b3887d2 · outbound

This paper cites Imagenet: A large-scale hierarchical image database.

UWAV: Uncertainty-weighted Weakly-supervised Audio-Visual Video Parsing Imagenet: A large-scale hierarchical image database

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T21:31:59.212123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:31:59.212123Z digest=sha256:c721d8cba4d16c9675869aebe962c98f4b55bc58c4a2d203ce484e71256a2fce

Observation 090619de-dd9f-46e9-b03a-c30d137e603f · outbound

This paper cites Revisit weakly-supervised audio-visual video parsing from the lan- guage perspective.

UWAV: Uncertainty-weighted Weakly-supervised Audio-Visual Video Parsing Revisit weakly-supervised audio-visual video parsing from the lan- guage perspective

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:32:01.858907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T21:31:59.220812Z digest=sha256:992e36ecf110de68116cf86f92150c08615bcd52c6dd1c2486d05da3dcd77c37

Observation c08ac07f-28e2-4ca2-a3ee-222dd4a14819 · outbound

This paper cites Col- lecting cross-modal presence-absence evidence for weakly- supervised audio-visual event perception.

UWAV: Uncertainty-weighted Weakly-supervised Audio-Visual Video Parsing Col- lecting cross-modal presence-absence evidence for weakly- supervised audio-visual event perception

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:32:01.831909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T21:31:59.229412Z digest=sha256:3de168ab41c1ea9a3278e643600d6576bba7a759da4025c4d66734443d899883

Observation fe254114-4723-468c-a3f6-94a658a3de52 · outbound

This paper cites Audio set: An ontology and human- labeled dataset for audio events.

UWAV: Uncertainty-weighted Weakly-supervised Audio-Visual Video Parsing Audio set: An ontology and human- labeled dataset for audio events

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:32:01.778052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T21:31:59.243269Z digest=sha256:3640244b4e052352f402412c46ec3751bb99a0a27e27b6d9e8ae86fa4fec40c2

Observation 2f7b1528-1a7e-48e4-b24f-16e9d5a583dc · outbound

This paper cites Dynamic graph representation learning for video dialog via multi-modal shuffled transformers.

UWAV: Uncertainty-weighted Weakly-supervised Audio-Visual Video Parsing Dynamic graph representation learning for video dialog via multi-modal shuffled transformers

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:32:01.736573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T21:31:59.259637Z digest=sha256:183d8a6aea84c2640dd6b3789ded2b43b91485f8fea6a226091278c05c83a1f0

Observation fe142bdd-512d-4679-8fc4-c6c61e38b418 · outbound

This paper cites Dense-localizing audio-visual events in untrimmed videos: A large-scale benchmark and baseline.

UWAV: Uncertainty-weighted Weakly-supervised Audio-Visual Video Parsing Dense-localizing audio-visual events in untrimmed videos: A large-scale benchmark and baseline

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:32:01.717286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T21:31:59.277061Z digest=sha256:9177e338a0bc8ee2f578e84651d6b0022a90c5a5aed2775e2d3510dc5eb2a0de

Observation 68fce9a2-ea9a-41e2-b3e4-1bea7f5d9aa1 · outbound

This paper cites Deep residual learning for image recognition.

UWAV: Uncertainty-weighted Weakly-supervised Audio-Visual Video Parsing Deep residual learning for image recognition

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T21:31:59.291432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:31:59.291432Z digest=sha256:a824fd88c484f4930ec672d6b257f0179a88b29c3d7c09c0d72f3452692a72f7

Observation bb944c0c-1732-4e18-aa01-4b392bd883a6 · outbound

This paper cites Cnn ar- chitectures for large-scale audio classification.

UWAV: Uncertainty-weighted Weakly-supervised Audio-Visual Video Parsing Cnn ar- chitectures for large-scale audio classification

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T21:31:59.298919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:31:59.298919Z digest=sha256:f7dd038b9f06e157ea2744769ec77282ce5c676bd043b654cdc10ace2ca3f2a2

Observation d34a55a7-ae8d-47d5-a862-11ddb19ae7c7 · outbound

This paper cites Mix and local- ize: Localizing sound sources in mixtures.

UWAV: Uncertainty-weighted Weakly-supervised Audio-Visual Video Parsing Mix and local- ize: Localizing sound sources in mixtures

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:32:01.568671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T21:31:59.305654Z digest=sha256:5a01f2e9780be3093e4f74b6bdbd4eef1bafb3da45c80837c32e10152999c8e0

Observation 5bece086-1497-43f5-be45-5bcf098d5c53 · outbound

This paper cites Egocentric audio-visual object localization.

UWAV: Uncertainty-weighted Weakly-supervised Audio-Visual Video Parsing Egocentric audio-visual object localization

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:32:01.534997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T21:31:59.315804Z digest=sha256:4c2fc14e7f844eeda2058331100ee5c9c6590bbd26ef4824c634ee59afb61d8b

Observation b951281e-da2c-40a5-8611-83f72d791519 · outbound

This paper cites The Kinetics Human Action Video Dataset.

UWAV: Uncertainty-weighted Weakly-supervised Audio-Visual Video Parsing The Kinetics Human Action Video Dataset

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T21:31:59.332086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:31:59.332086Z digest=sha256:536f1d23b5d330934c8dc6781be5931d23279e3680d0d6f30f622219f6ad78ba

Observation c27e4f5b-d117-476e-98cb-d2253818e872 · outbound

This paper cites Modality-independent teachers meet weakly-supervised audio-visual event parser.

UWAV: Uncertainty-weighted Weakly-supervised Audio-Visual Video Parsing Modality-independent teachers meet weakly-supervised audio-visual event parser

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:32:01.500340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T21:31:59.342042Z digest=sha256:a125ee61f0ac9e4b2afbd68698364869a70b32f3ba30e1400c78f18398d42913

Observation 6264a68f-0e20-4446-8754-edf576e16446 · outbound

This paper cites Learning to answer questions in dy- namic audio-visual scenarios.

UWAV: Uncertainty-weighted Weakly-supervised Audio-Visual Video Parsing Learning to answer questions in dy- namic audio-visual scenarios

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:32:01.421408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T21:31:59.347030Z digest=sha256:3b0a1162c2dceef4b6ea2980330d7d47dde8c68d84be8ac604eb288cfe749a29

Observation 163b6095-6b3d-4375-b49d-a4bf3743196d · outbound

This paper cites Audio-visual segmentation via unlabeled frame exploitation.

UWAV: Uncertainty-weighted Weakly-supervised Audio-Visual Video Parsing Audio-visual segmentation via unlabeled frame exploitation

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:32:01.377681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T21:31:59.363415Z digest=sha256:97f56eb5f6d63e46a0b1815c83f507827327ddcaa0a74c1d672627d9befc083b

Observation 1fa7b556-f9cf-4a63-867a-006f66a06aba · outbound

This paper cites Decoupled weight decay regularization.

UWAV: Uncertainty-weighted Weakly-supervised Audio-Visual Video Parsing Decoupled weight decay regularization

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:32:01.342048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T21:31:59.368589Z digest=sha256:26f28b495fcc6fa50a0fbd11f7cbe0ffc45031c8a1af897ed4476426fc032102

Observation b152b6d9-26e4-4c29-95ac-78f5b4b006da · outbound

This paper cites Move2hear: Active audio-visual source separation.

UWAV: Uncertainty-weighted Weakly-supervised Audio-Visual Video Parsing Move2hear: Active audio-visual source separation

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:32:01.309656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T21:31:59.372939Z digest=sha256:2d1eb2846ce675112ae6d733e79662b3cede13fc753e4894104c034eb43bb1ba

Observation 6106a30e-c430-4c17-bd60-6fc1420bb1df · outbound

This paper cites Multimodal variational auto-encoder based audio-visual segmentation.

UWAV: Uncertainty-weighted Weakly-supervised Audio-Visual Video Parsing Multimodal variational auto-encoder based audio-visual segmentation

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:32:01.284809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T21:31:59.380749Z digest=sha256:8ba1fcc0296ad075622100f0bc6168b1a387a6c7354c1285c7c70ef62972eecb

Observation 71eb37a7-0b4c-4dd2-9082-fb9ffd963bd8 · outbound

This paper cites Multi-modal grouping net- work for weakly-supervised audio-visual video parsing.

UWAV: Uncertainty-weighted Weakly-supervised Audio-Visual Video Parsing Multi-modal grouping net- work for weakly-supervised audio-visual video parsing

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:32:01.252579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T21:31:59.385132Z digest=sha256:123be97e950e9ecadb08ce1298a9732827f0f4b4a155de5eb730b51e42f256b8

Observation cdf0f1a2-e6b9-4436-a190-687c4a493698 · outbound

This paper cites Audio-visual grouping net- work for sound localization from mixtures.

UWAV: Uncertainty-weighted Weakly-supervised Audio-Visual Video Parsing Audio-visual grouping net- work for sound localization from mixtures

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:32:01.226954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T21:31:59.389258Z digest=sha256:2a1fd617bb2652b9d1f957b31095e3ee5a166a74fadc4cc37f3172087e863657

Observation 3c7e0dc6-c86d-4907-b81f-5ded0880a5c6 · outbound

This paper cites Beyond mono to binaural: Generating binaural au- dio from mono audio with depth and cross modal attention.

UWAV: Uncertainty-weighted Weakly-supervised Audio-Visual Video Parsing Beyond mono to binaural: Generating binaural au- dio from mono audio with depth and cross modal attention

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:32:01.184770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T21:31:59.396275Z digest=sha256:2fef5fafadac5dd8bc3bb3a27addc09fc8ee1de38f6b0747103fe854d59b387f

Observation 8dbc4278-1faa-41cd-a5c2-82859a0ea444 · outbound

This paper cites Weakly-supervised audio-visual video parsing with prototype-based pseudo-labeling.

UWAV: Uncertainty-weighted Weakly-supervised Audio-Visual Video Parsing Weakly-supervised audio-visual video parsing with prototype-based pseudo-labeling

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:32:01.147245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T21:31:59.429483Z digest=sha256:2aaee69bde94c4c19a5c958b9d915cd7ddab60cb67776fad2ea1d1aba4947244

Observation 761423ff-a33b-4083-a772-9c4c2b002f52 · outbound

This paper cites Boosting positive seg- ments for weakly-supervised audio-visual video parsing.

UWAV: Uncertainty-weighted Weakly-supervised Audio-Visual Video Parsing Boosting positive seg- ments for weakly-supervised audio-visual video parsing

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:32:01.090896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T21:31:59.434730Z digest=sha256:69587f008b990ba5bf448bcbf1f352a1a204b3f40c798d28f73e02a20fe290ce

Observation 5cf628b7-382a-44c6-87fb-bd9e24f24e21 · outbound

This paper cites Learn- ing transferable visual models from natural language super- vision.

UWAV: Uncertainty-weighted Weakly-supervised Audio-Visual Video Parsing Learn- ing transferable visual models from natural language super- vision

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:32:01.033205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T21:31:59.439188Z digest=sha256:26f817533bd4f9f20ecbd216aeca0be5a6ebf2df858ecfda80630351c34e49d0

Observation a13ecf75-1ec7-488a-a428-eb6e560ec6ff · outbound

This paper cites Dual perspective network for audio-visual event localization.

UWAV: Uncertainty-weighted Weakly-supervised Audio-Visual Video Parsing Dual perspective network for audio-visual event localization

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:32:01.001001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T21:31:59.468592Z digest=sha256:18971181611e4a91466601f195414078e6ccf16993289a35cd89c816e582020d

Observation 67a44cfd-909c-4cb8-a67f-39f93bb07ad7 · outbound

This paper cites Mm-diffusion: Learning multi-modal diffusion models for joint audio and video generation.

UWAV: Uncertainty-weighted Weakly-supervised Audio-Visual Video Parsing Mm-diffusion: Learning multi-modal diffusion models for joint audio and video generation

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:32:00.958757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T21:31:59.475114Z digest=sha256:018855abb1322eaebf864642a3d452ed51363e48ef53d4063a0cac0ae08b14d4

Observation bc626cb0-5c41-4f34-8a6b-c9f22cfa6885 · outbound

This paper cites Coleaf: A contrastive-collaborative learning framework for weakly supervised audio-visual video pars- ing.

UWAV: Uncertainty-weighted Weakly-supervised Audio-Visual Video Parsing Coleaf: A contrastive-collaborative learning framework for weakly supervised audio-visual video pars- ing

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:32:00.921513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T21:31:59.496529Z digest=sha256:618c81e2ee616f1fd8d51ba1c3be6f4e6e1443d118615aa5a4f5485b0d89875c

Observation b66d66a1-83f6-4cfb-95c9-a4988a96b0a8 · outbound

This paper cites Sound source local- ization is all about cross-modal alignment.

UWAV: Uncertainty-weighted Weakly-supervised Audio-Visual Video Parsing Sound source local- ization is all about cross-modal alignment

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:32:00.882516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T21:31:59.510634Z digest=sha256:a7385b8e24da5ae1b20a4994be58affef24b5bbfc467b5ff334752d35406b03c

Observation 39069cb5-3202-4f52-b530-63d8a6e60d8c · outbound

This paper cites Audio-visual scene-aware dialog and reasoning using audio- visual transformers with joint student-teacher learning.

UWAV: Uncertainty-weighted Weakly-supervised Audio-Visual Video Parsing Audio-visual scene-aware dialog and reasoning using audio- visual transformers with joint student-teacher learning

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:32:00.844801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T21:31:59.517011Z digest=sha256:175ba7bc3dc5553e089ceb138a5b0fd89bb96add6ec60ab7c097acdfd7c80b5e

Observation 33bc6e40-3c3b-4999-bebb-2fcb21874f9a · outbound

This paper cites Audio-visual event localization in unconstrained videos.

UWAV: Uncertainty-weighted Weakly-supervised Audio-Visual Video Parsing Audio-visual event localization in unconstrained videos

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:32:00.782620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T21:31:59.544863Z digest=sha256:5879929f706b99c638f5310f8d0ebd735b6647352a01290a67a17e2c018893cc

Observation d555cb39-b25b-4639-8748-664d0ed50552 · outbound

This paper cites Unified mul- tisensory perception: Weakly-supervised audio-visual video parsing.

UWAV: Uncertainty-weighted Weakly-supervised Audio-Visual Video Parsing Unified mul- tisensory perception: Weakly-supervised audio-visual video parsing

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:32:00.739320Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T21:31:59.549590Z digest=sha256:faee18891afa6fd440de055e3bba15aa0cf568d520054804566eea9bb00a8e7a

Observation 196cbc9c-f071-4256-ad58-fb113cd3019d · outbound

This paper cites Cyclic co-learning of sounding object visual grounding and sound separation.

UWAV: Uncertainty-weighted Weakly-supervised Audio-Visual Video Parsing Cyclic co-learning of sounding object visual grounding and sound separation

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:32:00.704073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T21:31:59.554817Z digest=sha256:11e90238a52d7a861dfb31d8144c6d30291a4ada5eb2f85bad2e30e53725b986

Observation b3f7fa74-d24f-4a39-9e40-e07ca16d0b9d · outbound

This paper cites A closer look at spatiotemporal convolutions for action recognition.

UWAV: Uncertainty-weighted Weakly-supervised Audio-Visual Video Parsing A closer look at spatiotemporal convolutions for action recognition

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:32:00.649971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T21:31:59.560734Z digest=sha256:4c406702ba5f62fb1aa2e3fca368789fed55e9b508640a84dbc9cb72c1aec9ec

Observation c4e3e517-544c-4018-aa24-8cbcd43df6e2 · outbound

This paper cites Attention is all you need.

UWAV: Uncertainty-weighted Weakly-supervised Audio-Visual Video Parsing Attention is all you need

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:32:00.626134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T21:31:59.570161Z digest=sha256:04e07761543846237cc2dfc253953c2d330c33d763913086880aa8f85bb4ff4f

Observation edaf838b-100d-4cdd-8cbc-0d1834ba3ea9 · outbound

This paper cites Exploring heterogeneous clues for weakly-supervised audio-visual video parsing.

UWAV: Uncertainty-weighted Weakly-supervised Audio-Visual Video Parsing Exploring heterogeneous clues for weakly-supervised audio-visual video parsing

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:32:00.607326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T21:31:59.575253Z digest=sha256:3ba5ebd5bd100e8261195c0221bbafd981b865b36f49199325eee3abde7b1545

Observation 45bb1053-d8b5-436f-9616-617208fbd3d7 · outbound

This paper cites Large-scale con- trastive language-audio pretraining with feature fusion and keyword-to-caption augmentation.

UWAV: Uncertainty-weighted Weakly-supervised Audio-Visual Video Parsing Large-scale con- trastive language-audio pretraining with feature fusion and keyword-to-caption augmentation

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:32:00.576855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T21:31:59.583316Z digest=sha256:e8ab9eda8e10207e289ab0f6693f1d43a0c175be4fe0f4102e76fdd0630a2477

Observation 1bf0d9e2-95d8-423a-a3c6-d352f6cd8879 · outbound

This paper cites Seeing and hearing: Open-domain visual- audio generation with diffusion latent aligners.

UWAV: Uncertainty-weighted Weakly-supervised Audio-Visual Video Parsing Seeing and hearing: Open-domain visual- audio generation with diffusion latent aligners

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-15T21:31:59.598988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:31:59.598988Z digest=sha256:f01a18e7e61fbc324d1fb5d5fe686253ab288f43d1e661dffea2b8ef96b73807

Observation 82134964-1dfc-4804-9934-d0845d4bd6a6 · outbound

This paper cites Cooperation does matter: Exploring multi-order bilateral relations for audio- visual segmentation.

UWAV: Uncertainty-weighted Weakly-supervised Audio-Visual Video Parsing Cooperation does matter: Exploring multi-order bilateral relations for audio- visual segmentation

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:32:00.542247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T21:31:59.646291Z digest=sha256:8e304f9e3f861097eb239bc4ccf4696bf1e8b3aca232023d7a56b4bc8e884adb

Observation 0100c19c-7e21-403e-8631-718a5bdea093 · outbound

This paper cites Lavss: Location-guided audio-visual spatial audio separation.

UWAV: Uncertainty-weighted Weakly-supervised Audio-Visual Video Parsing Lavss: Location-guided audio-visual spatial audio separation

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:32:00.519802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T21:31:59.659106Z digest=sha256:ba1b5d4bf9a2054739420dcb9e41add5761161753f6844c1974c38d67017007c

Observation 5a19288c-a658-4258-bb17-e851198d851f · outbound

This paper cites Catch me if you hear me: Audio-visual navigation in complex unmapped environments with moving sounds.

UWAV: Uncertainty-weighted Weakly-supervised Audio-Visual Video Parsing Catch me if you hear me: Audio-visual navigation in complex unmapped environments with moving sounds

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:32:00.481224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T21:31:59.665582Z digest=sha256:446ee5a7456532f3cdafb7faca8c79116a32f3e3840274d39922936e829a296a

Observation 48f0b0b8-d54c-45cc-847f-6fc3a3f72bdc · outbound

This paper cites Mm-pyramid: Multimodal pyramid attentional network for audio-visual event localization and video pars- ing.

UWAV: Uncertainty-weighted Weakly-supervised Audio-Visual Video Parsing Mm-pyramid: Multimodal pyramid attentional network for audio-visual event localization and video pars- ing

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:32:00.458147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T21:31:59.680361Z digest=sha256:ac9782f6a71790e5e99ca96bbdbbd1771829eb266e6170ac6183253074e644d1

Observation 00510dc9-60a4-4deb-9779-4177ba41a438 · outbound

This paper cites Sound adversarial audio- visual navigation.

UWAV: Uncertainty-weighted Weakly-supervised Audio-Visual Video Parsing Sound adversarial audio- visual navigation

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:32:00.439982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T21:31:59.687456Z digest=sha256:e1d44f92d9de74ff98d04a813e4ee79691a0321531e72fb6ebc5466482ed0aa1

Observation b6efe3de-7ec8-49d2-ae8f-a1293f41bc3a · outbound

This paper cites Pano-avqa: Grounded audio-visual question answering on 360deg videos.

UWAV: Uncertainty-weighted Weakly-supervised Audio-Visual Video Parsing Pano-avqa: Grounded audio-visual question answering on 360deg videos

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:32:00.409591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T21:31:59.694973Z digest=sha256:c01e467e8735e649e142cc47e261657a7eb60d4b9392054cb13f81a71871bf2b

Observation a2b0b969-271b-4913-a179-57b8c7571bc8 · outbound

This paper cites Positive sample propagation along the audio- visual event line.

UWAV: Uncertainty-weighted Weakly-supervised Audio-Visual Video Parsing Positive sample propagation along the audio- visual event line

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:32:00.376580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T21:31:59.701281Z digest=sha256:af08b47ebd5d0830be62292814f4dd1f7afa5e3731ac48d30e0f6f559afab99e

Observation 94f99248-d5bc-4395-8011-5d55daa45612 · outbound

This paper cites Audio–visual segmentation.

UWAV: Uncertainty-weighted Weakly-supervised Audio-Visual Video Parsing Audio–visual segmentation

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:32:00.348668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T21:31:59.716403Z digest=sha256:5b412d0ce9bff32a2ee32128536ae32ec6d85087055b21d6a21f1e058aaac69c

Observation f87879e6-6f2a-495e-8bd2-0b465a6438e2 · outbound

This paper cites Improving Audio-Visual Video Parsing with Pseudo Visual Labels.

UWAV: Uncertainty-weighted Weakly-supervised Audio-Visual Video Parsing Improving Audio-Visual Video Parsing with Pseudo Visual Labels

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-15T21:31:59.723469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:31:59.723469Z digest=sha256:a6267673f2c4d22611b364402d0c160f67ccb5deeefe370d6cecaf3f9a13fc47

Observation 72850119-9bc4-4d97-87de-4c4e324e65c4 · outbound

This paper cites Label-anticipated Event Disentanglement for Audio-Visual Video Parsing.

UWAV: Uncertainty-weighted Weakly-supervised Audio-Visual Video Parsing Label-anticipated Event Disentanglement for Audio-Visual Video Parsing

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-15T21:31:59.735886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:31:59.735886Z digest=sha256:f64fca64b4ca0ef6b185c70bf0d8f7b85c8452e3a1c6268986738c46a9631d47

Observation 8a537a66-5792-44a6-ab2a-254b00464927 · outbound

This paper cites an unresolved cited work.

UWAV: Uncertainty-weighted Weakly-supervised Audio-Visual Video Parsing Unresolved cited work

Reference 55

Resolution
unresolved
raw_fallback, observed 2026-08-15T21:32:00.327917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T21:31:59.743569Z digest=sha256:3c8b1062cd0724e3992f2dc39f1bd58560d6f6938c8d3e9e5ef6c03a9c591fbc

Observation 4b3a8ba1-756b-41e5-a620-bff3dcf9668a · outbound

This paper cites Infer- ence Time.

UWAV: Uncertainty-weighted Weakly-supervised Audio-Visual Video Parsing Infer- ence Time

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:32:00.298247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T21:31:59.755392Z digest=sha256:19d4e209604e4f10bdaeb6f4a8d6d407c16a7ebe00adca570678bdbb5fa1650f

Observation 67290ba5-3189-4ca7-b194-62dfb489fb92 · outbound

This paper cites an unresolved cited work.

UWAV: Uncertainty-weighted Weakly-supervised Audio-Visual Video Parsing Unresolved cited work

Reference 57

Resolution
unresolved
raw_fallback, observed 2026-08-15T21:32:00.271922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T21:31:59.763712Z digest=sha256:729bafd96ca1d9868734a4ae8cbdbe30f3b7693a31479827a55aefccabf48a0c

Observation 9c48b870-ba84-4ea8-8eec-045753a1093c · outbound

This paper cites CoLeaf [34] on the LLP dataset [38].

UWAV: Uncertainty-weighted Weakly-supervised Audio-Visual Video Parsing CoLeaf [34] on the LLP dataset [38]

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:32:00.250972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T21:31:59.771495Z digest=sha256:d439fc6ff39d1ab91c6155dc543bebc58b01712efff4aa4b58556867fce857e1

Observation f81f2b27-f1a3-4e5c-b61e-95abfd2fbe03 · outbound

This paper cites CLIP and CLAP as visual and audio feature backbones.

UWAV: Uncertainty-weighted Weakly-supervised Audio-Visual Video Parsing CLIP and CLAP as visual and audio feature backbones

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:32:00.227161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T21:31:59.777320Z digest=sha256:baa9ff5f89c0dea852efae3efa23093c606e9963e8f86ec7fe87f868ebf1af2e

Observation 59f87c5f-b8a7-42f9-9ca2-1d80021ad4b5 · outbound

This paper cites When α is adjusted, class-balanced loss re-weighting is not applied.

UWAV: Uncertainty-weighted Weakly-supervised Audio-Visual Video Parsing When α is adjusted, class-balanced loss re-weighting is not applied

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:32:00.204557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T21:31:59.784903Z digest=sha256:f362a41dce7dccd4a2d4dcaed0af5f9668d326a78b333f8390a6724f748c5861

Observation c9fabd4b-fd37-4de4-a131-cdec963d5ac6 · outbound

This paper cites The scalability of UW A V.

UWAV: Uncertainty-weighted Weakly-supervised Audio-Visual Video Parsing The scalability of UW A V

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:32:00.153559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T21:31:59.795193Z digest=sha256:8cf517ed7926ae045853b27dcafe7f48db895cad3aeec26b9b6abd0000af2e85

Observation 02239d69-91e6-41d2-a4ce-e5836b830be4 · outbound

This paper cites Binary” denotes training with binary pseudo-labels. “Soft.

UWAV: Uncertainty-weighted Weakly-supervised Audio-Visual Video Parsing Binary” denotes training with binary pseudo-labels. “Soft

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:32:00.128379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T21:31:59.801428Z digest=sha256:e709342f485c8825d1d5768d6167f7bdca9d47c2e1fd0160a6e49b7e58f1f7cc

Observation d1d2606a-24bb-491e-ad49-8dfcaff1a07a · outbound

This paper cites Figure A8 shows the same, for sample videos on the A VE dataset [37].

UWAV: Uncertainty-weighted Weakly-supervised Audio-Visual Video Parsing Figure A8 shows the same, for sample videos on the A VE dataset [37]

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:32:00.074340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T21:31:59.807964Z digest=sha256:783b2cca499667f0617b05c956edfbd318c3c8bd910b7370e4602a1df3fc406a

Pith citing papers

No inbound Pith citation observations are available.