Pith. sign in

Paper Citation Record · LEDGER

LINK: Adaptive Modality Interaction for Audio-Visual Video Parsing

As of 20 August 2026, this Paper Citation Record lists 25 of 25 outbound references and 1 inbound Pith citation observation for arXiv:2412.20872.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.20872 v2

Coverage vector

measured 25 of 25 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T23:14:31.426091Z

measured 26 of 26 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T01:07:11.142298Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-16T01:07:11.189974Z

Reference resolution

25 of 25 outbound references displayed

  • verified exact2
  • verified fuzzy23
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 027ffa93-b264-4561-9501-c1dbc7f1e54b · outbound

This paper cites Audio- visual event localization in unconstrained videos,.

LINK: Adaptive Modality Interaction for Audio-Visual Video Parsing Audio- visual event localization in unconstrained videos,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:14:31.960670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T23:14:31.338358Z digest=sha256:bf23d3df0f436c249363c8b8145ae4c62b5c0b3c4980449a824246cd2452dc7f

Observation cb88e4eb-532b-4cd9-bbb8-27a098845744 · outbound

This paper cites Dual attention matching for audio-visual event localization,.

LINK: Adaptive Modality Interaction for Audio-Visual Video Parsing Dual attention matching for audio-visual event localization,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:14:31.950588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T23:14:31.342825Z digest=sha256:0b41e327ecb60aa4eef4982918c8a6ddfd6b8e24791263c0bfbbdddabdd6898f

Observation a80ea30b-4a9b-47e5-bc61-38e74334137d · outbound

This paper cites Learning to answer questions in dynamic audio-visual scenarios,.

LINK: Adaptive Modality Interaction for Audio-Visual Video Parsing Learning to answer questions in dynamic audio-visual scenarios,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:14:31.940052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T23:14:31.346742Z digest=sha256:807915966cf194548bf65cb47d076f1ef872868604d18660f9e15e0a4c1fd492

Observation 291f353f-4670-4556-b5a5-2814a3299113 · outbound

This paper cites Listen to look: Action recognition by previewing audio,.

LINK: Adaptive Modality Interaction for Audio-Visual Video Parsing Listen to look: Action recognition by previewing audio,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:14:31.927958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T23:14:31.350756Z digest=sha256:0b612d0b3ad1a798de8314c11185e8146a08735d5fc3a77b77e07444762d9377

Observation 084c3cf5-50bc-4850-9bab-73e7b19f63e7 · outbound

This paper cites Unified multisensory per- ception: Weakly-supervised audio-visual video parsing,.

LINK: Adaptive Modality Interaction for Audio-Visual Video Parsing Unified multisensory per- ception: Weakly-supervised audio-visual video parsing,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:14:31.917109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T23:14:31.354817Z digest=sha256:d99ea503afc413449d883f177bf762121a039da84d380cf479fe024fbd73b26b

Observation 78b33c2e-ab8e-4b1a-9698-0c5370b29680 · outbound

This paper cites MM-Pyramid: Multimodal pyramid attentional network for audio-visual event localization and video parsing,.

LINK: Adaptive Modality Interaction for Audio-Visual Video Parsing MM-Pyramid: Multimodal pyramid attentional network for audio-visual event localization and video parsing,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:14:31.905241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T23:14:31.358620Z digest=sha256:3ce553c2cb5dbd22bab4e942dca8760ec1c5da85116d6aa5102afe8cf811cee1

Observation 772b055e-a842-4dae-8880-c4f23906bcb5 · outbound

This paper cites Cross-modal Prompts: Adapting Large Pre-trained Models for Audio-Visual Downstream Tasks.

LINK: Adaptive Modality Interaction for Audio-Visual Video Parsing Cross-modal Prompts: Adapting Large Pre-trained Models for Audio-Visual Downstream Tasks

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-08-10T23:14:31.479730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T23:14:31.362775Z digest=sha256:8c1071239d108de112ed804d22210aca12a77715e1ac95c215102ae18a5d6e67

Observation 3678b30a-fd30-4b7a-8caf-757610b41371 · outbound

This paper cites Modality-independent teachers meet weakly-supervised audio-visual event parser.

LINK: Adaptive Modality Interaction for Audio-Visual Video Parsing Modality-independent teachers meet weakly-supervised audio-visual event parser

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:14:31.893531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T23:14:31.366623Z digest=sha256:2f995ed60f599d23a8c74cfe83a01235d660525dc54cae10a8d5400dabed394c

Observation b00ba2d6-0458-437d-baf4-b8376e7bfa00 · outbound

This paper cites Label-anticipated Event Disentanglement for Audio-Visual Video Parsing,.

LINK: Adaptive Modality Interaction for Audio-Visual Video Parsing Label-anticipated Event Disentanglement for Audio-Visual Video Parsing,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:14:31.878816Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T23:14:31.369985Z digest=sha256:d4c2afa1752ceee3c7a0d84615ab889a1fd42d4fe06fccc1f3fb1d3d0d6e3eb0

Observation 1f1ea22e-4445-4d3c-80e6-f73bb5246b79 · outbound

This paper cites Cbam: Convo- lutional block attention module,.

LINK: Adaptive Modality Interaction for Audio-Visual Video Parsing Cbam: Convo- lutional block attention module,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:14:31.651641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T23:14:31.373333Z digest=sha256:37f830b34af59b3bc95fd1090b1175f50eeebeff2565656b7396d29c90b20c6e

Observation 02f180ef-7605-4869-b5be-14b205afe4d5 · outbound

This paper cites Towards Efficient Audio-Visual Learners via Empowering Pre-trained Vi- sion Transformers with Cross-Modal Adaptation,.

LINK: Adaptive Modality Interaction for Audio-Visual Video Parsing Towards Efficient Audio-Visual Learners via Empowering Pre-trained Vi- sion Transformers with Cross-Modal Adaptation,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:14:31.642040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T23:14:31.376298Z digest=sha256:f62ccb6483223b9f2a759e42cfd95e1768fa0cd25513b9be06f440ce1c9b66b8

Observation f874c14f-f6c7-419f-a57f-5c741bf31e3a · outbound

This paper cites Learning transferable visual models from natural language supervision,.

LINK: Adaptive Modality Interaction for Audio-Visual Video Parsing Learning transferable visual models from natural language supervision,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:14:31.630531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T23:14:31.379469Z digest=sha256:23fab0579d8034c3a265fd0953456b2c606a330e11409ac4e739d07cca17d508

Observation c57e30cc-86a3-408f-84d0-533adb9a56ae · outbound

This paper cites Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation,.

LINK: Adaptive Modality Interaction for Audio-Visual Video Parsing Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:14:31.618846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T23:14:31.382375Z digest=sha256:608e9fe636867d8c1b1bdaf2f125e444f41d18238ac8227e621637d116d24393

Observation f9db92f0-8937-4978-a6c2-e1ab5c9da90a · outbound

This paper cites Balanced multimodal learning via on-the-fly gradient modulation,.

LINK: Adaptive Modality Interaction for Audio-Visual Video Parsing Balanced multimodal learning via on-the-fly gradient modulation,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:14:31.606881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T23:14:31.385120Z digest=sha256:6388fef35fb6945e8961b390d93d7ebf12d781cf8a9f6b8a5a079b8984f57a0a

Observation 86bc2f74-c6ed-4c96-823a-881dafcffd81 · outbound

This paper cites What makes training multi-modal classification networks hard?.

LINK: Adaptive Modality Interaction for Audio-Visual Video Parsing What makes training multi-modal classification networks hard?

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:14:31.596182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T23:14:31.388073Z digest=sha256:f9e1298134b20917eec94623f98f572024183f62d24daa83a18fb44d3d84f5d1

Observation d9fe477f-bdc7-40ed-a1da-392e5ae1f55b · outbound

This paper cites Scal- ing multimodal pre-training via cross-modality gradient harmonization,.

LINK: Adaptive Modality Interaction for Audio-Visual Video Parsing Scal- ing multimodal pre-training via cross-modality gradient harmonization,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:14:31.584694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T23:14:31.391510Z digest=sha256:e1b9c7cf604ee90ab6cf43efd3d47adbfbcfb49085ccf88e8363afdf6ea18510

Observation c8b9b600-26d9-4085-b466-68c64a65c0dd · outbound

This paper cites Text-IF: Leveraging Semantic Text Guidance for Degradation- Aware and Interactive Image Fusion,.

LINK: Adaptive Modality Interaction for Audio-Visual Video Parsing Text-IF: Leveraging Semantic Text Guidance for Degradation- Aware and Interactive Image Fusion,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:14:31.573528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T23:14:31.394816Z digest=sha256:c74e000fe2db45f43afa5e4aab141c7ce1d917cc59a608b84f36466aba949505

Observation 5dcd6f78-f15d-456f-ac84-61bcdc1c61ba · outbound

This paper cites Language- driven All-in-one Adverse Weather Removal,.

LINK: Adaptive Modality Interaction for Audio-Visual Video Parsing Language- driven All-in-one Adverse Weather Removal,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:14:31.563153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T23:14:31.398392Z digest=sha256:c510068c31f6bc68278ba7391d8c34da43ed310c5d78a4ca8e30f17d862afbc8

Observation 261f4f31-11f2-498b-ac28-e0c59cc642ac · outbound

This paper cites Multi-modal grouping network for weakly-supervised audio-visual video parsing,.

LINK: Adaptive Modality Interaction for Audio-Visual Video Parsing Multi-modal grouping network for weakly-supervised audio-visual video parsing,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:14:31.552666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T23:14:31.402479Z digest=sha256:0e09da063c2a513547afa62096c3070b71cb3e95616bdabaf1e67cd56b037901

Observation c695cf29-0b0d-4c24-b307-03af97c209ed · outbound

This paper cites Joint-modal label denoising for weakly-supervised audio-visual video parsing,.

LINK: Adaptive Modality Interaction for Audio-Visual Video Parsing Joint-modal label denoising for weakly-supervised audio-visual video parsing,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:14:31.541674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T23:14:31.406002Z digest=sha256:fe16a6631239dfd9a845239e4e8ae043f540547e6998ed24704e6fd75a1eadc7

Observation d4577133-bff7-4684-a3a1-cdca17f11bec · outbound

This paper cites Collecting cross-modal presence-absence evidence for weakly-supervised audio- visual event perception,.

LINK: Adaptive Modality Interaction for Audio-Visual Video Parsing Collecting cross-modal presence-absence evidence for weakly-supervised audio- visual event perception,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:14:31.531000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T23:14:31.409922Z digest=sha256:1f6d1a03ae6e71f734bb0ef93e44dd9b9775bb084d94ed199e3d6abe12df3304

Observation efe3f719-9e97-4b0f-867f-25133ee6c53b · outbound

This paper cites CoLeaF: A Contrastive-Collaborative Learning Framework for Weakly Supervised Audio-Visual Video Parsing.

LINK: Adaptive Modality Interaction for Audio-Visual Video Parsing CoLeaF: A Contrastive-Collaborative Learning Framework for Weakly Supervised Audio-Visual Video Parsing

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-08-10T23:14:31.462911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T23:14:31.413878Z digest=sha256:d3e19d9a80f3029fba35469929e346dcd5bff762047aa75fa43534ab34c7ce3d

Observation 854a2959-0109-4ff4-bb39-9fcbbd030c5a · outbound

This paper cites CM-PIE: Cross-modal perception for interactive-enhanced audio-visual video parsing,.

LINK: Adaptive Modality Interaction for Audio-Visual Video Parsing CM-PIE: Cross-modal perception for interactive-enhanced audio-visual video parsing,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:14:31.519353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T23:14:31.417511Z digest=sha256:4071b98ac8c320e18ae4395ce958f6ee072eae4828dbd237a390e886fa896d34

Observation bd2c0d7f-8b4b-4b52-a42f-fa13dda274b7 · outbound

This paper cites V ALOR: Vision- Audio-Language Omni-Perception Pretraining Model and Dataset,.

LINK: Adaptive Modality Interaction for Audio-Visual Video Parsing V ALOR: Vision- Audio-Language Omni-Perception Pretraining Model and Dataset,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:14:31.508137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T23:14:31.421175Z digest=sha256:6394a73d69b9e292441b300415f71874736b19776ef992e4670f1bdb4db2c6b2

Observation d7cf971f-fddf-4cb3-84f1-1c7240fbf6a6 · outbound

This paper cites Multi-grained representa- tion learning for cross-modal retrieval,.

LINK: Adaptive Modality Interaction for Audio-Visual Video Parsing Multi-grained representa- tion learning for cross-modal retrieval,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:14:31.494562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T23:14:31.426091Z digest=sha256:9804a27d851d24f1ed8ab56439af9f061c96b64ff4a9d30cb2d8c1b1b0b9a380

Pith citing papers

Observation 5ece17e2-c5a1-4539-b62f-01df159f5a7a · inbound

TeMTG: Text-Enhanced Multi-Hop Temporal Graph Modeling for Audio-Visual Video Parsing cites this paper.

TeMTG: Text-Enhanced Multi-Hop Temporal Graph Modeling for Audio-Visual Video Parsing LINK: Adaptive Modality Interaction for Audio-Visual Video Parsing

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-08-16T01:07:11.196864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T01:07:11.142298Z digest=sha256:af3a3a6555bbacca59ebc323d537a5ffd0a26aa6f4ab61b1a50cab3a2f2aa8fb