Pith. sign in

Paper Citation Record · LEDGER

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding

As of 16 August 2026, this Paper Citation Record lists 64 of 64 outbound references and 1 inbound Pith citation observation for arXiv:2506.23196.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.23196 v1

Coverage vector

measured 64 of 64 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T21:53:02.484752Z

measured 65 of 65 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-12T03:03:13.496634Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-12T03:06:19.413098Z

Reference resolution

64 of 64 outbound references displayed

  • verified exact1
  • verified fuzzy49
  • unresolved14
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6adf58fb-1397-49f7-ac85-0b38c8e91849 · outbound

This paper cites Maas: Multi-modal assignation for active speaker detection.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Maas: Multi-modal assignation for active speaker detection

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:12.228540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:52:56.383564Z digest=sha256:89e8ad65023a928ebf6cf40e9855ee45af4c3a6dd0e39d766893a48757c4613a

Observation 28ad82c1-9139-4562-a667-3d761d6b933e · outbound

This paper cites Hear Me Out: Fusional Approaches for Audio Augmented Temporal Action Localization.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Hear Me Out: Fusional Approaches for Audio Augmented Temporal Action Localization

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:56.439313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:56.439313Z digest=sha256:81fc0e4bb6969b6022399a835773bb01919ac2f5b95c2a9872667558da52e3fa

Observation d47090ca-d83c-438b-ae89-a2525b2859f1 · outbound

This paper cites Activitynet: A large-scale video benchmark for human activity understanding.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Activitynet: A large-scale video benchmark for human activity understanding

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:12.031192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:52:56.521994Z digest=sha256:dcded5dfbfc36d7aac3f3c79c9b536024ae3c77eb92e3bb5f89ebfaf56d5fe10

Observation 571de3b4-4ced-492a-bbb9-51b03117b901 · outbound

This paper cites Quo vadis, action recognition? a new model and the kinetics dataset.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Quo vadis, action recognition? a new model and the kinetics dataset

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:11.821632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:52:56.624292Z digest=sha256:42356ea301320c19a52c771bd59bfa98be44776b5cc1f1a26c1e08024b7b5e10

Observation 7803e0e6-ca3a-4bb7-bcb6-0a03ba5ed688 · outbound

This paper cites Augmented transformer with adaptive graph for tem- poral action proposal generation.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Augmented transformer with adaptive graph for tem- poral action proposal generation

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:11.636257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:52:56.734272Z digest=sha256:3d0e9e7bc85469b8c18fc4bac8feae3399d545f7cc4baf1dd4f83f5a1e599470

Observation 7951d87f-64e0-4430-b5da-69ee90e7f60b · outbound

This paper cites Re- thinking the faster r-cnn architecture for temporal action localization.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Re- thinking the faster r-cnn architecture for temporal action localization

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:11.462906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:52:56.792960Z digest=sha256:ba7741fc46b3f660693b50e01ff39075ed628233d51b69cb74ba517c6de43130

Observation e9949550-2b40-4f80-abfb-904784a08b92 · outbound

This paper cites Tallformer: Temporal ac- tion localization with a long-memory transformer.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Tallformer: Temporal ac- tion localization with a long-memory transformer

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:11.244143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:52:56.866148Z digest=sha256:1fa00a5196dc9e9d80103fc9d57be484b70f4948ea5edac9eda81fc0267aa419

Observation 86994997-3545-44b3-b9fa-b88eb1c620a3 · outbound

This paper cites Yolo-world: Real-time open-vocabulary object detection.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Yolo-world: Real-time open-vocabulary object detection

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:56.942447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:56.942447Z digest=sha256:1822458794d3f867bd92634afdcb50091c7f311b2ccdd64c8663e10f184db933

Observation ad95e9b2-49f4-46ee-8d25-d3c8b3e77441 · outbound

This paper cites Rescaling egocentric vision: Collection, pipeline and chal- lenges for epic-kitchens-100.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Rescaling egocentric vision: Collection, pipeline and chal- lenges for epic-kitchens-100

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:11.029369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:52:57.080544Z digest=sha256:6e3d2adaddc4fc78912e244263e765220074dd6f6c50be47daab02b2a6f69f13

Observation 839918ca-bc76-4068-b0b0-a4420566535b · outbound

This paper cites Slowfast networks for video recognition.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Slowfast networks for video recognition

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:10.855384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:52:57.209727Z digest=sha256:e727fb478f04674583df7e2bbe6e700caab9c979cb37081869acbb7a247fdacd

Observation 4945c444-97ff-41ba-8e53-54c9f5253efe · outbound

This paper cites Multi-Resolution Audio-Visual Feature Fusion for Temporal Action Localization.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Multi-Resolution Audio-Visual Feature Fusion for Temporal Action Localization

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-08-06T21:53:02.785904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:52:57.316945Z digest=sha256:ae3b16edee5d11b9326071a2cd1f577f7ee73f3204ec3ee96b618f4839f9441f

Observation 1ff4a861-5dc7-4132-94a8-9ed8a34b6d33 · outbound

This paper cites Audio set: An ontology and human- labeled dataset for audio events.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Audio set: An ontology and human- labeled dataset for audio events

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:10.698271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:52:57.408175Z digest=sha256:8df5f9d0eae80542d98da720d9726db28ccbbc080e1ebf7d231fb39ee2be5db3

Observation c02bd514-8400-48f0-aab9-74bc46808358 · outbound

This paper cites Dense-localizing audio-visual events in untrimmed videos: A large-scale benchmark and baseline.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Dense-localizing audio-visual events in untrimmed videos: A large-scale benchmark and baseline

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:10.463129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:52:57.495631Z digest=sha256:92ffba9bcaf292e6114fd7693a1e1713b060bfc5b48ffed656d6d3ecb35ccc21

Observation 8c37d1a1-7b8a-46d5-88bd-e30a619c88e8 · outbound

This paper cites Momentum contrast for unsupervised visual rep- resentation learning.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Momentum contrast for unsupervised visual rep- resentation learning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:57.588220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:57.588220Z digest=sha256:62a72c6f4c2b3bd7551fb374d946326cc29d408fe51b241df96f0f08e46d10f0

Observation 6c7fba88-18db-429b-b253-bbc1b260e576 · outbound

This paper cites Cnn archi- tectures for large-scale audio classification.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Cnn archi- tectures for large-scale audio classification

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:10.259205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:52:57.718925Z digest=sha256:f55205e6942a91940efa4dd2e1794811a245bf9bfcd1f07ee2c550b9eed25ab1

Observation 4dda55a1-49df-46a2-b808-68de94d62b04 · outbound

This paper cites Mix and local- ize: Localizing sound sources in mixtures.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Mix and local- ize: Localizing sound sources in mixtures

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:57.797874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:57.797874Z digest=sha256:349b9d4c3d067019235d14d29340ebedadf29015a6b7ee8f6a2c6f41b1b89683

Observation 45643fd4-60c9-4124-aee3-f8be220466a2 · outbound

This paper cites in the wild.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding in the wild

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:10.051074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:52:57.881569Z digest=sha256:b7a8282f8ab0695ab80501cb92efc74adff169110886f6a3fa92042a0e57b46a

Observation 80271b71-b850-4d92-9f4c-96a1906fe8c5 · outbound

This paper cites Causal inference meets deep learning: A compre- hensive survey.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Causal inference meets deep learning: A compre- hensive survey

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:09.820174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:52:58.000025Z digest=sha256:9967ed3d4a703415c9ad658e82b3f2cd121815b48431ba8c7efa0101666d6419

Observation b31a13d5-21c4-4f59-8c87-5aed19bb0808 · outbound

This paper cites Epic-fusion: Audio-visual temporal bind- ing for egocentric action recognition.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Epic-fusion: Audio-visual temporal bind- ing for egocentric action recognition

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:09.637366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:52:58.137657Z digest=sha256:44d75a8ed827a6daa3cf3412cd80e69cd5aa37f4b14895592bcf70f1bd1ca222

Observation 174d3af3-f042-43d5-b345-337e0e8e60ae · outbound

This paper cites MERT: Acoustic Music Understanding Model with Large-Scale Self-supervised Training.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding MERT: Acoustic Music Understanding Model with Large-Scale Self-supervised Training

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:58.255447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:58.255447Z digest=sha256:32ca600b6f6d423ea9baa30b3afafc17d901c7b260d2c464b8c7008605af1edc

Observation 67fe8343-b81f-444f-9065-b2f8139842fa · outbound

This paper cites Learning salient boundary feature for anchor- free temporal action localization.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Learning salient boundary feature for anchor- free temporal action localization

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:09.492907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:52:58.345630Z digest=sha256:a0f8ac19c9d0f7c51a637b42f724250964bf1d330b942b44aea8c43c732d4d23

Observation 9f0d2f44-002d-4240-addb-f9b5143a8927 · outbound

This paper cites Bsn: Boundary sensitive network for temporal action proposal generation.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Bsn: Boundary sensitive network for temporal action proposal generation

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:09.315122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:52:58.449901Z digest=sha256:249bc9766f40ba6ec7944e90061fe45a6d6f93cc03bda7d1b71231620219301a

Observation 4f7e9a09-c5c0-4f1b-b4aa-68115365b8dc · outbound

This paper cites Bmn: Boundary-matching network for temporal action pro- posal generation.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Bmn: Boundary-matching network for temporal action pro- posal generation

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:58.542307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:58.542307Z digest=sha256:3cc4234b1765cbd1beaa2a8639b685a2110c3d864f2dc6f3005bad7549a242de

Observation c11b35a9-a764-4db3-b9e1-42bc4f26be31 · outbound

This paper cites Progressive boundary refine- ment network for temporal action detection.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Progressive boundary refine- ment network for temporal action detection

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:09.184494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:52:58.669443Z digest=sha256:5fd6dccb52df90fbf209127ae9bba69fa16b4c6275f61297e84776bfa240da56

Observation 145a2007-89c2-451d-8ebb-17466280acb4 · outbound

This paper cites Dense modality interaction network for audio-visual event localization.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Dense modality interaction network for audio-visual event localization

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:09.038197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:52:58.774936Z digest=sha256:dbabfaf8f6d3dbeb22a9ef2a017dddc6c4347c1c84a494f436c5b5873ee4459b

Observation 8e6cbd72-b2e8-469b-be94-c96ae7f4cf77 · outbound

This paper cites Multi-shot temporal event localization: a benchmark.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Multi-shot temporal event localization: a benchmark

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:08.873567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:52:58.879243Z digest=sha256:350c38a0adcf5b0f0c278b75c20f6b51076ffbec95d7a848c7ad53aa27cff85f

Observation 26109f82-ccde-42db-88d9-595505edf9da · outbound

This paper cites End-to-end temporal ac- tion detection with transformer.IEEE Transactions on Image Processing, 31:5427–5441, 2022.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding End-to-end temporal ac- tion detection with transformer.IEEE Transactions on Image Processing, 31:5427–5441, 2022

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:08.703945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:52:58.983937Z digest=sha256:05fb00ac5ff927a014eaf25507efdf6af87dac4b427ffc8e904725ebaa763194

Observation 0a9edb4e-f849-4d1f-8cdc-95ea4c662f25 · outbound

This paper cites Gaussian temporal awareness networks for action localization.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Gaussian temporal awareness networks for action localization

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:08.524721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:52:59.119360Z digest=sha256:b6a204707083349e0cf27f595d1b459c9e9f6d254239a8b98528b824c4f5f5de

Observation cbf46e29-3e88-4069-84ca-3f61652542e3 · outbound

This paper cites Proposal-free temporal action detection via global segmen- tation mask learning.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Proposal-free temporal action detection via global segmen- tation mask learning

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:08.374265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:52:59.215515Z digest=sha256:0bac7d55556ef1084c2c43c801171fd6ce536388326f4c846ef27d2e0858f1e2

Observation f4bc45dc-7480-4baa-8f38-360c9f7af30d · outbound

This paper cites Attention bottlenecks for multimodal fusion.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Attention bottlenecks for multimodal fusion

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:08.185290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:52:59.303981Z digest=sha256:ac185855489f8c34f45db6f4a05d22e7bc22501c975996698b82c23e63446a19

Observation fc53abfb-da85-4cd7-a9d3-e13ef2b7475a · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding DINOv2: Learning Robust Visual Features without Supervision

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:59.408856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:59.408856Z digest=sha256:9adb3fcdabb2ad81313c4c7d41f160740edf568cb01dff9ee118f98eb0c9fb44

Observation 57daa4b0-d0df-4a9f-90cc-335b1600a2e1 · outbound

This paper cites Audio-visual scene analysis with self-supervised multisensory features.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Audio-visual scene analysis with self-supervised multisensory features

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:08.046715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:52:59.543003Z digest=sha256:0fd1e7bbacdb41a2fd584734bf1f5d6fbe1ffb88b3301e1f31b532ae4d03f05a

Observation ccab2e5c-8bc7-4cca-a040-28fc3e3f8836 · outbound

This paper cites A review of deep learning techniques in audio event recognition (aer) applications.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding A review of deep learning techniques in audio event recognition (aer) applications

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:07.877014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:52:59.622948Z digest=sha256:5707ec92e1911b21312bf8e67814723688b7a411c511d02721e4691afbe89f6b

Observation 7cd8b596-8223-4b2a-a5e4-5003dea88cea · outbound

This paper cites Owl (observe, watch, listen): Audiovisual temporal context for localizing actions in ego- centric videos.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Owl (observe, watch, listen): Audiovisual temporal context for localizing actions in ego- centric videos

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:07.703910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:52:59.736617Z digest=sha256:b6f235cb7cf70ac0ef7466435dedbec73826f4a8f39effd96e0c03c169b197a7

Observation d160e83d-f8e2-4f1a-887e-ef62cb617543 · outbound

This paper cites You only look once: Unified, real-time object de- tection.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding You only look once: Unified, real-time object de- tection

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:59.832187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:59.832187Z digest=sha256:72ac190e3c0aea367401ca05f0505a1e3c95172ee0e605e5b9336e96a34b46d5

Observation 9b5c229e-e276-4ca3-91a4-bba32b5d3587 · outbound

This paper cites Action sensitivity learning for temporal action localization.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Action sensitivity learning for temporal action localization

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:59.922852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:59.922852Z digest=sha256:2aee2a0b46e783ed905be002421bedca75f1056654009cfc0489f928d9a0fe10

Observation e4f8db01-5425-4c93-b34b-101ce7945aec · outbound

This paper cites Temporal Action Localization with Enhanced Instant Discriminability.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Temporal Action Localization with Enhanced Instant Discriminability

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T21:53:00.048092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:53:00.048092Z digest=sha256:52d3539e3ee3f4b17332d245352e5ba9fc016913abf26995af2736688139cc0a

Observation 6f57e020-44bd-4943-8662-b77ad44fd57b · outbound

This paper cites Tridet: Temporal action detection with relative boundary modeling.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Tridet: Temporal action detection with relative boundary modeling

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:07.463989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:53:00.153296Z digest=sha256:c421b65369d87453fcaed5a4f44066216f9ac049358b0fcc8d0082eb831db024

Observation 022467a2-d6ab-4a35-9887-f852a27e4e6e · outbound

This paper cites Re- laxed transformer decoders for direct action proposal gener- ation.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Re- laxed transformer decoders for direct action proposal gener- ation

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:07.134101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:53:00.248074Z digest=sha256:b67c0197dfa5aa4c658d470294f602e9568bf866f1a4594d3fe3ce6676dda2b5

Observation 3a408326-4418-43bf-9b67-396df641f1d3 · outbound

This paper cites Audio-visual event localization in unconstrained videos.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Audio-visual event localization in unconstrained videos

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:06.791765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:53:00.340814Z digest=sha256:edff4460606ad91f01998f645184f766d692fc9bacd08ee5443397293acc18f5

Observation 0dd7ef7a-81b6-4b5e-a317-d1eabb452cf0 · outbound

This paper cites Deep learning-based action detection in untrimmed videos: A survey.IEEE Transactions on Pattern Analysis and Machine Intelligence , 45(4):4302– 4320, 2022.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Deep learning-based action detection in untrimmed videos: A survey.IEEE Transactions on Pattern Analysis and Machine Intelligence , 45(4):4302– 4320, 2022

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:06.518398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:53:00.491979Z digest=sha256:3fd5f4720c9ae7cbe29e2b8a4d05336f7cac29c6ce82f9a65ccba83cdd471bc6

Observation 98b04cc2-17fc-4d29-882b-6bf07547d5c8 · outbound

This paper cites You only hear once: a yolo-like algorithm for audio segmentation and sound event detection.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding You only hear once: a yolo-like algorithm for audio segmentation and sound event detection

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:06.182668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:53:00.581690Z digest=sha256:d9fdd075ca98286471ee0123a133b02cd4b69151710231495d14cad7b1a306fd

Observation 7d18d4f0-4ac5-40a9-bb98-6fa3bfae5bf0 · outbound

This paper cites Temporal Action Proposal Generation with Transformers.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Temporal Action Proposal Generation with Transformers

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T21:53:00.679239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:53:00.679239Z digest=sha256:3dda808cee3b216e1a023d2ea33dffd1b5b8cdd96ef8e4885b9ae44d32d9bef3

Observation f8281b68-9c52-4f7c-b270-0664967836d9 · outbound

This paper cites Rcl: Recurrent continuous localization for temporal action detec- tion.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Rcl: Recurrent continuous localization for temporal action detec- tion

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:05.894285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:53:00.753442Z digest=sha256:2c660afb3fbb475736c0c53f10ad518e7f041534fe0c187a33cd825569063811

Observation 782a3c0a-83ca-4af2-a2aa-a91be51477db · outbound

This paper cites Internvideo2: Scaling foundation models for mul- timodal video understanding.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Internvideo2: Scaling foundation models for mul- timodal video understanding

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:05.706511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:53:00.861896Z digest=sha256:c48b8271d920efda96d97509ec1b23d13b8f584e4241c64d8ac0fde83d518fc9

Observation 75eb99e0-c396-48bb-8d86-7d12f6842739 · outbound

This paper cites An efficient spatio-temporal pyramid transformer for action detection.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding An efficient spatio-temporal pyramid transformer for action detection

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:05.526823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:53:00.939127Z digest=sha256:f50e4b73164aa05d0e45f40dfb764db98c529c35bb8f07a352485398b8729458

Observation e2ea7e57-8730-4fd1-b28f-78e1db40ad0e · outbound

This paper cites Dual attention matching for audio-visual event localization.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Dual attention matching for audio-visual event localization

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T21:53:01.011214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:53:01.011214Z digest=sha256:ae58d9e8639c60ad1301eb501978359ea1d0a8ea75033a5c09e7f0b1549010d6

Observation ac19a1c4-9689-4c5a-9d93-259124fe18b1 · outbound

This paper cites Dual relation network for temporal action localization.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Dual relation network for temporal action localization

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:05.267675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:53:01.089484Z digest=sha256:c84241c63a5189d9d43da60174f89e2b0b184d29660893115a9e6351666c2648

Observation 91a4d0c6-7b42-4adf-85ea-1b8476481198 · outbound

This paper cites Learning to refactor action and co-occurrence fea- tures for temporal action localization.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Learning to refactor action and co-occurrence fea- tures for temporal action localization

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:04.930022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:53:01.205808Z digest=sha256:82a91c8b3b33385bd987d8f3f6fe3b985dfc1ab306fb1231006f7cfc75775b7c

Observation eeeb6297-1c94-4403-b82a-165461958cee · outbound

This paper cites Audiovisual SlowFast Networks for Video Recognition.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Audiovisual SlowFast Networks for Video Recognition

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T21:53:01.288755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:53:01.288755Z digest=sha256:b38142bb5b581eda11e12761935101ba1a36b1c90383fc0266ea44f4a09656d9

Observation f2d36044-0dc2-4d7f-bd5c-48f4b73bdce9 · outbound

This paper cites G-tad: Sub-graph localization for tempo- ral action detection.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding G-tad: Sub-graph localization for tempo- ral action detection

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:04.649130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:53:01.409621Z digest=sha256:88dc623a1a309f89d92d8994b6ae358ccc5f5600e50f9775d2ab10da6ca1a6b9

Observation ace9aa87-5799-4b7e-a854-2a1df9253e52 · outbound

This paper cites Audio-visual event localization by learning spatial and semantic co-attention.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Audio-visual event localization by learning spatial and semantic co-attention

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:04.300311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:53:01.497261Z digest=sha256:0a8fc7e8dec1d9a069ec6018f3f3597e759179acdcd670d2a317e0f4b276f951

Observation 98644f61-cfbb-4f39-b9fd-1f9d39f146d3 · outbound

This paper cites Temporal pyramid network for action recognition.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Temporal pyramid network for action recognition

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:04.087532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:53:01.572693Z digest=sha256:4d889e8887749075cff5d32a16f6291c5572967b675f77f97f97c772d74e8fea

Observation 3e4f63c9-fc4c-4de1-b7bc-31f74b47d100 · outbound

This paper cites Revisiting anchor mechanisms for temporal ac- tion localization.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Revisiting anchor mechanisms for temporal ac- tion localization

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:03.961697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:53:01.667259Z digest=sha256:3c1d3842dcc37037efe9e899f7d205debd5fa6898e025c6f85cd7712888c3f0a

Observation da56df42-9e95-4be0-ae0e-8dd3e2543d4c · outbound

This paper cites Mpn: Multimodal parallel network for audio-visual event localization.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Mpn: Multimodal parallel network for audio-visual event localization

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:03.839549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:53:01.739029Z digest=sha256:d3b573352f6738d61577f8dd2ff9c22bc56ebba3ab0f66e255f2ce5816bbd653

Observation eec95379-296b-495c-a05a-27b0704c6cfc · outbound

This paper cites Mm-pyramid: Multimodal pyramid attentional network for audio-visual event localization and video pars- ing.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Mm-pyramid: Multimodal pyramid attentional network for audio-visual event localization and video pars- ing

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:03.739506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:53:01.822512Z digest=sha256:42eac1d779a662ce0cc282230c8b59365d5439dee0c58fe3191f11bfb11e54f6

Observation d9106c28-a30f-4748-a608-e214a731aa28 · outbound

This paper cites Graph con- volutional networks for temporal action localization.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Graph con- volutional networks for temporal action localization

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:03.628158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:53:01.887537Z digest=sha256:7b250077ff24a290291f633d69c966dbf833d28baaf09c6960ec8398e3c2ef83

Observation ee01785a-809a-4097-9844-f5cb33925766 · outbound

This paper cites Actionformer: Lo- calizing moments of actions with transformers.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Actionformer: Lo- calizing moments of actions with transformers

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T21:53:01.984274Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:53:01.984274Z digest=sha256:04860f9a83cd98b67fd1ad7e0d9cd0eb610a45c1279144260bc868fe749c44a6

Observation d8015066-5343-4a61-a764-7c223f5fe782 · outbound

This paper cites Video self- stitching graph network for temporal action localization.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Video self- stitching graph network for temporal action localization

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:03.522826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:53:02.073232Z digest=sha256:b4cfe38774ed41072eaf7f1ca4e54b1d0c9b44da55d8471e573f5280f51fc66f

Observation b311c305-b820-4278-8bce-a976fe78e285 · outbound

This paper cites Bottom-up temporal action localization with mutual regularization.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Bottom-up temporal action localization with mutual regularization

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:03.412100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:53:02.139404Z digest=sha256:3e6c3363f2a05b1218dabf3adf3b384d5a7e8b52b42adefe2371b19ffb7ec2ec

Observation 25fc7dc9-2928-4cbc-b6dd-d57993c85d7b · outbound

This paper cites Enriching local and global contexts for temporal action localization.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Enriching local and global contexts for temporal action localization

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:03.286231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:53:02.210024Z digest=sha256:c52df01e705304ab8d6f78001501ff6345546e1590a567ae5411ebbba05263bf

Observation 9b8f3925-1ac2-4f58-af5d-74b122c37715 · outbound

This paper cites Our code provides further information.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Our code provides further information

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:03.160024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:53:02.296518Z digest=sha256:c8a301a82b4263b0dc023026dd6193888a1ec3f6683d17cbc7bd8be88c46bc54

Observation bbc44088-7e4e-420f-a1c0-02c4353f1649 · outbound

This paper cites Performance was measured us- ing mAP@[0.5:0.05:0.95], along with the average mAP.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Performance was measured us- ing mAP@[0.5:0.05:0.95], along with the average mAP

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:03.035271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:53:02.400314Z digest=sha256:97fc842c2fcedf107b34d76ab509e04a6c53a133428c868f2d0a24f2ad0c1c86

Observation cda74189-f55c-4927-a1f0-11c380a57b4c · outbound

This paper cites This approach ensures greater feature consis- tency across different temporal resolutions, ultimately im- proving the regression of an event’s location.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding This approach ensures greater feature consis- tency across different temporal resolutions, ultimately im- proving the regression of an event’s location

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:02.924236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:53:02.484752Z digest=sha256:a1ca0c6e462c077ecf5c79bcf3153e77bd8fd0964e004003ed001ed26ac3808c

Pith citing papers

Observation 47b14a0d-d100-4117-a088-aa437b827869 · inbound

EAR: Enhancing Uni-Modal Representations for Weakly Supervised Audio-Visual Video Parsing cites this paper.

EAR: Enhancing Uni-Modal Representations for Weakly Supervised Audio-Visual Video Parsing DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-12T03:06:19.415238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-12T03:03:13.496634Z digest=sha256:035f6f6c7091ddfbddb8e4108abf24ad37c4ef75c6601fe1a207f7f8b95ba5e2