Pith. sign in

Paper Citation Record · LEDGER

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding

As of 8 August 2026, this Paper Citation Record lists 64 of 64 outbound references and 1 inbound Pith citation observation for arXiv:2506.23196.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.23196 v1

Coverage vector

measured 64 of 64 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T21:53:02.484752Z

measured 65 of 65 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-12T03:03:13.496634Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-12T03:06:19.413098Z

Reference resolution

64 of 64 outbound references displayed

  • verified exact1
  • verified fuzzy49
  • unresolved14
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6adf58fb-1397-49f7-ac85-0b38c8e91849 · outbound

This paper cites Maas: Multi-modal assignation for active speaker detection.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Maas: Multi-modal assignation for active speaker detection

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:12.228540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:52:56.383564Z digest=sha256:b1597a6b6abb5bd6c1d68bfe5d9841de4baafb3290089659d113325200fb9e7c

Observation 28ad82c1-9139-4562-a667-3d761d6b933e · outbound

This paper cites Hear Me Out: Fusional Approaches for Audio Augmented Temporal Action Localization.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Hear Me Out: Fusional Approaches for Audio Augmented Temporal Action Localization

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:56.439313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:56.439313Z digest=sha256:84bb28726756182a860d8937b65ada6ceec0be36e1b174b99d76ac101ef48153

Observation d47090ca-d83c-438b-ae89-a2525b2859f1 · outbound

This paper cites Activitynet: A large-scale video benchmark for human activity understanding.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Activitynet: A large-scale video benchmark for human activity understanding

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:12.031192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:52:56.521994Z digest=sha256:e97572c236f79d0a2a1f74d7716fa99ed07b60ce9531e5df93047aa62c43ceb6

Observation 571de3b4-4ced-492a-bbb9-51b03117b901 · outbound

This paper cites Quo vadis, action recognition? a new model and the kinetics dataset.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Quo vadis, action recognition? a new model and the kinetics dataset

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:11.821632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:52:56.624292Z digest=sha256:35649ec3e703dccc19b7f9d052b8c7a08d0de3882e32d09714e8709df1d6aeb2

Observation 7803e0e6-ca3a-4bb7-bcb6-0a03ba5ed688 · outbound

This paper cites Augmented transformer with adaptive graph for tem- poral action proposal generation.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Augmented transformer with adaptive graph for tem- poral action proposal generation

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:11.636257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:52:56.734272Z digest=sha256:7d8fba2fdcd29fddef56671be32b2847079d602709a2ea0f189148d24820bbc5

Observation 7951d87f-64e0-4430-b5da-69ee90e7f60b · outbound

This paper cites Re- thinking the faster r-cnn architecture for temporal action localization.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Re- thinking the faster r-cnn architecture for temporal action localization

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:11.462906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:52:56.792960Z digest=sha256:60fe743b2f3f04d2594bc2920ed6ff0fb21a638af103a7e4877c7df5cda7c46c

Observation e9949550-2b40-4f80-abfb-904784a08b92 · outbound

This paper cites Tallformer: Temporal ac- tion localization with a long-memory transformer.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Tallformer: Temporal ac- tion localization with a long-memory transformer

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:11.244143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:52:56.866148Z digest=sha256:1ae475e73b4288cd5c6caad69177e3766d251778f5e0942212056707f678b02d

Observation 86994997-3545-44b3-b9fa-b88eb1c620a3 · outbound

This paper cites Yolo-world: Real-time open-vocabulary object detection.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Yolo-world: Real-time open-vocabulary object detection

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:56.942447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:56.942447Z digest=sha256:5e25ed64bc001e2559f951af5c7f2f986fb99b881147bf7be0319f6dc3aba479

Observation ad95e9b2-49f4-46ee-8d25-d3c8b3e77441 · outbound

This paper cites Rescaling egocentric vision: Collection, pipeline and chal- lenges for epic-kitchens-100.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Rescaling egocentric vision: Collection, pipeline and chal- lenges for epic-kitchens-100

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:11.029369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:52:57.080544Z digest=sha256:57e5fcd2571259eb36a62675737bb82e3126b09a396a373c07379c829d0b3015

Observation 839918ca-bc76-4068-b0b0-a4420566535b · outbound

This paper cites Slowfast networks for video recognition.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Slowfast networks for video recognition

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:10.855384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:52:57.209727Z digest=sha256:a779aa21565c8662b9252664e6ad5fd76db7cfa3bd1d7bbf6477f7ced868decb

Observation 4945c444-97ff-41ba-8e53-54c9f5253efe · outbound

This paper cites Multi-Resolution Audio-Visual Feature Fusion for Temporal Action Localization.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Multi-Resolution Audio-Visual Feature Fusion for Temporal Action Localization

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-08-06T21:53:02.785904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:52:57.316945Z digest=sha256:f54b3a4d20dab0dac50fe9a1017fb1fb60a0beb1fcc5fbb8fe8a34f43905f4b3

Observation 1ff4a861-5dc7-4132-94a8-9ed8a34b6d33 · outbound

This paper cites Audio set: An ontology and human- labeled dataset for audio events.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Audio set: An ontology and human- labeled dataset for audio events

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:10.698271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:52:57.408175Z digest=sha256:5295647e80674d2829e291dfbc998aa38bd628766da72f016ad3d2adaefeb2e0

Observation c02bd514-8400-48f0-aab9-74bc46808358 · outbound

This paper cites Dense-localizing audio-visual events in untrimmed videos: A large-scale benchmark and baseline.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Dense-localizing audio-visual events in untrimmed videos: A large-scale benchmark and baseline

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:10.463129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:52:57.495631Z digest=sha256:6df878ffd1dea4137fc51f39528cfcdc2df4fe3e65521c5cdd1e5a90781bebe7

Observation 8c37d1a1-7b8a-46d5-88bd-e30a619c88e8 · outbound

This paper cites Momentum contrast for unsupervised visual rep- resentation learning.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Momentum contrast for unsupervised visual rep- resentation learning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:57.588220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:57.588220Z digest=sha256:16378c132553c03c4c27b3a79769b7f3f7340647f4b7652fa019247e79277612

Observation 6c7fba88-18db-429b-b253-bbc1b260e576 · outbound

This paper cites Cnn archi- tectures for large-scale audio classification.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Cnn archi- tectures for large-scale audio classification

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:10.259205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:52:57.718925Z digest=sha256:514339a28bb5069ca6920989a9fdedae3e7bc6b34ef27dc95815063701830652

Observation 4dda55a1-49df-46a2-b808-68de94d62b04 · outbound

This paper cites Mix and local- ize: Localizing sound sources in mixtures.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Mix and local- ize: Localizing sound sources in mixtures

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:57.797874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:57.797874Z digest=sha256:3233ba497cffe6d0a693b0ca52d76079bf2e65b6fef0618cf90aa01d3493cae3

Observation 45643fd4-60c9-4124-aee3-f8be220466a2 · outbound

This paper cites in the wild.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding in the wild

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:10.051074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:52:57.881569Z digest=sha256:5d61ae678f6aaf321358ff4fbb05582cf2fc9a1ff71f46a5e4f304d1db8bc5ee

Observation 80271b71-b850-4d92-9f4c-96a1906fe8c5 · outbound

This paper cites Causal inference meets deep learning: A compre- hensive survey.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Causal inference meets deep learning: A compre- hensive survey

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:09.820174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:52:58.000025Z digest=sha256:59b65dee3da28d41727f232d8213090239857e57156e1fb47bd4118f46c74c03

Observation b31a13d5-21c4-4f59-8c87-5aed19bb0808 · outbound

This paper cites Epic-fusion: Audio-visual temporal bind- ing for egocentric action recognition.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Epic-fusion: Audio-visual temporal bind- ing for egocentric action recognition

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:09.637366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:52:58.137657Z digest=sha256:9676f3a5d4f2e64ccdb6a823d17a15b29b8d5b444cca53b4b5c58d2e3072f541

Observation 174d3af3-f042-43d5-b345-337e0e8e60ae · outbound

This paper cites MERT: Acoustic Music Understanding Model with Large-Scale Self-supervised Training.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding MERT: Acoustic Music Understanding Model with Large-Scale Self-supervised Training

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:58.255447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:58.255447Z digest=sha256:f4885a7269d724bb305c0f90ccde9833d459698d67cdd7c1d873a2fec201a9e5

Observation 67fe8343-b81f-444f-9065-b2f8139842fa · outbound

This paper cites Learning salient boundary feature for anchor- free temporal action localization.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Learning salient boundary feature for anchor- free temporal action localization

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:09.492907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:52:58.345630Z digest=sha256:10ea5c59cf53ec9356b651d55ec4d9e65ce7f79934020d2634ef11871791e586

Observation 9f0d2f44-002d-4240-addb-f9b5143a8927 · outbound

This paper cites Bsn: Boundary sensitive network for temporal action proposal generation.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Bsn: Boundary sensitive network for temporal action proposal generation

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:09.315122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:52:58.449901Z digest=sha256:68485fefede61896430c05de4807f360b04b8ad8c250b176fdcd21e4270871a8

Observation 4f7e9a09-c5c0-4f1b-b4aa-68115365b8dc · outbound

This paper cites Bmn: Boundary-matching network for temporal action pro- posal generation.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Bmn: Boundary-matching network for temporal action pro- posal generation

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:58.542307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:58.542307Z digest=sha256:164e464f12bd163087e8253d2f82667c99d1c539747099f3967e4bb495d5fab7

Observation c11b35a9-a764-4db3-b9e1-42bc4f26be31 · outbound

This paper cites Progressive boundary refine- ment network for temporal action detection.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Progressive boundary refine- ment network for temporal action detection

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:09.184494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:52:58.669443Z digest=sha256:089e9dd028c95c9757924f2f0da95016d9d0a4b674fbb6bf7cb86030a1ab4983

Observation 145a2007-89c2-451d-8ebb-17466280acb4 · outbound

This paper cites Dense modality interaction network for audio-visual event localization.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Dense modality interaction network for audio-visual event localization

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:09.038197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:52:58.774936Z digest=sha256:e3ecd778a3790acb7c63b7efd09cb0ef9f098f631b80395112234899654a5f24

Observation 8e6cbd72-b2e8-469b-be94-c96ae7f4cf77 · outbound

This paper cites Multi-shot temporal event localization: a benchmark.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Multi-shot temporal event localization: a benchmark

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:08.873567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:52:58.879243Z digest=sha256:3eac44699860d3f68d440c3d0e5f4b341455a2fa78c465b1ba801302e02384b6

Observation 26109f82-ccde-42db-88d9-595505edf9da · outbound

This paper cites End-to-end temporal ac- tion detection with transformer.IEEE Transactions on Image Processing, 31:5427–5441, 2022.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding End-to-end temporal ac- tion detection with transformer.IEEE Transactions on Image Processing, 31:5427–5441, 2022

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:08.703945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:52:58.983937Z digest=sha256:8952a7179d3a89932de21fe9bae1d6c80cd60cdebfe77d172dd8f0bbc5b5a355

Observation 0a9edb4e-f849-4d1f-8cdc-95ea4c662f25 · outbound

This paper cites Gaussian temporal awareness networks for action localization.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Gaussian temporal awareness networks for action localization

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:08.524721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:52:59.119360Z digest=sha256:4d1c49e70c5c85cb66b0256b6260f11e28a3e7a40fb5fe2482648d66dfdec7ae

Observation cbf46e29-3e88-4069-84ca-3f61652542e3 · outbound

This paper cites Proposal-free temporal action detection via global segmen- tation mask learning.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Proposal-free temporal action detection via global segmen- tation mask learning

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:08.374265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:52:59.215515Z digest=sha256:39eaacaee0a74e0d965ece3fa403cc172a6d4e3e55ccc084459c1ad9b2d71642

Observation f4bc45dc-7480-4baa-8f38-360c9f7af30d · outbound

This paper cites Attention bottlenecks for multimodal fusion.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Attention bottlenecks for multimodal fusion

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:08.185290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:52:59.303981Z digest=sha256:a231bab016db5aa579c6897808cdb698e5bb400925f1bfe396d7ea088f52a61a

Observation fc53abfb-da85-4cd7-a9d3-e13ef2b7475a · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding DINOv2: Learning Robust Visual Features without Supervision

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:59.408856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:59.408856Z digest=sha256:5a60082ffbf46debb68ce5b239d6f59b60332c9574a355161e0f52129c2d2cd9

Observation 57daa4b0-d0df-4a9f-90cc-335b1600a2e1 · outbound

This paper cites Audio-visual scene analysis with self-supervised multisensory features.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Audio-visual scene analysis with self-supervised multisensory features

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:08.046715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:52:59.543003Z digest=sha256:a30ac171334840e5e7478c0a0fce4e314784d33aa3bd31d4d15b6f331eac6b3a

Observation ccab2e5c-8bc7-4cca-a040-28fc3e3f8836 · outbound

This paper cites A review of deep learning techniques in audio event recognition (aer) applications.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding A review of deep learning techniques in audio event recognition (aer) applications

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:07.877014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:52:59.622948Z digest=sha256:6fed80ce1c944f6e432d70cceecc57d86b9a83e03aac864950c534c744d79c4e

Observation 7cd8b596-8223-4b2a-a5e4-5003dea88cea · outbound

This paper cites Owl (observe, watch, listen): Audiovisual temporal context for localizing actions in ego- centric videos.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Owl (observe, watch, listen): Audiovisual temporal context for localizing actions in ego- centric videos

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:07.703910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:52:59.736617Z digest=sha256:d413d804712b4740aa35fadb6ec4ad74243e74b50cc61d06cf1f0f430111aa43

Observation d160e83d-f8e2-4f1a-887e-ef62cb617543 · outbound

This paper cites You only look once: Unified, real-time object de- tection.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding You only look once: Unified, real-time object de- tection

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:59.832187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:59.832187Z digest=sha256:f7519d979c52e726ea08c4d29a26a338f2a933644c355cd3aa30fdbc13bbf69f

Observation 9b5c229e-e276-4ca3-91a4-bba32b5d3587 · outbound

This paper cites Action sensitivity learning for temporal action localization.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Action sensitivity learning for temporal action localization

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:59.922852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:59.922852Z digest=sha256:0d7f7864b3ae7c200784c94187c2fb1fba2d784046f2a280a613c152966187d0

Observation e4f8db01-5425-4c93-b34b-101ce7945aec · outbound

This paper cites Temporal Action Localization with Enhanced Instant Discriminability.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Temporal Action Localization with Enhanced Instant Discriminability

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T21:53:00.048092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:53:00.048092Z digest=sha256:03295ea1730a2d0115396c1023bb37f3aa2ee9938ac66856f6fa7c2c4e66f2eb

Observation 6f57e020-44bd-4943-8662-b77ad44fd57b · outbound

This paper cites Tridet: Temporal action detection with relative boundary modeling.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Tridet: Temporal action detection with relative boundary modeling

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:07.463989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:53:00.153296Z digest=sha256:0edbda77f27e714213c8f49124c952c503910eeecbe9f59d88b19cdc3548db30

Observation 022467a2-d6ab-4a35-9887-f852a27e4e6e · outbound

This paper cites Re- laxed transformer decoders for direct action proposal gener- ation.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Re- laxed transformer decoders for direct action proposal gener- ation

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:07.134101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:53:00.248074Z digest=sha256:ee4dc410a0cda69f05562cda94f84f1ac8e5d7632ee61c78ff49be31f1d8af62

Observation 3a408326-4418-43bf-9b67-396df641f1d3 · outbound

This paper cites Audio-visual event localization in unconstrained videos.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Audio-visual event localization in unconstrained videos

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:06.791765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:53:00.340814Z digest=sha256:9dedd309e6234eff5021d27cb732a3569278e400c6a783f8bc0a596f81e0c89c

Observation 0dd7ef7a-81b6-4b5e-a317-d1eabb452cf0 · outbound

This paper cites Deep learning-based action detection in untrimmed videos: A survey.IEEE Transactions on Pattern Analysis and Machine Intelligence , 45(4):4302– 4320, 2022.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Deep learning-based action detection in untrimmed videos: A survey.IEEE Transactions on Pattern Analysis and Machine Intelligence , 45(4):4302– 4320, 2022

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:06.518398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:53:00.491979Z digest=sha256:c6d7b0582392604545dfd5e2eb1856aaa24e7dab83974d679992fe74c5daaa6b

Observation 98b04cc2-17fc-4d29-882b-6bf07547d5c8 · outbound

This paper cites You only hear once: a yolo-like algorithm for audio segmentation and sound event detection.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding You only hear once: a yolo-like algorithm for audio segmentation and sound event detection

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:06.182668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:53:00.581690Z digest=sha256:91e3ffea109249ca3cb60238a475450e05814ed4f3591ccbacc22483018e123f

Observation 7d18d4f0-4ac5-40a9-bb98-6fa3bfae5bf0 · outbound

This paper cites Temporal Action Proposal Generation with Transformers.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Temporal Action Proposal Generation with Transformers

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T21:53:00.679239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:53:00.679239Z digest=sha256:29716ef788eb9a45eddb8fcf00617e89eaf7a07625bb51fdff82c102b8cbcdae

Observation f8281b68-9c52-4f7c-b270-0664967836d9 · outbound

This paper cites Rcl: Recurrent continuous localization for temporal action detec- tion.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Rcl: Recurrent continuous localization for temporal action detec- tion

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:05.894285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:53:00.753442Z digest=sha256:d3c3e47e7a033d367046bd4a9605e349ca389a597c1fdad4fa58341564152203

Observation 782a3c0a-83ca-4af2-a2aa-a91be51477db · outbound

This paper cites Internvideo2: Scaling foundation models for mul- timodal video understanding.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Internvideo2: Scaling foundation models for mul- timodal video understanding

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:05.706511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:53:00.861896Z digest=sha256:932a5204b7a9cb6273726302e47f02af1d142c9cfbda6ddac590595cff80d194

Observation 75eb99e0-c396-48bb-8d86-7d12f6842739 · outbound

This paper cites An efficient spatio-temporal pyramid transformer for action detection.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding An efficient spatio-temporal pyramid transformer for action detection

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:05.526823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:53:00.939127Z digest=sha256:44d9253f51f292453cf410a676be5087be4c6500a18d80479c6aa72a71eb8e91

Observation e2ea7e57-8730-4fd1-b28f-78e1db40ad0e · outbound

This paper cites Dual attention matching for audio-visual event localization.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Dual attention matching for audio-visual event localization

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T21:53:01.011214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:53:01.011214Z digest=sha256:d5d7ab841443211d27fd095de21d41fd854d5b2879ca2261abda2cbc17d7a0ed

Observation ac19a1c4-9689-4c5a-9d93-259124fe18b1 · outbound

This paper cites Dual relation network for temporal action localization.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Dual relation network for temporal action localization

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:05.267675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:53:01.089484Z digest=sha256:fc8fcf387b2a21af2c029d18441e3d0638c464296da362dc63d8c362784b0d7a

Observation 91a4d0c6-7b42-4adf-85ea-1b8476481198 · outbound

This paper cites Learning to refactor action and co-occurrence fea- tures for temporal action localization.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Learning to refactor action and co-occurrence fea- tures for temporal action localization

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:04.930022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:53:01.205808Z digest=sha256:d195ff03a462498976c008e214a288e17f00e63e8adc2a50d943dc32f3bea522

Observation eeeb6297-1c94-4403-b82a-165461958cee · outbound

This paper cites Audiovisual SlowFast Networks for Video Recognition.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Audiovisual SlowFast Networks for Video Recognition

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T21:53:01.288755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:53:01.288755Z digest=sha256:651a653fcf29428f8553a116cd8f4aa4316f739a8b2e9f50897facabb04d0102

Observation f2d36044-0dc2-4d7f-bd5c-48f4b73bdce9 · outbound

This paper cites G-tad: Sub-graph localization for tempo- ral action detection.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding G-tad: Sub-graph localization for tempo- ral action detection

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:04.649130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:53:01.409621Z digest=sha256:79b7439171bc24fde00e4b49b44e128860116fd04b512413a07e8782b0b5c4f5

Observation ace9aa87-5799-4b7e-a854-2a1df9253e52 · outbound

This paper cites Audio-visual event localization by learning spatial and semantic co-attention.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Audio-visual event localization by learning spatial and semantic co-attention

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:04.300311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:53:01.497261Z digest=sha256:41dc11beef42e8eaef8f592e6d24c551e46bac657abc5fa15c0125390a7eac5e

Observation 98644f61-cfbb-4f39-b9fd-1f9d39f146d3 · outbound

This paper cites Temporal pyramid network for action recognition.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Temporal pyramid network for action recognition

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:04.087532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:53:01.572693Z digest=sha256:d30cd650735a73264e30833d251a0b01be1e25f1cc68da45a99bc1dd4bcf8534

Observation 3e4f63c9-fc4c-4de1-b7bc-31f74b47d100 · outbound

This paper cites Revisiting anchor mechanisms for temporal ac- tion localization.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Revisiting anchor mechanisms for temporal ac- tion localization

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:03.961697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:53:01.667259Z digest=sha256:cca815818069a9ecd2d97d1521fd8a29c77269d612f98c692570c084c3ceb118

Observation da56df42-9e95-4be0-ae0e-8dd3e2543d4c · outbound

This paper cites Mpn: Multimodal parallel network for audio-visual event localization.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Mpn: Multimodal parallel network for audio-visual event localization

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:03.839549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:53:01.739029Z digest=sha256:0db26a032f03c430b308a4ad2c1258823d65bf4297e243b8ceba5bb01439f6dc

Observation eec95379-296b-495c-a05a-27b0704c6cfc · outbound

This paper cites Mm-pyramid: Multimodal pyramid attentional network for audio-visual event localization and video pars- ing.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Mm-pyramid: Multimodal pyramid attentional network for audio-visual event localization and video pars- ing

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:03.739506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:53:01.822512Z digest=sha256:7a770d1b3712c99fdbef713db476417ed2d6d224f42aed23d258d8e4272733ba

Observation d9106c28-a30f-4748-a608-e214a731aa28 · outbound

This paper cites Graph con- volutional networks for temporal action localization.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Graph con- volutional networks for temporal action localization

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:03.628158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:53:01.887537Z digest=sha256:fbef023757fcd548fe3610c000652853469166145fb6ffc68e6099dd6c3d0603

Observation ee01785a-809a-4097-9844-f5cb33925766 · outbound

This paper cites Actionformer: Lo- calizing moments of actions with transformers.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Actionformer: Lo- calizing moments of actions with transformers

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T21:53:01.984274Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:53:01.984274Z digest=sha256:27dddf9f631d3634f309f2319ee9b8d8359caaf31edc26fc8a64e2c251821b83

Observation d8015066-5343-4a61-a764-7c223f5fe782 · outbound

This paper cites Video self- stitching graph network for temporal action localization.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Video self- stitching graph network for temporal action localization

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:03.522826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:53:02.073232Z digest=sha256:0f7ace6a899b8f59c96d9e32f36f78a4ec396af46f227857a6e34e35a2646da2

Observation b311c305-b820-4278-8bce-a976fe78e285 · outbound

This paper cites Bottom-up temporal action localization with mutual regularization.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Bottom-up temporal action localization with mutual regularization

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:03.412100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:53:02.139404Z digest=sha256:969895b257e513a7e65fafb058e6a7fd1334e7127a3868d4f74e6d854ffdd6e2

Observation 25fc7dc9-2928-4cbc-b6dd-d57993c85d7b · outbound

This paper cites Enriching local and global contexts for temporal action localization.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Enriching local and global contexts for temporal action localization

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:03.286231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:53:02.210024Z digest=sha256:f95cec802287cefb5bc6192e9d4170bb7b8433073dd9e6e8d94dd02d20e3563c

Observation 9b8f3925-1ac2-4f58-af5d-74b122c37715 · outbound

This paper cites Our code provides further information.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Our code provides further information

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:03.160024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:53:02.296518Z digest=sha256:aee64fae7b4de8af82625fbcb3811c671e2b716a9730a9b89688884a256a14da

Observation bbc44088-7e4e-420f-a1c0-02c4353f1649 · outbound

This paper cites Performance was measured us- ing mAP@[0.5:0.05:0.95], along with the average mAP.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding Performance was measured us- ing mAP@[0.5:0.05:0.95], along with the average mAP

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:03.035271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:53:02.400314Z digest=sha256:2e1416cdc067a37e74fedf264716931d4e403176a91ecdfa31d0cb83f75df55e

Observation cda74189-f55c-4927-a1f0-11c380a57b4c · outbound

This paper cites This approach ensures greater feature consis- tency across different temporal resolutions, ultimately im- proving the regression of an event’s location.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding This approach ensures greater feature consis- tency across different temporal resolutions, ultimately im- proving the regression of an event’s location

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:53:02.924236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:53:02.484752Z digest=sha256:38be3899a48409656780712c51d3eca225c986baa34dba251c74305368e3ca5a

Pith citing papers

Observation 47b14a0d-d100-4117-a088-aa437b827869 · inbound

EAR: Enhancing Uni-Modal Representations for Weakly Supervised Audio-Visual Video Parsing cites this paper.

EAR: Enhancing Uni-Modal Representations for Weakly Supervised Audio-Visual Video Parsing DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-12T03:06:19.415238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-12T03:03:13.496634Z digest=sha256:88c6c265ecf72b4156ce5898fa48a77591f34d62cb1a0fe83f69adb23f743dfc