Pith. sign in

Paper Citation Record · LEDGER

Hierarchical Event Memory for Accurate and Low-latency Online Video Temporal Grounding

As of 9 August 2026, this Paper Citation Record lists 60 of 60 outbound references and 0 inbound Pith citation observations for arXiv:2508.04546.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.04546 v1

Coverage vector

measured 60 of 60 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T00:00:03.806881Z

measured 60 of 60 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

60 of 60 outbound references displayed

  • verified exact0
  • verified fuzzy44
  • unresolved15
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c8c106e4-32ed-4feb-a878-f068044377f2 · outbound

This paper cites A new surveil- lance and security alert system based on real-time motion de- tection.Journal of Smart Systems Research, 4:31–47, 2023.

Hierarchical Event Memory for Accurate and Low-latency Online Video Temporal Grounding A new surveil- lance and security alert system based on real-time motion de- tection.Journal of Smart Systems Research, 4:31–47, 2023

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:00:04.376583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T00:00:03.626468Z digest=sha256:dbcd3e8f64b0841e6fadbb408a381378937e1820e536b1381197a0d8473eef00

Observation 37c3ec9e-6e40-4b02-8e1f-346f489af6e7 · outbound

This paper cites Miniroad: Minimal rnn framework for online action detection.

Hierarchical Event Memory for Accurate and Low-latency Online Video Temporal Grounding Miniroad: Minimal rnn framework for online action detection

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:00:04.368076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T00:00:03.630495Z digest=sha256:c6181367b19f21c773454578b33b64ee59cb166b12e745a50703aa34ca302ad7

Observation 4102aee0-96d1-4b0d-afb5-ce00ecd9c7a2 · outbound

This paper cites E2e-load: end-to-end long-form online action de- tection.

Hierarchical Event Memory for Accurate and Low-latency Online Video Temporal Grounding E2e-load: end-to-end long-form online action de- tection

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:00:04.359095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T00:00:03.633772Z digest=sha256:7b3048227b07750dd464854eb23b98b9d7e8311b771e0bc0fc92d4b77bb9febd

Observation 68de790d-8497-4e7d-998c-dcd1201b4843 · outbound

This paper cites Gatehub: Gated history unit with background sup- pression for online action detection.

Hierarchical Event Memory for Accurate and Low-latency Online Video Temporal Grounding Gatehub: Gated history unit with background sup- pression for online action detection

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:00:04.350553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T00:00:03.638144Z digest=sha256:cb41a550fcd19333a970d83598059dc1069f38cd081683d4c80df5fd2f3de693

Observation c63d8f15-d39a-484b-a3b9-b56a872f5d50 · outbound

This paper cites Enhancing Long Video Understanding via Hierarchical Event-Based Memory.

Hierarchical Event Memory for Accurate and Low-latency Online Video Temporal Grounding Enhancing Long Video Understanding via Hierarchical Event-Based Memory

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T00:00:03.642025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:00:03.642025Z digest=sha256:5c70738abb6238581a288e1c36000deabbdcea2825cce575091acb3edaf56657

Observation 5766dccb-66ec-4831-9344-7f1a5f42bcf4 · outbound

This paper cites Empirical Evaluation of Gated Recurrent Neural Networks on Sequence Modeling.

Hierarchical Event Memory for Accurate and Low-latency Online Video Temporal Grounding Empirical Evaluation of Gated Recurrent Neural Networks on Sequence Modeling

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T00:00:03.645941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:00:03.645941Z digest=sha256:ee0f9ae8c28e26eae663ed9e8205976163079d79fe7e080f9a683e0a8842f9ee

Observation f4b83441-b569-49c3-9940-ef9418adec9c · outbound

This paper cites Moment detection in long tutorial videos.

Hierarchical Event Memory for Accurate and Low-latency Online Video Temporal Grounding Moment detection in long tutorial videos

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:00:04.341043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T00:00:03.650510Z digest=sha256:4d5b44554b1b6a3c359997b4869eda03be8a26b3573d0940bd4af92dfea46f00

Observation 4d64475c-9df3-4c21-8876-73943611cce8 · outbound

This paper cites Online ac- tion detection.

Hierarchical Event Memory for Accurate and Low-latency Online Video Temporal Grounding Online ac- tion detection

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:00:04.330770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T00:00:03.653383Z digest=sha256:ccd9147969f8f69eceb1bd496097078637179f594ea99802e08a010407c66aa2

Observation e3a56d9d-92f6-423a-82f3-9463e0ace12b · outbound

This paper cites Learning to discriminate information for online action detection.

Hierarchical Event Memory for Accurate and Low-latency Online Video Temporal Grounding Learning to discriminate information for online action detection

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:00:04.322006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T00:00:03.656452Z digest=sha256:b333e79f8084c42a03b00d7fa33bdcc11f118c63299d0806ac7a09fe4ad57e3b

Observation 9ae8c4fb-bdc2-4f11-bae0-81e5fe594f52 · outbound

This paper cites Temporal sentence grounding in streaming videos.

Hierarchical Event Memory for Accurate and Low-latency Online Video Temporal Grounding Temporal sentence grounding in streaming videos

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:00:04.313071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T00:00:03.659145Z digest=sha256:ac384801fba08d585521d936e36853481455c8d4e810d722394c871d04f7f445

Observation 1bc21709-7208-4148-878d-b399b0cc003b · outbound

This paper cites Tall: Temporal activity localization via language query.

Hierarchical Event Memory for Accurate and Low-latency Online Video Temporal Grounding Tall: Temporal activity localization via language query

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:00:04.303986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T00:00:03.661894Z digest=sha256:8974e02e41d9ef4236ac3b3fcdaa3eb0a23e15f5d57fec0a96e06e882115822b

Observation ddbaed9e-9cbd-46e4-b495-5ca378a84ecd · outbound

This paper cites Prentice Hall PTR, 1994.

Hierarchical Event Memory for Accurate and Low-latency Online Video Temporal Grounding Prentice Hall PTR, 1994

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:00:04.294359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T00:00:03.665577Z digest=sha256:44b384ce4684af8405a108bb318dcedba4e4409131af0143c152e12f56f0d102

Observation 22d561b4-8f13-4005-81e5-f64b2b7b4880 · outbound

This paper cites Ma-lmm: Memory-augmented large multimodal model for long-term video understanding.

Hierarchical Event Memory for Accurate and Low-latency Online Video Temporal Grounding Ma-lmm: Memory-augmented large multimodal model for long-term video understanding

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:00:04.284042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T00:00:03.668777Z digest=sha256:458299ad0767b145e7d99f6a5950024b9af0c586b508eb3facaef836d0db1050

Observation 7701116d-ee02-4b81-99de-e9b508b57309 · outbound

This paper cites Video activity localisation with uncertainties in temporal boundary.

Hierarchical Event Memory for Accurate and Low-latency Online Video Temporal Grounding Video activity localisation with uncertainties in temporal boundary

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:00:04.273144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T00:00:03.671586Z digest=sha256:718bb999161b9d6f410a78f3c42641fbcbdcffb7dccbfd966550986495d3ca32

Observation 96fe039f-1857-4ac6-aba6-d6b8fe041f4f · outbound

This paper cites Cag-qil: Context-aware actionness grouping via q im- itation learning for online temporal action localization.

Hierarchical Event Memory for Accurate and Low-latency Online Video Temporal Grounding Cag-qil: Context-aware actionness grouping via q im- itation learning for online temporal action localization

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:00:04.263894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T00:00:03.674278Z digest=sha256:86a0d1e58f7092f0d3e1dcf9e89ae035d61aadc7039d284afd1994e0a782a9c2

Observation b347f31d-2c9f-4030-81bd-a2a602ebd70a · outbound

This paper cites A sliding window scheme for online temporal action localization.

Hierarchical Event Memory for Accurate and Low-latency Online Video Temporal Grounding A sliding window scheme for online temporal action localization

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:00:04.252811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T00:00:03.677553Z digest=sha256:5f5fc6432e2d0ff6de24120e2663671400da420540d2c0f3b95435fb5e565fa2

Observation a38911d8-ad59-4ccc-9e24-2b367f69981d · outbound

This paper cites Dense-captioning events in videos.

Hierarchical Event Memory for Accurate and Low-latency Online Video Temporal Grounding Dense-captioning events in videos

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:00:04.243918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T00:00:03.681058Z digest=sha256:033e95fb52bb3f3f3f8f8f4637af62de50e319126bbe8e57a3d63cc51477d977

Observation 61f5090b-c327-43a5-bd3e-bdccb56cdc8c · outbound

This paper cites Efficient adaptive human-object inter- action detection with concept-guided memory, 2023.

Hierarchical Event Memory for Accurate and Low-latency Online Video Temporal Grounding Efficient adaptive human-object inter- action detection with concept-guided memory, 2023

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T00:00:03.684597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:00:03.684597Z digest=sha256:b4778f56a5451343a2a740571e1c729c45c3a992fa02848e3eafd3afb09d61e8

Observation e705e4b5-20a3-4d7d-9d95-b2cb8646400c · outbound

This paper cites Cross modal adaptive few-shot learning based on task depen- dence.Chinese Journal of Electronics, 32(1):85–96, 2023.

Hierarchical Event Memory for Accurate and Low-latency Online Video Temporal Grounding Cross modal adaptive few-shot learning based on task depen- dence.Chinese Journal of Electronics, 32(1):85–96, 2023

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:00:04.228389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T00:00:03.688413Z digest=sha256:7a24b3bb87747b007a7e9a7126620ccbb84c047c9cbe41e3ac88e74194ecd9e4

Observation 416f8863-437c-4ce6-ab09-062868f6a599 · outbound

This paper cites G2l: Semantically aligned and uniform video grounding via geodesic and game theory.

Hierarchical Event Memory for Accurate and Low-latency Online Video Temporal Grounding G2l: Semantically aligned and uniform video grounding via geodesic and game theory

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:00:04.218186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T00:00:03.691376Z digest=sha256:b0bb6efe7c40284a1a1ab365c10aa0937eb478296b5fe10df5b4c30204e34ac4

Observation 92a0772e-0c2a-4ed5-8f40-ed514094ea8d · outbound

This paper cites Mo- mentdiff: Generative video moment retrieval from random to real.Advances in neural information processing systems, 36, 2024.

Hierarchical Event Memory for Accurate and Low-latency Online Video Temporal Grounding Mo- mentdiff: Generative video moment retrieval from random to real.Advances in neural information processing systems, 36, 2024

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:00:04.208553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T00:00:03.694304Z digest=sha256:d64021d46f82100a5dc0d7ce4524b477d23d80bfe76f6de9d86a2595d2e0c630

Observation ea08fb25-3a90-4b9d-a5e6-d6e13bc7ceb8 · outbound

This paper cites Focal Loss for Dense Object Detection.

Hierarchical Event Memory for Accurate and Low-latency Online Video Temporal Grounding Focal Loss for Dense Object Detection

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T00:00:03.698296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:00:03.698296Z digest=sha256:b2b6d41ec204b4687fc60d4298e83138aca0dee2e7bc317bf2abf30c6d52f231

Observation 90fbad3b-e4c0-40fc-897e-d323c11b0485 · outbound

This paper cites Towards balanced alignment: Modal-enhanced semantic modeling for video moment re- trieval.

Hierarchical Event Memory for Accurate and Low-latency Online Video Temporal Grounding Towards balanced alignment: Modal-enhanced semantic modeling for video moment re- trieval

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:00:04.198675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T00:00:03.701578Z digest=sha256:d9f3a12f640f86c23368dcd7341771371c79839dc8ef3b1b92b9b74e3008a6b9

Observation 586018d7-d057-490a-9fb5-f9c59ebb7acd · outbound

This paper cites Decoupled Weight Decay Regularization.

Hierarchical Event Memory for Accurate and Low-latency Online Video Temporal Grounding Decoupled Weight Decay Regularization

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T00:00:03.704250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:00:03.704250Z digest=sha256:5ee073c958af366e08d91bca3d2fc526b9f1ad149fe4a50c0e1068c87e75ecbc

Observation bbc6bdd9-1d23-428d-83a0-cc781903a67b · outbound

This paper cites Towards generalisable video moment retrieval: Visual-dynamic injection to image-text pre-training.

Hierarchical Event Memory for Accurate and Low-latency Online Video Temporal Grounding Towards generalisable video moment retrieval: Visual-dynamic injection to image-text pre-training

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:00:04.187780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T00:00:03.707413Z digest=sha256:de5f64163a97e048f561d6129efae9e820ca0e5df568756b3d2309f5f065a636

Observation 682839cf-a0e6-4a5e-bb05-b3cf8e05db07 · outbound

This paper cites Zero-shot video moment retrieval from frozen vision-language models.

Hierarchical Event Memory for Accurate and Low-latency Online Video Temporal Grounding Zero-shot video moment retrieval from frozen vision-language models

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:00:04.178388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T00:00:03.709898Z digest=sha256:f418164fd0207751e2a70313e440d470be81e5a3d679d0bd10e7887f7d3b51f7

Observation 1c511538-94b7-4b1d-be38-a696d8d0e853 · outbound

This paper cites Snag: Scalable and accurate video grounding.

Hierarchical Event Memory for Accurate and Low-latency Online Video Temporal Grounding Snag: Scalable and accurate video grounding

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:00:04.167028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T00:00:03.712473Z digest=sha256:99e9b5fd278184da0f664dad60a2f7c94c8acfcca36f7b5d9074154da2f32d5f

Observation 29c00aae-dbda-4bac-8ab3-666817837939 · outbound

This paper cites Local- global video-text interactions for temporal grounding.

Hierarchical Event Memory for Accurate and Low-latency Online Video Temporal Grounding Local- global video-text interactions for temporal grounding

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:00:04.157203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T00:00:03.715246Z digest=sha256:da3046d732979660d0882bbda750ebf0343430ff5485a981f0b0811f08490c2a

Observation 62a2d395-520d-4e59-bda1-26624444bab3 · outbound

This paper cites Scanning Only Once: An End-to-end Framework for Fast Temporal Grounding in Long Videos.

Hierarchical Event Memory for Accurate and Low-latency Online Video Temporal Grounding Scanning Only Once: An End-to-end Framework for Fast Temporal Grounding in Long Videos

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T00:00:03.717769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:00:03.717769Z digest=sha256:ab0f4cb5da6b297e52752c6314b7a61287c1625585f3f611c3d5075d7881b8eb

Observation 55159b80-fd16-46c9-923d-d042b8b564fe · outbound

This paper cites An overview of cross-media retrieval: Concepts, methodologies, bench- marks, and challenges.IEEE Transactions on Circuits and Systems for Video Technology, 28:2372–2385, 2017.

Hierarchical Event Memory for Accurate and Low-latency Online Video Temporal Grounding An overview of cross-media retrieval: Concepts, methodologies, bench- marks, and challenges.IEEE Transactions on Circuits and Systems for Video Technology, 28:2372–2385, 2017

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:00:04.147813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T00:00:03.720559Z digest=sha256:3f03f0683332c61aa224b4ad81c332b9e819e642690c41cbf02c9d37c28b930c

Observation ecc23ea0-ba85-4e35-965a-b952cf0ef620 · outbound

This paper cites Streaming long video understanding with large language models.Advances in Neu- ral Information Processing Systems, 37:119336–119360,.

Hierarchical Event Memory for Accurate and Low-latency Online Video Temporal Grounding Streaming long video understanding with large language models.Advances in Neu- ral Information Processing Systems, 37:119336–119360,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T00:00:03.723182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:00:03.723182Z digest=sha256:3640001fdc02aa76df33787b498380709b015e6fd90d3a4c571c6b0185da400b

Observation de278ce3-8708-465c-b82b-22be9f626f49 · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

Hierarchical Event Memory for Accurate and Low-latency Online Video Temporal Grounding Learning transferable visual models from natural language supervi- sion

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T00:00:03.726376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:00:03.726376Z digest=sha256:b45a91174fb65297a3bc76e014bfa45545699daae2dc9d94965cffeb28982dc8

Observation ed936847-2d8d-4196-afb2-639f8251bbd4 · outbound

This paper cites Hat: History-augmented anchor transformer for on- line temporal action localization.

Hierarchical Event Memory for Accurate and Low-latency Online Video Temporal Grounding Hat: History-augmented anchor transformer for on- line temporal action localization

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:00:04.126061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T00:00:03.729869Z digest=sha256:a46835767ea01a6897a04d6424657df9dbf7de6e8497d5e1b67a6c6d31aa33e8

Observation 77504ab3-10f2-404f-9263-8aeda83c81e6 · outbound

This paper cites Coherent multi-sentence video description with variable level of detail.

Hierarchical Event Memory for Accurate and Low-latency Online Video Temporal Grounding Coherent multi-sentence video description with variable level of detail

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:00:04.116314Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T00:00:03.732731Z digest=sha256:742fd58ad7b5f4abdf6c15b313187b396bce3d8d70b2d3299cdc13f23f6b837b

Observation d59f1833-1610-415e-a2cd-397bb2018bc6 · outbound

This paper cites Online action detection in untrimmed, streaming videos-modeling and evaluation.

Hierarchical Event Memory for Accurate and Low-latency Online Video Temporal Grounding Online action detection in untrimmed, streaming videos-modeling and evaluation

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:00:04.106866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T00:00:03.735226Z digest=sha256:4d58ebd3bd3bd86c2327a1e21778c88ef3ee53eedb58d86b70d1a1a951594d49

Observation 567aa239-1c2b-4965-bc43-e1b8b8bde4f6 · outbound

This paper cites Mad: A scalable dataset for language grounding in videos from movie audio descriptions.

Hierarchical Event Memory for Accurate and Low-latency Online Video Temporal Grounding Mad: A scalable dataset for language grounding in videos from movie audio descriptions

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:00:04.097528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T00:00:03.737976Z digest=sha256:204c00a016c6d5354955aabf44e06763a64de2cb2262f6cd520fa6a519bfbeef

Observation 59b7e413-3c40-45c9-9f79-845b88d3bf28 · outbound

This paper cites Moviechat: From dense token to sparse memory for long video understanding.

Hierarchical Event Memory for Accurate and Low-latency Online Video Temporal Grounding Moviechat: From dense token to sparse memory for long video understanding

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:00:04.087010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T00:00:03.740748Z digest=sha256:ac1d25ce98592a334211852e7bbd41b6ebf5bd2cb259545d9de4c47f67e5f4dd

Observation e1182f0a-4318-4e2e-ad82-76ff4e3e519e · outbound

This paper cites Learning spatiotemporal features with 3d convolutional networks.

Hierarchical Event Memory for Accurate and Low-latency Online Video Temporal Grounding Learning spatiotemporal features with 3d convolutional networks

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T00:00:03.743312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:00:03.743312Z digest=sha256:2d9ef45f7cd64ed93dc75c63a525f8bfa46c07131987b9bfc4878d523c10948f

Observation a033cea5-7d45-4ca8-b740-712720ff790e · outbound

This paper cites Attention is all you need.Advances in neural information processing systems, 30, 2017.

Hierarchical Event Memory for Accurate and Low-latency Online Video Temporal Grounding Attention is all you need.Advances in neural information processing systems, 30, 2017

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:00:04.072095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T00:00:03.746238Z digest=sha256:2b660f05a0a453b54e730ad556efc98a86842faede446ba74a507dc803e56924

Observation 0e9005ce-0fdf-4021-94d6-b52be4d1e664 · outbound

This paper cites Oadtr: Online action detection with transformers.

Hierarchical Event Memory for Accurate and Low-latency Online Video Temporal Grounding Oadtr: Online action detection with transformers

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:00:04.062710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T00:00:03.748889Z digest=sha256:4b284abecf382a2d82b6dbc5ef38539023530fc7f4616d94fd9367868948ee5e

Observation 6e01456c-dabb-4883-965f-2ae7c1926e44 · outbound

This paper cites Efficient temporal extrapolation of multimodal large language models with temporal grounding bridge.

Hierarchical Event Memory for Accurate and Low-latency Online Video Temporal Grounding Efficient temporal extrapolation of multimodal large language models with temporal grounding bridge

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:00:04.053554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T00:00:03.751591Z digest=sha256:722bd64093e962d09b9808c195c54815f9d19d3cc327bf0517d22449d9ce2d6e

Observation 6c2b59d2-197a-469b-83ea-a207d082543a · outbound

This paper cites VideoLLaMB: Long Streaming Video Understanding with Recurrent Memory Bridges.

Hierarchical Event Memory for Accurate and Low-latency Online Video Temporal Grounding VideoLLaMB: Long Streaming Video Understanding with Recurrent Memory Bridges

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T00:00:03.754369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:00:03.754369Z digest=sha256:4073ea9b4669b43bf1e2187cfe1a6fb9f13a15a953038f246638f544c49b6ae8

Observation 781b5f8a-ddd4-4500-9748-5d442a38d4f9 · outbound

This paper cites Negative sample matters: A renaissance of met- ric learning for temporal grounding.

Hierarchical Event Memory for Accurate and Low-latency Online Video Temporal Grounding Negative sample matters: A renaissance of met- ric learning for temporal grounding

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:00:04.044174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T00:00:03.757493Z digest=sha256:b3fb06103b733325944ef6a3e20eb2abe62e9aa30836dccd297272d9fb394fe8

Observation 3537a7a4-eb0e-4060-908c-938ad3258cb0 · outbound

This paper cites Longvlm: Efficient long video understand- ing via large language models.

Hierarchical Event Memory for Accurate and Low-latency Online Video Temporal Grounding Longvlm: Efficient long video understand- ing via large language models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T00:00:03.760159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:00:03.760159Z digest=sha256:8fecdac993db4a848c0e319ee6ccd59abb7a47b2f4ea35ecf0b846bddb100f18

Observation 70e4c475-a2c6-4257-af6c-805eba6e4563 · outbound

This paper cites Bridging the gap: A unified video comprehension framework for mo- ment retrieval and highlight detection.

Hierarchical Event Memory for Accurate and Low-latency Online Video Temporal Grounding Bridging the gap: A unified video comprehension framework for mo- ment retrieval and highlight detection

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:00:04.026870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T00:00:03.762870Z digest=sha256:f024328a4dad8497ee87be43c2e2a3295d8ab68396b2201cf225b292c8533cbf

Observation 221385eb-fec1-4fb8-a76c-18145febdba6 · outbound

This paper cites Temporal recurrent networks for online action detection.

Hierarchical Event Memory for Accurate and Low-latency Online Video Temporal Grounding Temporal recurrent networks for online action detection

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T00:00:03.765400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:00:03.765400Z digest=sha256:d1e6c81046e34f79368ed476b4e6a160343acb2c4580892acbeafa359fd0b76c

Observation 4d79a847-4f31-4660-a963-303c645526f9 · outbound

This paper cites Long short-term trans- former for online action detection.Advances in Neural In- formation Processing Systems, 34:1086–1099, 2021.

Hierarchical Event Memory for Accurate and Low-latency Online Video Temporal Grounding Long short-term trans- former for online action detection.Advances in Neural In- formation Processing Systems, 34:1086–1099, 2021

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:00:04.011062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T00:00:03.768204Z digest=sha256:93d253a016957d2da51059889ef6117d1c80a3aec1fe26d98addc5b52c6ba4f1

Observation 746f2438-8cdd-40bd-8e19-b901c06f55d0 · outbound

This paper cites Active object detection with knowledge aggregation and distillation from large models.

Hierarchical Event Memory for Accurate and Low-latency Online Video Temporal Grounding Active object detection with knowledge aggregation and distillation from large models

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:00:04.001334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T00:00:03.770695Z digest=sha256:5bdb9bea229b3e98644af5c41b986b1e214d50675fda2fdded8b944b7ec5b9cc

Observation 6c0b1b3b-d15a-4bca-9b8d-3d113aa5ee20 · outbound

This paper cites Flowgananomaly: Flow-based anomaly network intrusion detection with adver- sarial learning.Chinese Journal of Electronics, 33(1):58–71,.

Hierarchical Event Memory for Accurate and Low-latency Online Video Temporal Grounding Flowgananomaly: Flow-based anomaly network intrusion detection with adver- sarial learning.Chinese Journal of Electronics, 33(1):58–71,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:00:03.992050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T00:00:03.773233Z digest=sha256:85dffe98243915ca67bdabb28a7f448570227036a836f5a43a729b7c8f57f0cd

Observation 2ba91c3e-102d-4de4-a5c1-c99f9b4b88de · outbound

This paper cites Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams.

Hierarchical Event Memory for Accurate and Low-latency Online Video Temporal Grounding Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T00:00:03.777349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:00:03.777349Z digest=sha256:0c6f8e3d185512727665a76c281c36f2d56cf5132712d2d6f4d85e25210c4fb6

Observation 2e867451-fc1f-4ef8-b0fb-a92529a3c4eb · outbound

This paper cites Learning 2d temporal adjacent networks for moment local- ization with natural language.

Hierarchical Event Memory for Accurate and Low-latency Online Video Temporal Grounding Learning 2d temporal adjacent networks for moment local- ization with natural language

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T00:00:03.780650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:00:03.780650Z digest=sha256:a92112d8539571303d5dec86e750f1f2c9991b8bd4cd40c43cb1a301e6c3b72f

Observation 72ce5440-5286-4afc-aff1-fc9f3b1b4b4a · outbound

This paper cites Progressive privileged knowledge distillation for on- line action detection.Pattern Recognition, 129:108741,.

Hierarchical Event Memory for Accurate and Low-latency Online Video Temporal Grounding Progressive privileged knowledge distillation for on- line action detection.Pattern Recognition, 129:108741,

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T00:00:03.783553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:00:03.783553Z digest=sha256:d648a4d3bd69f7aa3d31009feb3fc7b2300495b256df72637f50d8946b473b99

Observation 37c5946a-cc81-4186-8ffd-5bfe6b5459ba · outbound

This paper cites Real-time online video detection with temporal smoothing transformers.

Hierarchical Event Memory for Accurate and Low-latency Online Video Temporal Grounding Real-time online video detection with temporal smoothing transformers

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:00:03.970000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T00:00:03.786834Z digest=sha256:7a040f8a94769f4781f29b96a03636f967831b164254b2d496115b1d3fc53e9b

Observation 194be584-97ee-4265-a4a1-44f2a41d1b06 · outbound

This paper cites Unsupervised cross-media hashing learning via knowledge graph.Chinese Journal of Electronics, 31(6):1081–1091, 2022.

Hierarchical Event Memory for Accurate and Low-latency Online Video Temporal Grounding Unsupervised cross-media hashing learning via knowledge graph.Chinese Journal of Electronics, 31(6):1081–1091, 2022

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:00:03.959228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T00:00:03.790152Z digest=sha256:101279f8439acc7d5289e9be747afbf4f3160ce5b337cf1744c9aaedbd553854

Observation 667feeca-3c4f-479d-a0f3-a69890e413fd · outbound

This paper cites Weakly supervised video moment localization with con- trastive negative sample mining.

Hierarchical Event Memory for Accurate and Low-latency Online Video Temporal Grounding Weakly supervised video moment localization with con- trastive negative sample mining

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:00:03.947534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T00:00:03.792730Z digest=sha256:78c32cebe01993eea3cf29a5d6401dcc1b865d18a963b5b11fae9cd6d774806f

Observation 17cf5ad9-acdc-4bdd-a50b-57c3bc1917ae · outbound

This paper cites Weakly supervised temporal sentence grounding with gaussian-based contrastive proposal learn- ing.

Hierarchical Event Memory for Accurate and Low-latency Online Video Temporal Grounding Weakly supervised temporal sentence grounding with gaussian-based contrastive proposal learn- ing

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:00:03.935291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T00:00:03.795687Z digest=sha256:4f55713892e221c116456b0215c0d7d725dc2ca46132193f67bc47656a9be038

Observation 74e5305e-d75a-4166-8060-ac84589ef354 · outbound

This paper cites Generating structured pseudo labels for noise- resistant zero-shot video sentence localization.

Hierarchical Event Memory for Accurate and Low-latency Online Video Temporal Grounding Generating structured pseudo labels for noise- resistant zero-shot video sentence localization

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:00:03.925397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T00:00:03.798721Z digest=sha256:6b547488a80b540fe899cf6ff7a0c0be79a2ddde4c4c2b13a31687a11d89b66d

Observation 786ca839-0407-4b2b-abe6-e4301fabdf27 · outbound

This paper cites Phrase-level temporal relationship mining for temporal sentence localization.

Hierarchical Event Memory for Accurate and Low-latency Online Video Temporal Grounding Phrase-level temporal relationship mining for temporal sentence localization

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:00:03.915647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T00:00:03.801411Z digest=sha256:67d92c91ee33970fe834947ee0833f95000338735b5089dc414dc00991d27de1

Observation f0a2fcde-42ce-43b0-937a-15b51ae2dc2a · outbound

This paper cites Training-free video temporal grounding usinglarge-scale pre-trained models.

Hierarchical Event Memory for Accurate and Low-latency Online Video Temporal Grounding Training-free video temporal grounding usinglarge-scale pre-trained models

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:00:03.905828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T00:00:03.804048Z digest=sha256:e2e0a4e9ae98704633d0a33ba844b668052ae54da2e8185ba5c36ad876300333

Observation 9d37da64-71e8-46f9-b5ea-5b159f06b07d · outbound

This paper cites Faster and better learning for bounding box regres- sion., 2020, 34.DOI: https://doi.

Hierarchical Event Memory for Accurate and Low-latency Online Video Temporal Grounding Faster and better learning for bounding box regres- sion., 2020, 34.DOI: https://doi

Reference 60

Resolution
malformed identifier
no resolver link, observed 2026-08-06T00:00:03.806881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:00:03.806881Z digest=sha256:880a2589a3915aefce3df50141fef227ceba70ee6be8e685b615ffeca9859be5

Pith citing papers

No inbound Pith citation observations are available.