Pith. sign in

Paper Citation Record · LEDGER

Span-based Localizing Network for Natural Language Video Localization

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 11 inbound Pith citation observations for arXiv:2004.13931.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2004.13931 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 11 of 11 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 11 of 11 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T04:45:26.462185Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T20:56:14.529226Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation efadc4be-a648-4ab5-bd4d-166249e945d8 · inbound

DemaFormer: Damped Exponential Moving Average Transformer with Energy-Based Modeling for Temporal Language Grounding cites this paper.

DemaFormer: Damped Exponential Moving Average Transformer with Energy-Based Modeling for Temporal Language Grounding Span-based Localizing Network for Natural Language Video Localization

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-24T04:53:54.016914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-24T04:51:07.439654Z digest=sha256:bc5eb5732b17e8c1f2f8bd71ddcc8032795dee56ce976fe2e26f5fd774cf3daa

Observation 15d63241-d3b1-4a46-a0c3-6eb7bfb71829 · inbound

Multi-Scale Contrastive Learning for Video Temporal Grounding cites this paper.

Multi-Scale Contrastive Learning for Video Temporal Grounding Span-based Localizing Network for Natural Language Video Localization

Reference 67

Resolution
verified exact
arxiv_id, observed 2026-05-23T07:32:43.053634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-23T07:29:46.573683Z digest=sha256:3180a9c3c3f6573f9afaabbb3845e45cbeb3f6b2c253d9ff6d76e65eddc4374b

Observation d09e348f-ea91-49d7-a6bc-4a78a13b70bb · inbound

Moment of Untruth: Dealing with Negative Queries in Video Moment Retrieval cites this paper.

Moment of Untruth: Dealing with Negative Queries in Video Moment Retrieval Span-based Localizing Network for Natural Language Video Localization

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-08T04:45:26.462185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:45:26.462185Z digest=sha256:789754f4a0662a54efbbce7f7bccbddc9db9420ce014312c0dfc325b39089bfc

Observation 1f2ab522-d356-44f2-9857-32b249d58841 · inbound

DeCafNet: Delegate and Conquer for Efficient Temporal Grounding in Long Videos cites this paper.

DeCafNet: Delegate and Conquer for Efficient Temporal Grounding in Long Videos Span-based Localizing Network for Natural Language Video Localization

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T15:07:13.924824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:07:13.924824Z digest=sha256:77d1615bd7fada2450227e71afe7c4b04fd7cb0cd6b5f7b1a805d81fd66d0c8f

Observation 8fa31332-25ce-4b36-930f-88143aa3d190 · inbound

MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering cites this paper.

MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering Span-based Localizing Network for Natural Language Video Localization

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T23:29:28.693277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:29:28.693277Z digest=sha256:066016ecfc522097995b11c17152cba655d84627743c088565ff19f9057e1360

Observation 69d7761d-49f7-4f9d-bb83-9b7e0e341bfc · inbound

MS-DETR: Towards Effective Video Moment Retrieval and Highlight Detection by Joint Motion-Semantic Learning cites this paper.

MS-DETR: Towards Effective Video Moment Retrieval and Highlight Detection by Joint Motion-Semantic Learning Span-based Localizing Network for Natural Language Video Localization

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-06T17:01:56.789076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:01:56.789076Z digest=sha256:393d760baa6ef2a65b391f45f801d78e6721c0dd6b6cc740036dee5d0155735a

Observation 8d2e8ea1-3eb3-4cce-b704-39e7b0c783fe · inbound

HIPPO-Video: Simulating Watch Histories with Large Language Models for Personalized Video Highlighting cites this paper.

HIPPO-Video: Simulating Watch Histories with Large Language Models for Personalized Video Highlighting Span-based Localizing Network for Natural Language Video Localization

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T15:17:58.184421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:17:58.184421Z digest=sha256:c7ff56ee86cbbfd15463c357f66519bd84b702e9b741609794295f00631e7f0f

Observation 5dc1d1af-fa7d-466a-a405-990cb0ba80c6 · inbound

UniversalVTG: A Universal and Lightweight Foundation Model for Video Temporal Grounding cites this paper.

UniversalVTG: A Universal and Lightweight Foundation Model for Video Temporal Grounding Span-based Localizing Network for Natural Language Video Localization

Reference 62

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T05:51:02.385460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T17:55:56.385801Z digest=sha256:d3a9c0f2ae3b29b4c8fa3754b793c2135d999125eaa394e7120357f9a461ac07

Observation da339d9c-afb7-4394-ad2a-e7de88a1e12b · inbound

MASRA: MLLM-Assisted Semantic-Relational Consistent Alignment for Video Temporal Grounding cites this paper.

MASRA: MLLM-Assisted Semantic-Relational Consistent Alignment for Video Temporal Grounding Span-based Localizing Network for Natural Language Video Localization

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:51:31.166366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-08T01:30:17.546118Z digest=sha256:3f7ee3d4f45a35602c0e9a76f7fc447c341d80700ffbc604536be13af4ffc05c

Observation 633bdd30-49df-46b2-bf1a-5b28286d07f1 · inbound

EvoGround: Self-Evolving Video Agents for Video Temporal Grounding cites this paper.

EvoGround: Self-Evolving Video Agents for Video Temporal Grounding Span-based Localizing Network for Natural Language Video Localization

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-14T19:32:52.529810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-14T19:29:47.356665Z digest=sha256:0f708f432f95803995184b914d3c4b81a6a89e163a697a63f9fd3d718597346e

Observation 260a7219-af8f-4735-87bc-eae83eed1ee5 · inbound

CoSTL: Comprehensive Spatial-Temporal Representation Learning for Moment Retrieval and Highlight Detection cites this paper.

CoSTL: Comprehensive Spatial-Temporal Representation Learning for Moment Retrieval and Highlight Detection Span-based Localizing Network for Natural Language Video Localization

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-07-01T20:56:14.530960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T17:34:40.317688Z digest=sha256:fd152e7955b14e5f84634cd961fe6d314efca073e439a0f1ccd2161f9fd829e8