Pith. sign in

Paper Citation Record · LEDGER

The AVA-Kinetics Localized Human Actions Video Dataset

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 7 inbound Pith citation observations for arXiv:2005.00214.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2005.00214 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 7 of 7 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T13:18:19.238341Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T09:45:39.584394Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 305de34c-d1fc-4b85-9f6a-546e3a9cb41d · inbound

InternVideo: General Video Foundation Models via Generative and Discriminative Learning cites this paper.

InternVideo: General Video Foundation Models via Generative and Discriminative Learning The AVA-Kinetics Localized Human Actions Video Dataset

Reference 84

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T00:36:53.369430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-17T00:36:53.235740Z digest=sha256:dcb54c6729e337093090da87a966b3d162960706c99fc7bf86825c49374e674c

Observation 48000244-df96-43b4-9096-5426b9012c8f · inbound

What to Say and When to Say it: Live Fitness Coaching as a Testbed for Situated Interaction cites this paper.

What to Say and When to Say it: Live Fitness Coaching as a Testbed for Situated Interaction The AVA-Kinetics Localized Human Actions Video Dataset

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-23T22:53:33.150965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-23T22:51:08.753650Z digest=sha256:1d0392e7d0fea17946dbeecd4c44e2c85442bd9a358e270699d4f448955a64b1

Observation b0333049-2055-49d2-8586-450b11b41444 · inbound

Enhancing Video Understanding: Deep Neural Networks for Spatiotemporal Analysis cites this paper.

Enhancing Video Understanding: Deep Neural Networks for Spatiotemporal Analysis The AVA-Kinetics Localized Human Actions Video Dataset

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-08T13:18:19.238341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:18:19.238341Z digest=sha256:859e7b88eb5294d069eea65edd8544b4ba58ace7680b2136c67681e75f3aeff6

Observation 86d7a5a5-92a7-47e3-bac2-c46bd06143ca · inbound

DeepFake Doctor: Diagnosing and Treating Audio-Video Fake Detection cites this paper.

DeepFake Doctor: Diagnosing and Treating Audio-Video Fake Detection The AVA-Kinetics Localized Human Actions Video Dataset

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:38.362189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:38.362189Z digest=sha256:a6218abb67f6a0cc47669128819b8ad8627ca4f33a76fd372d8c30b8f7c0700d

Observation db8411f1-fb83-4b46-b0ce-257ad16ea095 · inbound

From Vision To Language through Graph of Events in Space and Time: An Explainable Self-supervised Approach cites this paper.

From Vision To Language through Graph of Events in Space and Time: An Explainable Self-supervised Approach The AVA-Kinetics Localized Human Actions Video Dataset

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T19:43:48.160980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:43:48.160980Z digest=sha256:3ff36b8854e6fc863ffeda934c5c1a2e529d4b5c5b4ec2d47f62dd3400ba1c34

Observation 48262cbb-4f0e-42c7-bee7-2b80bc12f169 · inbound

Video Understanding by Design: How Datasets Shape Video Models cites this paper.

Video Understanding by Design: How Datasets Shape Video Models The AVA-Kinetics Localized Human Actions Video Dataset

Reference 181

Resolution
unresolved
no resolver link, observed 2026-08-04T19:37:42.079446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:37:42.079446Z digest=sha256:c9eda63e000137b24a7e3167d57392e71f2c83d86f6d66820b10c62820cebf91

Observation 0cf0de00-af71-4821-b323-90da2cf1af18 · inbound

Learning to Deny: Action Denial in Multimodal Large Language Models cites this paper.

Learning to Deny: Action Denial in Multimodal Large Language Models The AVA-Kinetics Localized Human Actions Video Dataset

Reference 37

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T09:45:39.585810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-01T06:21:09.996386Z digest=sha256:e5f58e5639a468d81574e96fba395b2c05e99c959266d00910f12aa7dc98ce4b