Pith. sign in

Paper Citation Record · LEDGER

Human Action CLIPs: Detecting AI-generated Human Motion

As of 19 August 2026, this Paper Citation Record lists 17 of 17 outbound references and 3 inbound Pith citation observations for arXiv:2412.00526.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.00526 v2

Coverage vector

measured 17 of 17 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T05:23:20.684103Z

measured 20 of 20 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T21:12:23.019873Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-20T15:13:25.052578Z

Reference resolution

17 of 17 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved17
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 85c9dbce-8aa6-46ec-9002-eca253eaa487 · outbound

This paper cites People are poorly equipped to detect AI-powered voice clones.

Human Action CLIPs: Detecting AI-generated Human Motion People are poorly equipped to detect AI-powered voice clones

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T05:23:20.615674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:23:20.615674Z digest=sha256:7cde3ac519c03a23636a65f81c0116a11d5fecac57db4f88d24fc200e120dc27

Observation 9f07d460-e61c-4e4d-bdc7-9dcf7f770f24 · outbound

This paper cites Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets.

Human Action CLIPs: Detecting AI-generated Human Motion Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T05:23:20.620799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:23:20.620799Z digest=sha256:2bad68a2d81981a04d792f0592cf09ef8dadaf15806ccb00a417f56beeed57aa

Observation ed0b7a9b-2803-4c74-a8e5-10033bf0e925 · outbound

This paper cites Playing for 3D Human Recovery.

Human Action CLIPs: Detecting AI-generated Human Motion Playing for 3D Human Recovery

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T05:23:20.625984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:23:20.625984Z digest=sha256:3aadd76e095291d0096537b8e8a2a3b0980effcad938ab081aab6d187cf9c132

Observation fd524c48-284d-4e88-8b0d-6c44be0146b5 · outbound

This paper cites PaLI: A Jointly-Scaled Multilingual Language-Image Model.

Human Action CLIPs: Detecting AI-generated Human Motion PaLI: A Jointly-Scaled Multilingual Language-Image Model

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T05:23:20.631282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:23:20.631282Z digest=sha256:73b59a2d1397aa3f7a0a1cdb2838c86841a0c7009a85f6b5c16622b55f0d8bc9

Observation 9cdfbb66-0a0b-4d1b-a09b-b090c93b4404 · outbound

This paper cites Exposing Lip-syncing Deepfakes from Mouth Inconsistencies.

Human Action CLIPs: Detecting AI-generated Human Motion Exposing Lip-syncing Deepfakes from Mouth Inconsistencies

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T05:23:20.636834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:23:20.636834Z digest=sha256:8f24013b38d57af09832d3d24b830ac9752cbdd83e689e4f8095d6aea07fdf11

Observation b1a2d967-7dcc-487f-a665-7a7c303f991f · outbound

This paper cites Exploring the Adversarial Robustness of CLIP for AI-generated Image Detection.

Human Action CLIPs: Detecting AI-generated Human Motion Exploring the Adversarial Robustness of CLIP for AI-generated Image Detection

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T05:23:20.641475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:23:20.641475Z digest=sha256:91466a5ad040e06bd3de0ab9ed4017a35b69b7633c9907a1526f34570c6909f4

Observation 6442f9c9-2048-40d6-9bda-349bc54598f1 · outbound

This paper cites VideoPoet: A Large Language Model for Zero-Shot Video Generation.

Human Action CLIPs: Detecting AI-generated Human Motion VideoPoet: A Large Language Model for Zero-Shot Video Generation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T05:23:20.646151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:23:20.646151Z digest=sha256:72249478a3f5781dddc8fc3caf80ac3aebcf2e2e4d6cf0a2f0d7ff9ae5a81528

Observation 71d4e8e1-5848-4744-aeb1-65df39c7d66b · outbound

This paper cites AnimateDiff-Lightning: Cross-Model Diffusion Distillation.

Human Action CLIPs: Detecting AI-generated Human Motion AnimateDiff-Lightning: Cross-Model Diffusion Distillation

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T05:23:20.655486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:23:20.655486Z digest=sha256:da48a00020432b3e99bc31dbf9e088a83156078c02c6ae338649434fd212bb2a

Observation 1f938311-0eef-4ee4-86e1-db65f2fa704b · outbound

This paper cites DeCLIP: Decoding CLIP representations for deepfake localization.

Human Action CLIPs: Detecting AI-generated Human Motion DeCLIP: Decoding CLIP representations for deepfake localization

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T05:23:20.670086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:23:20.670086Z digest=sha256:28d43f9b99262545ca3c5fe5e1751f098fd67453951d6196db9d80e2718a6bfc

Observation 15b8295a-9ff7-47af-bd7f-be6e5e19c222 · outbound

This paper cites CLIP Models are Few-shot Learners: Empirical Studies on VQA and Visual Entailment.

Human Action CLIPs: Detecting AI-generated Human Motion CLIP Models are Few-shot Learners: Empirical Studies on VQA and Visual Entailment

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T05:23:20.674694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:23:20.674694Z digest=sha256:a7669ce8f7e0151fbb3b3cf4b36c39f9e2b885b54504004e3b62541d1f285825

Observation 6e7baf02-b2af-48f4-8b46-e556881897fa · outbound

This paper cites 4Real: Towards Photorealistic 4D Scene Generation via Video Diffusion Models.

Human Action CLIPs: Detecting AI-generated Human Motion 4Real: Towards Photorealistic 4D Scene Generation via Video Diffusion Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T05:23:20.684103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:23:20.684103Z digest=sha256:1976cc34946ef979d461fa680a5d8bcc665b70b4fe57d4a99ece1595f9ee0de2

Observation 81aa69e6-e2e5-4fa3-baf2-59251b8623b2 · outbound

This paper cites CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer.

Human Action CLIPs: Detecting AI-generated Human Motion CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer

Reference 2001

Resolution
unresolved
no resolver link, observed 2026-08-12T05:23:20.679515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:23:20.679515Z digest=sha256:197c9282395a41cc9952eb76ece31419c2848609654fb71920254fafb725a1f0

Observation 2931710b-9f08-4443-9809-b6230cff4857 · outbound

This paper cites Vidu: a Highly Consistent, Dynamic and Skilled Text-to-Video Generator with Diffusion Models.

Human Action CLIPs: Detecting AI-generated Human Motion Vidu: a Highly Consistent, Dynamic and Skilled Text-to-Video Generator with Diffusion Models

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-12T05:23:20.604699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:23:20.604699Z digest=sha256:1d9bbee16a346502cb43c2630f418498485e2dd26e64c2e0288041ea9bd1fefc

Observation 0de5dad8-3326-4922-b73d-f82e5443a8b0 · outbound

This paper cites LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs.

Human Action CLIPs: Detecting AI-generated Human Motion LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-12T05:23:20.660085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:23:20.660085Z digest=sha256:0670f1ca05f28bf87c51079175fba97af03a27dc6b80da97f774b888fdc890e8

Observation d0fa75b4-7dfa-451b-bcfd-ee04f69247cc · outbound

This paper cites How Much Can CLIP Benefit Vision-and-Language Tasks?.

Human Action CLIPs: Detecting AI-generated Human Motion How Much Can CLIP Benefit Vision-and-Language Tasks?

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-12T05:23:20.665056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:23:20.665056Z digest=sha256:b8b88aa15a7f8117b47f83fe55d424997747a735362a449e74511a0740cef075

Observation 3997a9ce-1f6b-45bc-a7c4-e62e457061ee · outbound

This paper cites Jina CLIP: Your CLIP Model Is Also Your Text Retriever.

Human Action CLIPs: Detecting AI-generated Human Motion Jina CLIP: Your CLIP Model Is Also Your Text Retriever

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-12T05:23:20.650879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:23:20.650879Z digest=sha256:aae01aded0a25580af74e66b8f472a40f7923cd6fcd100be8bed6dd714e754d8

Observation 993eeb03-e35c-4bb5-8cb1-4e4ccf310ffc · outbound

This paper cites Lumiere: A Space-Time Diffusion Model for Video Generation.

Human Action CLIPs: Detecting AI-generated Human Motion Lumiere: A Space-Time Diffusion Model for Video Generation

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-12T05:23:20.610451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:23:20.610451Z digest=sha256:81518521b16b18a3f4655d0b08314cc1feca371e1e514c9cb2d4dd725a96f308

Pith citing papers

Observation d2dd6fc4-7569-4714-bdf3-b51737b727bd · inbound

Survey on AI-Generated Media Detection: From Non-MLLM to MLLM cites this paper.

Survey on AI-Generated Media Detection: From Non-MLLM to MLLM Human Action CLIPs: Detecting AI-generated Human Motion

Reference 141

Resolution
unresolved
no resolver link, observed 2026-08-08T21:12:23.019873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:12:23.019873Z digest=sha256:3aff686dc378ebc6aff644cf2f457a5f15844e857ce4ba1abfbd0db113c0fdc1

Observation 1fac0a74-999a-4860-9013-1b5e30fddcdd · inbound

SynthForensics: Benchmarking and Evaluating People-Centric Synthetic Video Deepfakes cites this paper.

SynthForensics: Benchmarking and Evaluating People-Centric Synthetic Video Deepfakes Human Action CLIPs: Detecting AI-generated Human Motion

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-16T07:37:33.016331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-16T07:32:30.580643Z digest=sha256:5d2529f529bf7e24736b0acbf9775d31ea4f1522293e7c5ca94238beed013e3b

Observation 23b91fb0-a20f-4881-a85c-ab2c1b4876da · inbound

CAM-VFD: Cross-Attention Multimodal Video Forgery Detection cites this paper.

CAM-VFD: Cross-Attention Multimodal Video Forgery Detection Human Action CLIPs: Detecting AI-generated Human Motion

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-20T15:13:25.054227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T15:08:25.309094Z digest=sha256:11d7b72520dab0f5ec434f7cefa107453ca767f50719a117209174af492cc74a