Pith. sign in

Paper Citation Record · LEDGER

Human Action CLIPs: Detecting AI-generated Human Motion

As of 19 August 2026, this Paper Citation Record lists 17 of 17 outbound references and 3 inbound Pith citation observations for arXiv:2412.00526.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.00526 v2

Coverage vector

measured 17 of 17 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T05:23:20.684103Z

measured 20 of 20 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T21:12:23.019873Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-20T15:13:25.052578Z

Reference resolution

17 of 17 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved17
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 85c9dbce-8aa6-46ec-9002-eca253eaa487 · outbound

This paper cites People are poorly equipped to detect AI-powered voice clones.

Human Action CLIPs: Detecting AI-generated Human Motion People are poorly equipped to detect AI-powered voice clones

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T05:23:20.615674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:23:20.615674Z digest=sha256:1ea845692f5b9f078bf2993d49277da8bb2b6b42c52d5db55f05ff0af9a98be5

Observation 9f07d460-e61c-4e4d-bdc7-9dcf7f770f24 · outbound

This paper cites Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets.

Human Action CLIPs: Detecting AI-generated Human Motion Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T05:23:20.620799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:23:20.620799Z digest=sha256:a093c1bd4cdc4119706464077dfa572229e32802e243ffa83c8eb14f571bfab3

Observation ed0b7a9b-2803-4c74-a8e5-10033bf0e925 · outbound

This paper cites Playing for 3D Human Recovery.

Human Action CLIPs: Detecting AI-generated Human Motion Playing for 3D Human Recovery

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T05:23:20.625984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:23:20.625984Z digest=sha256:e353d4c8400265244916e8979275f4418dda02004c3aeabdd2f05f5f9447ccaf

Observation fd524c48-284d-4e88-8b0d-6c44be0146b5 · outbound

This paper cites PaLI: A Jointly-Scaled Multilingual Language-Image Model.

Human Action CLIPs: Detecting AI-generated Human Motion PaLI: A Jointly-Scaled Multilingual Language-Image Model

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T05:23:20.631282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:23:20.631282Z digest=sha256:c96705075030388995fb49ed46f33338fea0efb84db86efc736eb3669fbacf3b

Observation 9cdfbb66-0a0b-4d1b-a09b-b090c93b4404 · outbound

This paper cites Exposing Lip-syncing Deepfakes from Mouth Inconsistencies.

Human Action CLIPs: Detecting AI-generated Human Motion Exposing Lip-syncing Deepfakes from Mouth Inconsistencies

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T05:23:20.636834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:23:20.636834Z digest=sha256:dfac090f77e9bb620f9fe2755fb45b99372a59ccbb027e61a14e7d953923439d

Observation b1a2d967-7dcc-487f-a665-7a7c303f991f · outbound

This paper cites Exploring the Adversarial Robustness of CLIP for AI-generated Image Detection.

Human Action CLIPs: Detecting AI-generated Human Motion Exploring the Adversarial Robustness of CLIP for AI-generated Image Detection

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T05:23:20.641475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:23:20.641475Z digest=sha256:393daa2a3e560b490bcf8a34302e572c5acf81dadebb163fdf9054c25cb43234

Observation 6442f9c9-2048-40d6-9bda-349bc54598f1 · outbound

This paper cites VideoPoet: A Large Language Model for Zero-Shot Video Generation.

Human Action CLIPs: Detecting AI-generated Human Motion VideoPoet: A Large Language Model for Zero-Shot Video Generation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T05:23:20.646151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:23:20.646151Z digest=sha256:b3abbe6434b021ac88c78a49c1280ac04dda7a1d1d891f275d1676d41a7d02bc

Observation 71d4e8e1-5848-4744-aeb1-65df39c7d66b · outbound

This paper cites AnimateDiff-Lightning: Cross-Model Diffusion Distillation.

Human Action CLIPs: Detecting AI-generated Human Motion AnimateDiff-Lightning: Cross-Model Diffusion Distillation

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T05:23:20.655486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:23:20.655486Z digest=sha256:66121a99f60857da0f966663f63414a68e23cd5fe3710b6942a646ed5a4cbb03

Observation 1f938311-0eef-4ee4-86e1-db65f2fa704b · outbound

This paper cites DeCLIP: Decoding CLIP representations for deepfake localization.

Human Action CLIPs: Detecting AI-generated Human Motion DeCLIP: Decoding CLIP representations for deepfake localization

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T05:23:20.670086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:23:20.670086Z digest=sha256:c72c0d95c24075675805b947f369f2ba81f2432a99c42edfba71cdb76de28686

Observation 15b8295a-9ff7-47af-bd7f-be6e5e19c222 · outbound

This paper cites CLIP Models are Few-shot Learners: Empirical Studies on VQA and Visual Entailment.

Human Action CLIPs: Detecting AI-generated Human Motion CLIP Models are Few-shot Learners: Empirical Studies on VQA and Visual Entailment

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T05:23:20.674694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:23:20.674694Z digest=sha256:ae393cd30ae6ccc6a881131ba0caa6d2c068d2a16c8f8d79227a3bd7de14d017

Observation 6e7baf02-b2af-48f4-8b46-e556881897fa · outbound

This paper cites 4Real: Towards Photorealistic 4D Scene Generation via Video Diffusion Models.

Human Action CLIPs: Detecting AI-generated Human Motion 4Real: Towards Photorealistic 4D Scene Generation via Video Diffusion Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T05:23:20.684103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:23:20.684103Z digest=sha256:4f074e34d7a3faa000220ea4350b366b6f802868000fb678c2904d07a66bb6b0

Observation 81aa69e6-e2e5-4fa3-baf2-59251b8623b2 · outbound

This paper cites CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer.

Human Action CLIPs: Detecting AI-generated Human Motion CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer

Reference 2001

Resolution
unresolved
no resolver link, observed 2026-08-12T05:23:20.679515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:23:20.679515Z digest=sha256:e9598d2f35dfe0ab19ae714e3a4f256b20d5b1d3ce8eb7f347e7f684dbf70e3d

Observation 2931710b-9f08-4443-9809-b6230cff4857 · outbound

This paper cites Vidu: a Highly Consistent, Dynamic and Skilled Text-to-Video Generator with Diffusion Models.

Human Action CLIPs: Detecting AI-generated Human Motion Vidu: a Highly Consistent, Dynamic and Skilled Text-to-Video Generator with Diffusion Models

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-12T05:23:20.604699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:23:20.604699Z digest=sha256:3c5d897ac2f2645ff9888edb6961fc2ce20ca8a8cf2691e7c2e276022f87123b

Observation 0de5dad8-3326-4922-b73d-f82e5443a8b0 · outbound

This paper cites LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs.

Human Action CLIPs: Detecting AI-generated Human Motion LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-12T05:23:20.660085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:23:20.660085Z digest=sha256:afe447be176abf456a74534805fa593ec020a1320dc8a6eb7509c5e090e1f607

Observation d0fa75b4-7dfa-451b-bcfd-ee04f69247cc · outbound

This paper cites How Much Can CLIP Benefit Vision-and-Language Tasks?.

Human Action CLIPs: Detecting AI-generated Human Motion How Much Can CLIP Benefit Vision-and-Language Tasks?

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-12T05:23:20.665056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:23:20.665056Z digest=sha256:141ea4fba57ab14f0498ac68d3aaa985707693aece27d83ee3662f562d2c48bb

Observation 3997a9ce-1f6b-45bc-a7c4-e62e457061ee · outbound

This paper cites Jina CLIP: Your CLIP Model Is Also Your Text Retriever.

Human Action CLIPs: Detecting AI-generated Human Motion Jina CLIP: Your CLIP Model Is Also Your Text Retriever

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-12T05:23:20.650879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:23:20.650879Z digest=sha256:71e88c8ea72bd5c345771a0060df945402907fb4c44bcb7cf802504c60f477c7

Observation 993eeb03-e35c-4bb5-8cb1-4e4ccf310ffc · outbound

This paper cites Lumiere: A Space-Time Diffusion Model for Video Generation.

Human Action CLIPs: Detecting AI-generated Human Motion Lumiere: A Space-Time Diffusion Model for Video Generation

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-12T05:23:20.610451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:23:20.610451Z digest=sha256:55a8bb70a67e2ec5f08a70df7b123a210432e0f4a32e838ce552c0f67c6bb9f2

Pith citing papers

Observation d2dd6fc4-7569-4714-bdf3-b51737b727bd · inbound

Survey on AI-Generated Media Detection: From Non-MLLM to MLLM cites this paper.

Survey on AI-Generated Media Detection: From Non-MLLM to MLLM Human Action CLIPs: Detecting AI-generated Human Motion

Reference 141

Resolution
unresolved
no resolver link, observed 2026-08-08T21:12:23.019873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:12:23.019873Z digest=sha256:538d3c968c112c19e02af3cd7a034885dafd3ef8c514e22a91e7e16d5bcfd3fb

Observation 1fac0a74-999a-4860-9013-1b5e30fddcdd · inbound

SynthForensics: Benchmarking and Evaluating People-Centric Synthetic Video Deepfakes cites this paper.

SynthForensics: Benchmarking and Evaluating People-Centric Synthetic Video Deepfakes Human Action CLIPs: Detecting AI-generated Human Motion

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-16T07:37:33.016331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-16T07:32:30.580643Z digest=sha256:3e0ab6f45f101492e50874ba8c42303f3da0ed4fadd5e71893a176b3a5a3e89f

Observation 23b91fb0-a20f-4881-a85c-ab2c1b4876da · inbound

CAM-VFD: Cross-Attention Multimodal Video Forgery Detection cites this paper.

CAM-VFD: Cross-Attention Multimodal Video Forgery Detection Human Action CLIPs: Detecting AI-generated Human Motion

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-20T15:13:25.054227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T15:08:25.309094Z digest=sha256:77b1ea5b5aa2aee4eaab648f9a65438120687c57000392cfcbc48efda1f78108