Pith. sign in

Paper Citation Record · LEDGER

ViTPose: Simple Vision Transformer Baselines for Human Pose Estimation

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 9 inbound Pith citation observations for arXiv:2204.12484.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2204.12484 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 9 of 9 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T22:43:35.766548Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T14:59:55.881195Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 9a2f4f26-73b5-4640-960f-ac26fa3166da · inbound

LiCamPose: Combining Multi-View LiDAR and RGB Cameras for Robust Single-timestamp 3D Human Pose Estimation cites this paper.

LiCamPose: Combining Multi-View LiDAR and RGB Cameras for Robust Single-timestamp 3D Human Pose Estimation ViTPose: Simple Vision Transformer Baselines for Human Pose Estimation

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-24T04:38:53.662646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-24T04:38:05.013487Z digest=sha256:789eeeafe4bf8ebb4f9c797d85e50aa51b6164b8e207778e2ff70dc626a05e06

Observation d8277b2a-9c79-4a82-8811-17f75591c3fd · inbound

IMASHRIMP: Automatic White Shrimp (Penaeus vannamei) Biometrical Analysis from Laboratory Images Using Computer Vision and Deep Learning cites this paper.

IMASHRIMP: Automatic White Shrimp (Penaeus vannamei) Biometrical Analysis from Laboratory Images Using Computer Vision and Deep Learning ViTPose: Simple Vision Transformer Baselines for Human Pose Estimation

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T20:34:11.858918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:34:11.858918Z digest=sha256:3c98d0f8da97eeeee418011a91c3ede408a8e4548267f28a17329c323de19b61

Observation c43331d7-32b4-49a9-b036-c96b36276fa0 · inbound

Prompt Engineering in Segment Anything Model: Methodologies, Applications, and Emerging Challenges cites this paper.

Prompt Engineering in Segment Anything Model: Methodologies, Applications, and Emerging Challenges ViTPose: Simple Vision Transformer Baselines for Human Pose Estimation

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-06T17:56:03.799771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:56:03.799771Z digest=sha256:aa39f9a84f34527c678ab4e05e6cc228e37eae6a2fb1a89039b6a39969715882

Observation 70691822-ca09-41a5-99ec-9f3115eea11a · inbound

Joint angle based learning to refine kinematic human pose estimation cites this paper.

Joint angle based learning to refine kinematic human pose estimation ViTPose: Simple Vision Transformer Baselines for Human Pose Estimation

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T17:23:34.997088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:23:34.997088Z digest=sha256:fed685c4144d0ba8951a0573af1cc0ccf005b02a0bd7c52e7ed538f4152a4b40

Observation f04c8987-4a2f-4206-9a41-85b60e932721 · inbound

Voice-guided Orchestrated Intelligence for Clinical Evaluation (VOICE): A Voice AI Agent System for Prehospital Stroke Assessment cites this paper.

Voice-guided Orchestrated Intelligence for Clinical Evaluation (VOICE): A Voice AI Agent System for Prehospital Stroke Assessment ViTPose: Simple Vision Transformer Baselines for Human Pose Estimation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T22:43:35.766548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:43:35.766548Z digest=sha256:8ee6b55750ae9237f99ae0127d7bf494cace0e3825464b1b4d78d543e764dec8

Observation d77ea705-6962-4e0f-b2a1-7babb4f8852b · inbound

Wan-S2V: Audio-Driven Cinematic Video Generation cites this paper.

Wan-S2V: Audio-Driven Cinematic Video Generation ViTPose: Simple Vision Transformer Baselines for Human Pose Estimation

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-05T16:24:59.404743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:24:59.404743Z digest=sha256:8f1ee5ec97a9f73ad487c757d0cd96b5cfe96d6f067e61a28fb330109b7c5dc6

Observation b73d90c3-7ac4-4752-9c61-c8ab89f6ed4b · inbound

MoViD: View-Invariant 3D Human Pose Estimation via Motion-View Disentanglement cites this paper.

MoViD: View-Invariant 3D Human Pose Estimation via Motion-View Disentanglement ViTPose: Simple Vision Transformer Baselines for Human Pose Estimation

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:28:00.269211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-14T21:23:13.451470Z digest=sha256:9055256b67dd2c149ce9702664f9e82afa2b3e61b4c19bf930d7ccb53ea9ef93

Observation 51cb353f-964f-46b9-9fcb-b1e314c05059 · inbound

TaskNPoint: How to Teach Your Humanoid to Hit a Backhand in Minutes cites this paper.

TaskNPoint: How to Teach Your Humanoid to Hit a Backhand in Minutes ViTPose: Simple Vision Transformer Baselines for Human Pose Estimation

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-07-04T14:59:55.882640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T01:57:35.170470Z digest=sha256:8d92bcdbff817dfe52ab15a1096457f1b3b91f44cb34410d98dc08a8df074409

Observation ffe45601-26d7-495c-9e3b-2e46511faa18 · inbound

Unleashing Infinite Motion: Scaling Expressive Quadrupedal Motion via Generative Video Priors cites this paper.

Unleashing Infinite Motion: Scaling Expressive Quadrupedal Motion via Generative Video Priors ViTPose: Simple Vision Transformer Baselines for Human Pose Estimation

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-07-01T17:05:51.239271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T04:10:56.435946Z digest=sha256:b6b80254660a7a9930e6c9b12fbfa58ba4439d4cbed37fbe66d8d09c9d793fe0