Pith. sign in

Paper Citation Record · LEDGER

ViTPose: Simple Vision Transformer Baselines for Human Pose Estimation

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 9 inbound Pith citation observations for arXiv:2204.12484.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2204.12484 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 9 of 9 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T22:43:35.766548Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T14:59:55.881195Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 9a2f4f26-73b5-4640-960f-ac26fa3166da · inbound

LiCamPose: Combining Multi-View LiDAR and RGB Cameras for Robust Single-timestamp 3D Human Pose Estimation cites this paper.

LiCamPose: Combining Multi-View LiDAR and RGB Cameras for Robust Single-timestamp 3D Human Pose Estimation ViTPose: Simple Vision Transformer Baselines for Human Pose Estimation

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-24T04:38:53.662646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-24T04:38:05.013487Z digest=sha256:64dc6a83e14cd8212451370924ddb24abae76f0d5db4ca84085d37f86e572414

Observation d8277b2a-9c79-4a82-8811-17f75591c3fd · inbound

IMASHRIMP: Automatic White Shrimp (Penaeus vannamei) Biometrical Analysis from Laboratory Images Using Computer Vision and Deep Learning cites this paper.

IMASHRIMP: Automatic White Shrimp (Penaeus vannamei) Biometrical Analysis from Laboratory Images Using Computer Vision and Deep Learning ViTPose: Simple Vision Transformer Baselines for Human Pose Estimation

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T20:34:11.858918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:34:11.858918Z digest=sha256:29e8ed207f31da60f386d140251f52d127c4ac0c4517b149ddd6bb47fe8b5f84

Observation c43331d7-32b4-49a9-b036-c96b36276fa0 · inbound

Prompt Engineering in Segment Anything Model: Methodologies, Applications, and Emerging Challenges cites this paper.

Prompt Engineering in Segment Anything Model: Methodologies, Applications, and Emerging Challenges ViTPose: Simple Vision Transformer Baselines for Human Pose Estimation

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-06T17:56:03.799771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:56:03.799771Z digest=sha256:59435f82e7620d1b561e353e45ddae2ea042e667361c000c745036de0e11ad92

Observation 70691822-ca09-41a5-99ec-9f3115eea11a · inbound

Joint angle based learning to refine kinematic human pose estimation cites this paper.

Joint angle based learning to refine kinematic human pose estimation ViTPose: Simple Vision Transformer Baselines for Human Pose Estimation

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T17:23:34.997088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:23:34.997088Z digest=sha256:12b45753b9004edaece673564f1d4a96dc87635207696511834ecc5543bf9d90

Observation f04c8987-4a2f-4206-9a41-85b60e932721 · inbound

Voice-guided Orchestrated Intelligence for Clinical Evaluation (VOICE): A Voice AI Agent System for Prehospital Stroke Assessment cites this paper.

Voice-guided Orchestrated Intelligence for Clinical Evaluation (VOICE): A Voice AI Agent System for Prehospital Stroke Assessment ViTPose: Simple Vision Transformer Baselines for Human Pose Estimation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T22:43:35.766548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:43:35.766548Z digest=sha256:031e4c5de9d6a3e4877267c59fbbae56d9018d1a20894fe630f50e401f4aa11c

Observation d77ea705-6962-4e0f-b2a1-7babb4f8852b · inbound

Wan-S2V: Audio-Driven Cinematic Video Generation cites this paper.

Wan-S2V: Audio-Driven Cinematic Video Generation ViTPose: Simple Vision Transformer Baselines for Human Pose Estimation

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-05T16:24:59.404743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:24:59.404743Z digest=sha256:2e86981cc24822c7f1723b3a218e700ca8c812054e4f4dc0afd6da9d7aed62a9

Observation b73d90c3-7ac4-4752-9c61-c8ab89f6ed4b · inbound

MoViD: View-Invariant 3D Human Pose Estimation via Motion-View Disentanglement cites this paper.

MoViD: View-Invariant 3D Human Pose Estimation via Motion-View Disentanglement ViTPose: Simple Vision Transformer Baselines for Human Pose Estimation

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:28:00.269211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-14T21:23:13.451470Z digest=sha256:213e637aaa250a3e8911fe89bd84b6bd8805926667518cd1e1cf06205324a2f6

Observation 51cb353f-964f-46b9-9fcb-b1e314c05059 · inbound

TaskNPoint: How to Teach Your Humanoid to Hit a Backhand in Minutes cites this paper.

TaskNPoint: How to Teach Your Humanoid to Hit a Backhand in Minutes ViTPose: Simple Vision Transformer Baselines for Human Pose Estimation

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-07-04T14:59:55.882640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T01:57:35.170470Z digest=sha256:f1adab8273e7b938ce554597165c535358f686724fa307abcc8b2d6d8775f2b3

Observation ffe45601-26d7-495c-9e3b-2e46511faa18 · inbound

Unleashing Infinite Motion: Scaling Expressive Quadrupedal Motion via Generative Video Priors cites this paper.

Unleashing Infinite Motion: Scaling Expressive Quadrupedal Motion via Generative Video Priors ViTPose: Simple Vision Transformer Baselines for Human Pose Estimation

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-07-01T17:05:51.239271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T04:10:56.435946Z digest=sha256:161887209844a9e1e913c1d8f6afd9e6c053941d98c6a247b1762bf3113a3af2