Pith. sign in

Paper Citation Record · LEDGER

Waterfall Transformer for Multi-person Pose Estimation

As of 17 August 2026, this Paper Citation Record lists 28 of 28 outbound references and 0 inbound Pith citation observations for arXiv:2411.18944.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.18944 v1

Coverage vector

measured 28 of 28 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T10:45:40.882265Z

measured 28 of 28 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

28 of 28 outbound references displayed

  • verified exact1
  • verified fuzzy20
  • unresolved7
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 49b2c69f-a853-4185-98c1-55aa3cb81cc3 · outbound

This paper cites OmniPose: A Multi-Scale Framework for Multi-Person Pose Estimation.

Waterfall Transformer for Multi-person Pose Estimation OmniPose: A Multi-Scale Framework for Multi-Person Pose Estimation

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-08-12T10:45:41.014544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T10:45:40.754930Z digest=sha256:a1aa7e950863189eebdb1465b8d0dbc2b3459dcbc019b50affe0881c58075624

Observation 3c1b5c4e-a886-4c86-a2ab-22003cca9fc9 · outbound

This paper cites Unipose+: A unified framework for 2d and 3d human pose es- timation in images and videos.

Waterfall Transformer for Multi-person Pose Estimation Unipose+: A unified framework for 2d and 3d human pose es- timation in images and videos

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:45:41.319831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T10:45:40.761083Z digest=sha256:26d927afa572a8c13ba130859f7647faaaa5a3dafdd064aed3053b56b9ce69e0

Observation 1b14281e-ec90-4655-9db5-e876def5fd98 · outbound

This paper cites BAPose: Bottom-up pose estimation with disentangled water- fall representations.

Waterfall Transformer for Multi-person Pose Estimation BAPose: Bottom-up pose estimation with disentangled water- fall representations

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:45:41.305921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T10:45:40.765786Z digest=sha256:2b73ae12ae011889d37019d86a81928f4f99be910c2157690b5fc122768415d7

Observation d239611b-80e3-46f0-92db-5cd14179a6aa · outbound

This paper cites Full-BAPose: Bottom up framework for full body pose estimation.

Waterfall Transformer for Multi-person Pose Estimation Full-BAPose: Bottom up framework for full body pose estimation

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:45:41.290399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T10:45:40.771217Z digest=sha256:fcb5245c61f62c4527f66b95c367039f6137a0fcb799ed0afd8b2daeac6f1261

Observation 1c42810f-6717-44a8-bbad-538185bc1cef · outbound

This paper cites Realtime multi-person 2d pose estimation us- ing part affinity fields.

Waterfall Transformer for Multi-person Pose Estimation Realtime multi-person 2d pose estimation us- ing part affinity fields

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:45:41.277295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T10:45:40.776181Z digest=sha256:5fd06060a05036a162f339010165121e0636b321abd1d33de8382b7d6e9a6001

Observation 8bf586b1-2653-4038-848a-11750dc2696f · outbound

This paper cites Yuille, and Xiaogang Wang.

Waterfall Transformer for Multi-person Pose Estimation Yuille, and Xiaogang Wang

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:45:41.263225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T10:45:40.780889Z digest=sha256:6d9c08cd815bc37f4e01f1a151b9c233d2078b309ed233166dc367833ce57963

Observation c2146e82-d257-474f-b143-50aff9440978 · outbound

This paper cites Openmmlab pose estimation toolbox and benchmark.

Waterfall Transformer for Multi-person Pose Estimation Openmmlab pose estimation toolbox and benchmark

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:45:41.246717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T10:45:40.786489Z digest=sha256:9e834a31278d088ed383e62c6e8eded73d18cc6467342d765f846bcfb1e75be7

Observation cedd7a66-f382-4208-bc40-5c12a5f30c62 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

Waterfall Transformer for Multi-person Pose Estimation An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T10:45:40.791075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:45:40.791075Z digest=sha256:a46df87c5cc4a9ee6a6ed886a6dd8d6a833fc6101ee504f0199f8424b91eeb55

Observation 6f70adb9-6405-4ea1-93da-b13d67a8cbab · outbound

This paper cites Dilated Neighborhood Attention Transformer.

Waterfall Transformer for Multi-person Pose Estimation Dilated Neighborhood Attention Transformer

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T10:45:40.795479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:45:40.795479Z digest=sha256:5ebe209987688c5f57fba1c2d24aed24fbc064ac9c0603d67ef4a252ae3e1e9a

Observation 6d842eda-a9d8-4aa0-965a-cf1d7799f921 · outbound

This paper cites Deep residual learning for image recognition.

Waterfall Transformer for Multi-person Pose Estimation Deep residual learning for image recognition

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:45:41.232533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T10:45:40.800136Z digest=sha256:b65fc7e1626636ec2c06b798371023628c25210087796921a3c51bea21f8c205

Observation 946f6bbe-7478-4852-9fd0-061826f8e1e1 · outbound

This paper cites Rethinking on Multi-Stage Networks for Human Pose Estimation.

Waterfall Transformer for Multi-person Pose Estimation Rethinking on Multi-Stage Networks for Human Pose Estimation

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T10:45:40.804871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:45:40.804871Z digest=sha256:9b250125d650aa072ac6601a3fd2ea41e6d61ee9ff1eda29f08b6356861d0883

Observation 44929b41-d367-4730-8037-b72c42a43728 · outbound

This paper cites Token- Pose: Learning keypoint tokens for human pose es- timation.

Waterfall Transformer for Multi-person Pose Estimation Token- Pose: Learning keypoint tokens for human pose es- timation

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:45:41.218227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T10:45:40.809858Z digest=sha256:fd5bbbcfea2df47315d6d59f7d2a11c2a5d52019ed1c5bd519e04094f1462859

Observation 0428547c-f1ef-4ccb-889c-ccd4a12be50c · outbound

This paper cites Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Doll´ar, and C.

Waterfall Transformer for Multi-person Pose Estimation Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Doll´ar, and C

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:45:41.204328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T10:45:40.814490Z digest=sha256:54e790c1818e8e625d49c1950f209e3fb06685947f50550bc8dc13e163fb3d60

Observation 51547b5c-ead1-4f7c-892f-42e806023fad · outbound

This paper cites Swin transformer: Hierarchical vision transformer using shifted windows.

Waterfall Transformer for Multi-person Pose Estimation Swin transformer: Hierarchical vision transformer using shifted windows

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:45:41.188848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T10:45:40.819136Z digest=sha256:461d938d53bc191b47d981fa7aedabf9dbae0c4412a6c3e9fd14c9db4360131f

Observation e14903af-0820-4d31-9532-4f5f4a3d2d0b · outbound

This paper cites Stacked hourglass networks for human pose estimation.

Waterfall Transformer for Multi-person Pose Estimation Stacked hourglass networks for human pose estimation

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:45:41.172765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T10:45:40.823482Z digest=sha256:699fc5b90872307a6fe4d0444e50c25c87227a1b07739a33f8252d009b621bec

Observation a8dec8df-d24b-4455-8348-d01eeae7c86d · outbound

This paper cites On the Convergence of Adam and Beyond.

Waterfall Transformer for Multi-person Pose Estimation On the Convergence of Adam and Beyond

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T10:45:40.827758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:45:40.827758Z digest=sha256:468adb205cd6112e5b9d78b6bfa520288c8c1ff10f4d8bf9a33923aa4a8be894

Observation 93cef3a9-45e3-4152-8840-d8d7348ef792 · outbound

This paper cites 15 keypoints is all you need.

Waterfall Transformer for Multi-person Pose Estimation 15 keypoints is all you need

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:45:41.158898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T10:45:40.832587Z digest=sha256:c122030cca473683bc20a66ed1b012e23e00a5a8a6195dfbfdca8ffd4214a236

Observation ef167870-978c-41a5-9cc3-17b8960760a3 · outbound

This paper cites Deep high-resolution representation learning for hu- man pose estimation.

Waterfall Transformer for Multi-person Pose Estimation Deep high-resolution representation learning for hu- man pose estimation

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:45:41.144722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T10:45:40.837133Z digest=sha256:ab6fb01340ebed38d0ec062eff251534da2faf31b0c2ab5c9c577baac942567b

Observation 850a7f89-1548-4479-a50a-d22121947c87 · outbound

This paper cites Deeply learned com- positional models for human pose estimation.

Waterfall Transformer for Multi-person Pose Estimation Deeply learned com- positional models for human pose estimation

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:45:41.129878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T10:45:40.841675Z digest=sha256:e5b5d317e07521b05c66496cb758ae2ca8b792b51c1855ec8ced86a3c006e278

Observation f00b7b4f-58f2-4949-aa41-eba39dbcf6d8 · outbound

This paper cites DeepPose: Human pose estimation via deep neural networks.

Waterfall Transformer for Multi-person Pose Estimation DeepPose: Human pose estimation via deep neural networks

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:45:41.115506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T10:45:40.846447Z digest=sha256:363be65a303d5e7d6b41de3e32b8bec2cd32c185136f5bb04a330d4a65b0cf16

Observation b6d9ea08-9453-47e2-9a25-da2b23b31e43 · outbound

This paper cites Attention is all you need.

Waterfall Transformer for Multi-person Pose Estimation Attention is all you need

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T10:45:40.851188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:45:40.851188Z digest=sha256:f6fec4b03cd3812238a06e0b2f82ed27a0f6ffac5f38aa9b78061f21860ff53b

Observation df18c7d0-eed9-413a-ab73-a4f08c1d7a0e · outbound

This paper cites Deep high- resolution representation learning for visual recogni- tion.

Waterfall Transformer for Multi-person Pose Estimation Deep high- resolution representation learning for visual recogni- tion

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:45:41.089839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T10:45:40.855807Z digest=sha256:f75ce0087e2640fb9ec27efcb3241c29ecde5e373566816185e0dd299161e784

Observation 106afd55-cb4b-4ac3-ba3a-149ba594317e · outbound

This paper cites Convolutional pose machines.

Waterfall Transformer for Multi-person Pose Estimation Convolutional pose machines

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:45:41.074381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T10:45:40.860167Z digest=sha256:9442b5f2c874bf768cf447b061c17a0f52b08485c595712a7b10833a7d919bcd

Observation a2362cc5-7c70-4a12-b54e-ec1b8cacb17c · outbound

This paper cites ViTPose: Simple Vision Transformer Baselines for Human Pose Estimation.

Waterfall Transformer for Multi-person Pose Estimation ViTPose: Simple Vision Transformer Baselines for Human Pose Estimation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T10:45:40.864432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:45:40.864432Z digest=sha256:b6cad987705562d87f408427b33c2eb729b98c554c25afe0f4a68b11038418ad

Observation 1d8b589b-af4d-4921-8095-9f669c78bed3 · outbound

This paper cites TransPose: Keypoint localization via transformer.

Waterfall Transformer for Multi-person Pose Estimation TransPose: Keypoint localization via transformer

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:45:41.060053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T10:45:40.869277Z digest=sha256:769dda7b8e9e005b987748c6af3550a82f326484e010ed051859d920eab1fdb9

Observation d48e2c71-5285-4b0b-9c18-d769ce02cad0 · outbound

This paper cites HRFormer: High-resolution vision transformer for dense predict.

Waterfall Transformer for Multi-person Pose Estimation HRFormer: High-resolution vision transformer for dense predict

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:45:41.044908Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T10:45:40.873458Z digest=sha256:4202a8c1f3cd7f74786d1c666b68cfe8e746815c04ccff1dab5503e16c400436

Observation a88c9097-a826-491e-9fbd-0f3d5403951a · outbound

This paper cites Human Pose Estimation with Spatial Contextual Information.

Waterfall Transformer for Multi-person Pose Estimation Human Pose Estimation with Spatial Contextual Information

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T10:45:40.877764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:45:40.877764Z digest=sha256:8b84683c3973729a3e03fca43b9814750c09fc641d9957fdd568e7d238772f05

Observation 52b4b00b-bf0e-453b-b94e-51f4350897ed · outbound

This paper cites 3d human pose estimation with spatial and temporal transform- ers.

Waterfall Transformer for Multi-person Pose Estimation 3d human pose estimation with spatial and temporal transform- ers

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:45:41.030614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T10:45:40.882265Z digest=sha256:84ce53c220386c0429809af3e5a933bb91ef086c2a9ad7f91cd104b89e44716b

Pith citing papers

No inbound Pith citation observations are available.