Pith. sign in

Paper Citation Record · LEDGER

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model

As of 7 August 2026, this Paper Citation Record lists 100 of 108 outbound references and 0 inbound Pith citation observations for arXiv:2507.11102.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.11102 v1

Coverage vector

measured 100 of 108 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T17:22:15.365213Z

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

100 of 108 outbound references displayed

  • verified exact1
  • verified fuzzy33
  • unresolved66
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1e575dc4-462e-4c0d-8792-da60801519e8 · outbound

This paper cites GPT-4 Technical Report.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T17:22:08.397914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:22:08.397914Z digest=sha256:d80e9ac1fbdba0bad966ff1482bc6101e8c4892991e6fc12dd121f4134a79430

Observation f21a9d29-8ac7-43af-8b26-d54d60991ba3 · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model Flamingo: a visual language model for few-shot learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T17:22:08.439737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:22:08.439737Z digest=sha256:981999304ef196ff9f21c165d3f7cbaea6fc234d10b70043f3711d65472b9bf8

Observation 7207cbbb-8cd1-42d0-b7aa-1c872101389e · outbound

This paper cites 2d human pose estimation: New benchmark and state of the art analysis.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model 2d human pose estimation: New benchmark and state of the art analysis

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T17:22:08.500514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:22:08.500514Z digest=sha256:31ff3c64a9c70f7c61cadc6a823d092c4900f536dadfc07634db30bedca71983

Observation 9af81807-c9b1-483e-a1f1-5bd556b02e07 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T17:22:08.549301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:22:08.549301Z digest=sha256:1fe3491a8ddbb91c81cb78966fc92a061b42d5c7f37402dc84006a15dc0e9275

Observation 52b6fe1e-6910-40d6-9465-debb48ae05f9 · outbound

This paper cites Language models are few-shot learners.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model Language models are few-shot learners

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T17:22:08.604576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:22:08.604576Z digest=sha256:317c1d3b64ac857b846b192ab96b18728ecc24c33c1f50fa728325a2c6f7a25a

Observation 3d5dcce1-3c34-4caf-bfc3-bc718347085b · outbound

This paper cites Cross-domain adaptation for animal pose estimation.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model Cross-domain adaptation for animal pose estimation

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T17:22:08.657396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:22:08.657396Z digest=sha256:ab09b143257f01b67cbd8c19020288a0a9d6353bbf619d1069b8ad0b612f8cdc

Observation 722e2f4f-6d8a-4e55-bd06-43be25a962dc · outbound

This paper cites Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T17:22:08.708494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:22:08.708494Z digest=sha256:0ca49a446bded2e9b102db70755aa4ae3b91fb800dcdc628f505c8590a84d223

Observation 41e14b63-2c7c-442a-a5df-bdb37a9e620a · outbound

This paper cites MotionLLM: Understanding Human Behaviors from Human Motions and Videos.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model MotionLLM: Understanding Human Behaviors from Human Motions and Videos

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T17:22:08.784420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:22:08.784420Z digest=sha256:dd9b46b4b7dde71236a3305b012346ae7dc3a62ed43af94d3833627104a9483d

Observation df60423c-1372-4f5a-a04d-352ff4cf9608 · outbound

This paper cites Cascaded pyramid network for multi-person pose estimation.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model Cascaded pyramid network for multi-person pose estimation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T17:22:08.873244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:22:08.873244Z digest=sha256:158b0f5b4e84c1d5bb3a2ad818c0903fa5057b4cafce242fe5abc97cd1a0fae9

Observation ad8c833b-17f8-450f-87cb-5f33596beff0 · outbound

This paper cites Higherhrnet: Scale-aware representation learning for bottom-up human pose estimation.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model Higherhrnet: Scale-aware representation learning for bottom-up human pose estimation

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T17:22:08.937648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:22:08.937648Z digest=sha256:6a14029a9fb073308ba63782ad6553c08e7bb1c3ff1a02823ccc71967c365473

Observation adab5577-926b-4db0-a236-928264f079de · outbound

This paper cites Palm: Scaling language modeling with pathways.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model Palm: Scaling language modeling with pathways

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T17:22:09.022376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:22:09.022376Z digest=sha256:533d756b52b88bc045809029c07e36b23ad6aedb6523f4f34feffa0e5930d2ba

Observation 81eae217-1435-45f6-aa0d-867eeb2773ad · outbound

This paper cites Instructblip: Towards general-purpose vision-language models with instruction tuning.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model Instructblip: Towards general-purpose vision-language models with instruction tuning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T17:22:09.091426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:22:09.091426Z digest=sha256:cec55e27dd1f93d17756d869a05a0fccbdf4e78c43e889574a0da5217c0f6f3e

Observation c30815ce-5e4e-4f70-85a7-6f2e35045019 · outbound

This paper cites Model-agnostic meta-learning for fast adaptation of deep networks.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model Model-agnostic meta-learning for fast adaptation of deep networks

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T17:22:09.153335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:22:09.153335Z digest=sha256:38b77f050ee1e88871bc55834a3272a83abd2f22bd551b1cb9ab96041cf8eaab

Observation 4273fe13-5bf9-4cbb-9c21-187d885d7cef · outbound

This paper cites Deepfashion2: A versatile benchmark for detection, pose estimation, segmentation and re-identification of clothing images.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model Deepfashion2: A versatile benchmark for detection, pose estimation, segmentation and re-identification of clothing images

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T17:22:09.221988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:22:09.221988Z digest=sha256:49263058dd3aff4f219533a1b4119a0bdf722048559cda007490ae4480eb5fb3

Observation 63dbaaf5-b091-4f5f-912c-d04fd5006345 · outbound

This paper cites LivePortrait: Efficient Portrait Animation with Stitching and Retargeting Control.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model LivePortrait: Efficient Portrait Animation with Stitching and Retargeting Control

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T17:22:09.273146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:22:09.273146Z digest=sha256:a804965713bbcc02847adea9810e1a58db698d910514a111bb8ba87e2a7c63c8

Observation 66112164-54e7-471d-afcc-d52f7409b51f · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model LoRA: Low-Rank Adaptation of Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T17:22:09.347197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:22:09.347197Z digest=sha256:920f5d6ddc775f06c6e82550380177f1b29d67259a2efde0234f95e8c4e98867

Observation e20c208d-7fd7-46dc-86ac-a750eb9ee0c6 · outbound

This paper cites Mistral 7B.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model Mistral 7B

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T17:22:09.421311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:22:09.421311Z digest=sha256:516dbec17548955becbfca66515a1e1c058a830152b4fbe81670681e376304e2

Observation edc05f21-d471-49ec-9075-97a12da249c4 · outbound

This paper cites Multi-person articulated tracking with spatial and temporal embeddings.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model Multi-person articulated tracking with spatial and temporal embeddings

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T17:22:09.503131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:22:09.503131Z digest=sha256:c2ca0c1ac722485357a08974d6e2f72b38f4391ea962fc6a88668d177ce79dc9

Observation 98051cca-7b82-4cd6-85fa-ce489709bbf6 · outbound

This paper cites Differentiable hierarchical graph grouping for multi-person pose estimation.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model Differentiable hierarchical graph grouping for multi-person pose estimation

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T17:22:09.546326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:22:09.546326Z digest=sha256:f5d555fa0f98eb381e03928bd0112f0e6ef9165e5e56ec8bd8e8f3ea323b4f23

Observation 28d9a720-b559-4054-85c2-507a171dcc03 · outbound

This paper cites Whole-body human pose estimation in the wild.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model Whole-body human pose estimation in the wild

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T17:22:09.607010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:22:09.607010Z digest=sha256:bf49e6de6ed7497b0df3995851d6932cde6f323d08b3672eb43518b80f9cdb1a

Observation acb0fec6-1929-4787-8c39-1d07c9d5ca6f · outbound

This paper cites Human-art: A versatile human-centric dataset bridging natural and artificial scenes.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model Human-art: A versatile human-centric dataset bridging natural and artificial scenes

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T17:22:09.665516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:22:09.665516Z digest=sha256:fd1b0fd9e57ff8ce3b033355aa6b705bd5c116147d0600cb0aede7d6a1b0395a

Observation aa15c8a4-e594-438d-89e3-6b7be211d618 · outbound

This paper cites Humansd: A native skeleton-guided diffusion model for human image generation.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model Humansd: A native skeleton-guided diffusion model for human image generation

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T17:22:09.732831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:22:09.732831Z digest=sha256:7445f54ec6954fddfad832749249fc80ed15417de246f18792e2e4339126f5c2

Observation 19cfd8aa-c4c9-4adf-aa5a-6c3e1a57a06f · outbound

This paper cites Animalweb: A large-scale hierarchical dataset of annotated animal faces.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model Animalweb: A large-scale hierarchical dataset of annotated animal faces

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T17:22:09.795152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:22:09.795152Z digest=sha256:a417ebc70b7431ce9900e69622d9376b49f344cf73a0bdbf0f082611d6425f07

Observation 965e575e-c9ea-4af7-b542-be3f8a7b3ca7 · outbound

This paper cites Segment anything.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model Segment anything

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T17:22:09.853259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:22:09.853259Z digest=sha256:9260070d397724258582af4c8fb301029ebd97bf8a1a665f0ab6cdfee4b45138

Observation 3e53591d-314e-4c55-a300-da9b080d313c · outbound

This paper cites in the wild.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model in the wild

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T17:22:09.915013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:22:09.915013Z digest=sha256:57649af6fd9be3a26057cc9222c6f74775216c5742fc14020efbd630594a0ba8

Observation 096df928-9178-48e5-8a28-324a82b4db17 · outbound

This paper cites LISA: Reasoning Segmentation via Large Language Model.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model LISA: Reasoning Segmentation via Large Language Model

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T17:22:09.963924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:22:09.963924Z digest=sha256:61da29557167dc188b8df0c531606557f0c42bfef8252fdc66ff116ce489ab46

Observation 41980a60-03c0-424b-914f-2f2352626e52 · outbound

This paper cites What matters when building vision-language models?.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model What matters when building vision-language models?

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T17:22:10.040901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:22:10.040901Z digest=sha256:05dd61fdf288f1c04531d90fe85974f0827825cc1eae45d2f85695c296d4e8a6

Observation e5a426ab-7b91-4d0e-8765-35979addef97 · outbound

This paper cites MIMIC-IT: Multi-Modal In-Context Instruction Tuning.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model MIMIC-IT: Multi-Modal In-Context Instruction Tuning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T17:22:10.082980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:22:10.082980Z digest=sha256:dfce20260858b1ba39775e16ff6b5945cac0192b06297018b6986ca2c57d57a5

Observation 0831afbf-fc6c-4d6d-a7f9-89bacda92b36 · outbound

This paper cites Crowdpose: Efficient crowded scenes pose estimation and a new benchmark.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model Crowdpose: Efficient crowded scenes pose estimation and a new benchmark

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T17:22:10.140423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:22:10.140423Z digest=sha256:f95188b12605857e43a136b821533695322cb985346865aa1b35d706154c4887

Observation 35498739-a94d-48f3-866a-72b9c207bf4d · outbound

This paper cites Human pose regression with residual log-likelihood estimation.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model Human pose regression with residual log-likelihood estimation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T17:22:10.212644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:22:10.212644Z digest=sha256:a5cca11152770d1312d23803e70062f933a429cb2d82166ed58e607036a09d12

Observation 7becd0fc-4cd5-4589-a604-332897af47e1 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T17:22:10.277017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:22:10.277017Z digest=sha256:3d9a002977fa320c4b54e5cf7a7f6cfa4b1dbbcf7b3e216abce6a3e190039f00

Observation aca1c1be-834d-41ae-aeb5-8c43a9a3e5ca · outbound

This paper cites TokenPose: Learning Keypoint Tokens for Human Pose Estimation.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model TokenPose: Learning Keypoint Tokens for Human Pose Estimation

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-08-06T17:22:16.549635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T17:22:10.347059Z digest=sha256:48a3f88c3a9ca82fb128b76b3055088829c9f3f53f02f2dd9c802378984e1903

Observation 891c5c39-6db8-43bc-8402-b29843605d37 · outbound

This paper cites Simcc: A simple coordinate classification perspective for human pose estimation.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model Simcc: A simple coordinate classification perspective for human pose estimation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T17:22:10.392295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:22:10.392295Z digest=sha256:2b638d7632207a1a0ebcce9c3616433eefcad1e6d3d99239c73ae6b3c3ef7534

Observation 6069bdca-5dab-4ef4-a61d-a758b62f7f20 · outbound

This paper cites Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T17:22:10.455266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:22:10.455266Z digest=sha256:6d40ca87a9333ba9bff971f15b4df6640fd3a38c086e10d12e58a0398c11fed8

Observation 08768bc4-df80-43f8-a3c7-0cdc7de55359 · outbound

This paper cites VILA: On Pre-training for Visual Language Models.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model VILA: On Pre-training for Visual Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T17:22:10.522038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:22:10.522038Z digest=sha256:9c9f702639b337c225800f2ef6e2b14748e340b8cff2237504a4dbd307750e05

Observation eefea3c9-1154-4999-9ecf-ae97a6a6659e · outbound

This paper cites Microsoft coco: Common objects in context.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model Microsoft coco: Common objects in context

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T17:22:10.573970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:22:10.573970Z digest=sha256:0f07e323525c29a68467d07a41d71013d46a3b572e9f721c306e7557747911d6

Observation 29dfe3b6-5bb1-4e61-a502-9567b7b7ff88 · outbound

This paper cites Improved baselines with visual instruction tuning, 2023 a.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model Improved baselines with visual instruction tuning, 2023 a

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T17:22:10.659686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:22:10.659686Z digest=sha256:f83b848b0a93508dcb16ce0a10a1b9ffd2fbfbf9cf0900598a25ec991439a18d

Observation ae8ffabf-0df9-44f0-8646-2f4000136edc · outbound

This paper cites Visual instruction tuning, 2023 b.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model Visual instruction tuning, 2023 b

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T17:22:10.712260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:22:10.712260Z digest=sha256:fefa22d6c4146cdaf711384e6c62038a7665fae31de70e69274e3849f0c79c57

Observation c9756c63-7999-4256-b83d-8bb282882a66 · outbound

This paper cites Llava-next: Improved reasoning, ocr, and world knowledge, January 2024.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model Llava-next: Improved reasoning, ocr, and world knowledge, January 2024

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T17:22:10.756326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:22:10.756326Z digest=sha256:e87fbb45c839b2264ddac30bd2eb17c86825ba425cb1395c682e84280c739659

Observation f30433d3-3016-47b7-90e4-eea774095811 · outbound

This paper cites Grounding dino: Marrying dino with grounded pre-training for open-set object detection.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model Grounding dino: Marrying dino with grounded pre-training for open-set object detection

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T17:22:10.818362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:22:10.818362Z digest=sha256:4e2d7f3dcc523e8d41ce09cc3441dfc39dae9d00e3746fbc53b3c43ac9eb1253

Observation daa8169b-1f48-4de4-ae48-fff91bc58e1a · outbound

This paper cites A convnet for the 2020s.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model A convnet for the 2020s

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T17:22:10.875282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:22:10.875282Z digest=sha256:a749e5ba978e3355a1679ae6b45f933510400258188c25c993bfbdc0b1609c07

Observation 939d3dcc-b1d7-4e5e-b825-3378cd0e8548 · outbound

This paper cites Deepseek-vl: Towards real-world vision-language understanding, 2024.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model Deepseek-vl: Towards real-world vision-language understanding, 2024

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T17:22:10.920288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:22:10.920288Z digest=sha256:0e72d4ac9f09ac6d1bbdd1d5a57db38ce1c12dd17667c6e1e937063b7f0449c0

Observation 6e3305b5-61ea-4571-a060-fea05536bf93 · outbound

This paper cites HumanTOMATO: Text-aligned Whole-body Motion Generation.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model HumanTOMATO: Text-aligned Whole-body Motion Generation

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T17:22:10.990900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:22:10.990900Z digest=sha256:d4994007d9349ab2f5505889401b909aeeb099fad8a38821ae16bce0b9842721

Observation b6d07928-854b-4ce8-931b-60c564c9f891 · outbound

This paper cites From keypoints to object landmarks via self-training correspondence: A novel approach to unsupervised landmark discovery.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model From keypoints to object landmarks via self-training correspondence: A novel approach to unsupervised landmark discovery

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:22:22.693855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T17:22:10.997048Z digest=sha256:2f97f85da7968828de432bda59af445578ce408d8da9f03cfb0ceec3dd573279

Observation a5b53bb2-96bb-493c-b041-ddd1971f8bb5 · outbound

This paper cites MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T17:22:11.164532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:22:11.164532Z digest=sha256:a315b7440e0282ae28c89497a73fc9ce76d5c08b0ba34b9c23ecbec78d742e60

Observation 37e83a97-d2bc-4ebe-8f31-31181833f824 · outbound

This paper cites Gemma: Open Models Based on Gemini Research and Technology.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model Gemma: Open Models Based on Gemini Research and Technology

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T17:22:11.281720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:22:11.281720Z digest=sha256:9f9df5c960e661a72deb5194dda88f1676af7f3eab85c17aaa75838ddb94fb6d

Observation f3bb7e68-4f45-46ec-bb30-0a496f696331 · outbound

This paper cites Interhand2.6m: A dataset and baseline for 3d interacting hand pose estimation from a single rgb image.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model Interhand2.6m: A dataset and baseline for 3d interacting hand pose estimation from a single rgb image

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:22:22.522833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T17:22:11.337039Z digest=sha256:06f56664cd46005774c3061771a9d5d626bc79ea555a274ab75803d438cf6aec

Observation 9dba3041-d485-46d1-ab8e-c9ffaa0219d1 · outbound

This paper cites Revisiting Fine-tuning for Few-shot Learning.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model Revisiting Fine-tuning for Few-shot Learning

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T17:22:11.496851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:22:11.496851Z digest=sha256:d6e482ef1df6a6765e3d125a9e126cc429d2297f5c5e23adb376098a4c2f7899

Observation dc1ef1a2-29a6-48dd-92fa-dec6eceb1cdd · outbound

This paper cites Stacked hourglass networks for human pose estimation.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model Stacked hourglass networks for human pose estimation

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:22:22.402703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T17:22:11.582583Z digest=sha256:03bbaedbccf39d466ca415e89bcb001de05e4ccd7e40ed52f6152c299d5b289a

Observation f6cdb89f-0137-4137-9b5e-dcc93f6750bb · outbound

This paper cites Animal kingdom: A large and diverse dataset for animal behavior understanding.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model Animal kingdom: A large and diverse dataset for animal behavior understanding

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:22:22.305728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T17:22:11.692051Z digest=sha256:d8c57446f0cdf459fb11baba52c1afaf30422c8a2f5b87de8410e05817cba2c2

Observation 9bacbe1e-dc12-4098-b34a-28ed5de5a620 · outbound

This paper cites Single-stage multi-person pose machines.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model Single-stage multi-person pose machines

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:22:22.202969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T17:22:11.803010Z digest=sha256:466a8201911e4b2c9f0801c29d94803bacb3994af29f31ea4e486d00e78bb6d2

Observation cec821b3-8e61-44ef-94e5-0bc0296fc372 · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model DINOv2: Learning Robust Visual Features without Supervision

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T17:22:11.891931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:22:11.891931Z digest=sha256:f32830fbf58304fe8d9fc90e71666d9e5b0acd45f1325f717d989f72f5defcde

Observation d5bacdeb-da28-496e-948a-7b8e26a9a10b · outbound

This paper cites Instruction Tuning with GPT-4.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model Instruction Tuning with GPT-4

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T17:22:12.010239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:22:12.010239Z digest=sha256:fa35d394e379213f9c8d306d51d48cc8fc0afd35e9bb26de294fd0bf20d353fd

Observation 2f35e897-2827-487e-a2d1-71341e8baec7 · outbound

This paper cites Kosmos-2: Grounding Multimodal Large Language Models to the World.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model Kosmos-2: Grounding Multimodal Large Language Models to the World

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T17:22:12.091020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:22:12.091020Z digest=sha256:c70ae067f53ac6d8a4f2f2c31e5bd552588695d25110445801874ffc54395129

Observation f0717450-6257-4e36-8e3a-1b0e092292d2 · outbound

This paper cites DetGPT: Detect What You Need via Reasoning.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model DetGPT: Detect What You Need via Reasoning

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T17:22:12.156699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:22:12.156699Z digest=sha256:cdb34f396307cd8c9f44bd7e706bfa9b7ce6af79a45bfbee6e8482d42ade32a1

Observation ac7378b4-d987-40b7-aad6-4158d99f5e3e · outbound

This paper cites PerceptionGPT: Effectively Fusing Visual Perception into LLM.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model PerceptionGPT: Effectively Fusing Visual Perception into LLM

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T17:22:12.222475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:22:12.222475Z digest=sha256:1d9c6ff0f1c0e6a37ecf919b24e1abaa633aaf1d49aba956a0e2411eeff943bc

Observation 00f6b4e1-e771-444b-8855-7ade893bbde3 · outbound

This paper cites Learning transferable visual models from natural language supervision.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model Learning transferable visual models from natural language supervision

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:22:22.125639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T17:22:12.301281Z digest=sha256:9404153d6b1b3ccd6e76d2b8c16b5e1514a6311a440cf6e9a5fbea6c8400961b

Observation ab397d07-183e-4853-b753-63ea015c68a6 · outbound

This paper cites Carfusion: Combining point tracking and part detection for dynamic 3d reconstruction of vehicles.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model Carfusion: Combining point tracking and part detection for dynamic 3d reconstruction of vehicles

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:22:22.053927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T17:22:12.360012Z digest=sha256:fdb3ed71fb8eef6fc8bace170014c01e9ce5693e5bfb2e7fc499b2a6642e40f9

Observation c809e9d7-db83-40d6-82ee-37b313a06f21 · outbound

This paper cites Zafeiriou, and M.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model Zafeiriou, and M

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:22:21.902725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T17:22:12.443453Z digest=sha256:efff755b987a529097f14c98630e65e3659974b8c722188c97c9be08134c557b

Observation 193d7a0f-d084-44a4-b50e-334e49b1ad21 · outbound

This paper cites Matching is not enough: A two-stage framework for category-agnostic pose estimation.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model Matching is not enough: A two-stage framework for category-agnostic pose estimation

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:22:21.697796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T17:22:12.518555Z digest=sha256:f0663a27f4abc3b119d3872b19ec1eeb048689894a573e8d87d05fc7d829d56f

Observation 0044afed-fb6e-4f5d-92b4-7bd38c1191e3 · outbound

This paper cites Prototypical networks for few-shot learning.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model Prototypical networks for few-shot learning

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:22:21.562706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T17:22:12.586243Z digest=sha256:9a4579e31ea413dfbffe9f4fac4447bcebbaf4cdc786d9b26e9babc204662903

Observation d707b8cd-7738-4732-931a-b313d78d12ae · outbound

This paper cites Self-supervised keypoint discovery in behavioral videos.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model Self-supervised keypoint discovery in behavioral videos

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:22:21.473712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T17:22:12.645354Z digest=sha256:ddafe1ad4405f90423efe4fee23000af3fb93bf5ac2ba479bcdcfbd427eb21d9

Observation 4a5fb059-5cab-44c9-8db4-85a02e3febdd · outbound

This paper cites Deep high-resolution representation learning for human pose estimation.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model Deep high-resolution representation learning for human pose estimation

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:22:21.322446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T17:22:12.715825Z digest=sha256:021e9e12445fc0a8776f139f896cf3601e357621b66689edc061f79228c9ece9

Observation 9d19d238-f1ea-4518-a951-d71afbf21c20 · outbound

This paper cites Compositional human pose regression.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model Compositional human pose regression

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:22:21.212839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T17:22:12.772097Z digest=sha256:6c156d855e805218baf369817985d3f77789b9c19a48490f7bc0e59a0fa4c63b

Observation 020a2b09-2843-435d-96e9-b60b71384085 · outbound

This paper cites Deeppose: Human pose estimation via deep neural networks.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model Deeppose: Human pose estimation via deep neural networks

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:22:20.996690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T17:22:12.848054Z digest=sha256:5b07048045f0ed5bedb879b7a42fe4a6d9bdbf2fff0f1c6c1064aec6485aebc3

Observation 11fd6029-7466-4628-9867-b98b512596c8 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model LLaMA: Open and Efficient Foundation Language Models

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-06T17:22:12.913739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:22:12.913739Z digest=sha256:ffe903f75a9dcc1e5573ab4e012795eb8e5d4de20e9c515fa178debfaed142a9

Observation ab97ce07-9584-4586-8107-d90835be4ba1 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-06T17:22:13.017689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:22:13.017689Z digest=sha256:b3b6c8afc79a3108af91049b790a6c8fc75b6f76c3d8d9f241ee97e3b508ea37

Observation c6e17968-63d8-4058-99ed-d85d64764396 · outbound

This paper cites Attention is all you need.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model Attention is all you need

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:22:20.889936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T17:22:13.097863Z digest=sha256:3d13a60da4ff1438a2b82b1b173fcf360f253de0dfbe9d8e9956db1d95f6a466

Observation 1682b3c2-1b56-4d90-954f-4680153e4297 · outbound

This paper cites Locllm: Exploiting generalizable human keypoint localization via large language model.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model Locllm: Exploiting generalizable human keypoint localization via large language model

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:22:20.760112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T17:22:13.162173Z digest=sha256:9e8bb51c75294aa0170722608877abab6a935c495d28c760686558b78cce4160

Observation 1093694e-ff0d-42cc-b21b-0d6b8cc879ef · outbound

This paper cites Visionllm: Large language model is also an open-ended decoder for vision-centric tasks.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model Visionllm: Large language model is also an open-ended decoder for vision-centric tasks

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:22:20.495122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T17:22:13.224523Z digest=sha256:6577ff5ee57b833dcf4455d0d8a676e48adf3b947111584f6493c9eb8ee4d719

Observation 54aec589-a372-47f7-a2a5-8c701f4d0232 · outbound

This paper cites Convolutional pose machines.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model Convolutional pose machines

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:22:20.249219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T17:22:13.325920Z digest=sha256:b3f3c3241e6490bf3624d60572f15a895a42bf28cdd48f86a3c78a8ccca9d656

Observation d43a7054-ee52-4ca0-8c0d-5cc729736004 · outbound

This paper cites VisionLLM v2: An End-to-End Generalist Multimodal Large Language Model for Hundreds of Vision-Language Tasks.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model VisionLLM v2: An End-to-End Generalist Multimodal Large Language Model for Hundreds of Vision-Language Tasks

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-06T17:22:13.366819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:22:13.366819Z digest=sha256:9ca980b5f7028580262834c38723bfb101a828ee17b9b3170d4a06848bcda32b

Observation 8029c7b6-d77d-49c4-a4dd-6d99ac6b047d · outbound

This paper cites F-LMM: Grounding Frozen Large Multimodal Models.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model F-LMM: Grounding Frozen Large Multimodal Models

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-06T17:22:13.424856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:22:13.424856Z digest=sha256:5ca08506d37ddc0513c4f9a3f75e23a089439783f65c03ec77e9a5790297f30b

Observation daa5e2e4-0e2d-460b-8b5b-bb93ef646c21 · outbound

This paper cites Look at boundary: A boundary-aware face alignment algorithm.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model Look at boundary: A boundary-aware face alignment algorithm

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:22:19.991106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T17:22:13.517966Z digest=sha256:8189c7f0303959fd6ba27a501a8d5b90849a893e0824a378f326ca48ad8a3282

Observation a8625e9d-9e14-4964-a116-79d40c55de5d · outbound

This paper cites Simple baselines for human pose estimation and tracking.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model Simple baselines for human pose estimation and tracking

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:22:19.816933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T17:22:13.609050Z digest=sha256:7ce2b7220d93da741714b505583ccac7d98fb269ffe44425d4bc6e61cade4039

Observation 4af29d8b-5f40-4ccb-91ba-6ff19c44019e · outbound

This paper cites Pixel-aligned language model.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model Pixel-aligned language model

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:22:19.613550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T17:22:13.656756Z digest=sha256:339ccd8cdecdc0973028676b3556dbd92ab4d766847c7da2018d80dca62531e9

Observation 1b189f69-544b-4dbb-b977-e3fc681299bf · outbound

This paper cites Vipnas: Efficient video pose estimation via neural architecture search.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model Vipnas: Efficient video pose estimation via neural architecture search

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:22:19.315519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T17:22:13.687889Z digest=sha256:d2311e46176894db461b9922cc4fe8476b870632e0b13a3b757c365a20ee31ec

Observation f51d56c7-31ae-4ef5-986b-2b22187b2ddc · outbound

This paper cites Pose for everything: Towards category-agnostic pose estimation.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model Pose for everything: Towards category-agnostic pose estimation

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:22:19.077852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T17:22:13.770149Z digest=sha256:8974fcfd55bf8cca588643b16ffc7e73c3eda3cd9fad1458e9776a981cdab2a9

Observation d665a2f6-f302-474a-9f64-27e72b9dfe03 · outbound

This paper cites Vitpose: Simple vision transformer baselines for human pose estimation.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model Vitpose: Simple vision transformer baselines for human pose estimation

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:22:18.920419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T17:22:13.820520Z digest=sha256:62dbbf37f09be30335e1ef9598f55181bf9c234e8e442a1a0cecb5fb359a3b6e

Observation 623d16ab-ef14-4aeb-8c7f-ac9bdd74423a · outbound

This paper cites Boosting Human-Object Interaction Detection with Text-to-Image Diffusion Model.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model Boosting Human-Object Interaction Detection with Text-to-Image Diffusion Model

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-06T17:22:13.857000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:22:13.857000Z digest=sha256:06287adaf1fd16ea1a9cf9dac1f5107670c5dd88cd51dbd7822a80c48a8e12c6

Observation e48fbddf-4a94-4d53-982f-5612c5b30cd0 · outbound

This paper cites Semantic human parsing via scalable semantic transfer over multiple label domains.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model Semantic human parsing via scalable semantic transfer over multiple label domains

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:22:18.797167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T17:22:13.935885Z digest=sha256:6202543566cc201af8c188cb293c6e889e1b2b8bd0987da0feecd85a991d67ea

Observation d92d8528-f971-418f-befa-a03528242823 · outbound

This paper cites Neural interactive keypoint detection.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model Neural interactive keypoint detection

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:22:18.660229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T17:22:13.997849Z digest=sha256:e72d41f1befdb77011c4b07af1431989926b108b79aeaeb6c090193d9d0b464b

Observation d6f0255e-3426-4a8b-b49c-dedc33e2547f · outbound

This paper cites Explicit Box Detection Unifies End-to-End Multi-Person Pose Estimation.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model Explicit Box Detection Unifies End-to-End Multi-Person Pose Estimation

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-06T17:22:14.093641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:22:14.093641Z digest=sha256:22c383aed3b96ddae44c02869aa9aaff8252cbeecf62f12d055f8041522d86fa

Observation 14404574-6dac-455e-868a-156d94e504ad · outbound

This paper cites X-Pose: Detecting Any Keypoints.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model X-Pose: Detecting Any Keypoints

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-06T17:22:14.151635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:22:14.151635Z digest=sha256:afec23a04fa273d603c9a062dca4f6e50b07521e808cb5d33b0403d9a9c5eae1

Observation fbdff743-f45d-4789-a08a-b3958541ca59 · outbound

This paper cites F-HOI: Toward Fine-grained Semantic-Aligned 3D Human-Object Interactions.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model F-HOI: Toward Fine-grained Semantic-Aligned 3D Human-Object Interactions

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-06T17:22:14.216807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:22:14.216807Z digest=sha256:b7d765bb9d349da02b7644a7b9d56d9ee4c27772ec38f7a215eb8c4b3b26e3f0

Observation 4ba0d7ec-2233-4007-aaac-5d034fe7b636 · outbound

This paper cites Kptllm: Unveiling the power of large language model for keypoint comprehension.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model Kptllm: Unveiling the power of large language model for keypoint comprehension

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:22:18.516527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T17:22:14.279793Z digest=sha256:9715eb4774281d149a722b98223c020f7fe777b975f56b10cf4b60451a40b963

Observation 99992b3d-816a-4194-87e7-bafe0445bb99 · outbound

This paper cites Ed-pose++: Enhanced explicit box detection for conventional and interactive multi-object keypoint detection.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model Ed-pose++: Enhanced explicit box detection for conventional and interactive multi-object keypoint detection

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:22:18.361369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T17:22:14.319317Z digest=sha256:9b68af3878d1e7bd91dd044ee94389e3f175dfe5db666f845611b5c9f22763dc

Observation c17f125e-98d2-4054-8b8b-f13350b71b81 · outbound

This paper cites Apt-36k: A large-scale benchmark for animal pose estimation and tracking.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model Apt-36k: A large-scale benchmark for animal pose estimation and tracking

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:22:18.199140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T17:22:14.386205Z digest=sha256:0c8f630e20668312f2466757426f81b97228dff03280e42265a1b07db6b2dcc0

Observation ec2be9e1-6a25-49e9-92ff-b38306b2aec1 · outbound

This paper cites mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-06T17:22:14.478484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:22:14.478484Z digest=sha256:54018c2f45b7473414c9931c846695e213d13f2df218609ba86f687b7c258c6d

Observation 53a9ac90-37f5-488d-9b45-56ab876e9f8c · outbound

This paper cites Ferret: Refer and Ground Anything Anywhere at Any Granularity.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model Ferret: Refer and Ground Anything Anywhere at Any Granularity

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-06T17:22:14.559435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:22:14.559435Z digest=sha256:36d0fc2e333581532851c81019ff4de5e6382c93646e788074f2712e05977128

Observation 4bc9b194-deaa-45f1-95cb-ef08bbd1047d · outbound

This paper cites AP-10K: A Benchmark for Animal Pose Estimation in the Wild.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model AP-10K: A Benchmark for Animal Pose Estimation in the Wild

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-06T17:22:14.658414Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:22:14.658414Z digest=sha256:b445f515c58182de3fa5472e8255f3afd895e26c4b99aec5e9a8d579cb8c4507

Observation 2d8289d9-32c5-4689-b2d0-f2de7b9836a2 · outbound

This paper cites HRFormer: High-Resolution Transformer for Dense Prediction.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model HRFormer: High-Resolution Transformer for Dense Prediction

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-06T17:22:14.730911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:22:14.730911Z digest=sha256:536ae3eed543ae2f6ff57a75af6b39c83abbcd8ca64a182465a42c6e065738c7

Observation 23390d0a-6e3d-4aa9-8eb2-366838c9140e · outbound

This paper cites Contextual Object Detection with Multimodal Large Language Models.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model Contextual Object Detection with Multimodal Large Language Models

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-06T17:22:14.824478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:22:14.824478Z digest=sha256:cc08c9a2957bd00f4a369c8d7f7aa0d9ae036e67248dcc31ecb11788ba210562

Observation 1378ab1f-f0eb-418a-bf65-923ae2222a24 · outbound

This paper cites Sigmoid loss for language image pre-training.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model Sigmoid loss for language image pre-training

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:22:18.057396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T17:22:14.952673Z digest=sha256:e8799746a7809c957fa24f29a1fe432f6caeccf4a58d1b1d37c1552d90e8c2fb

Observation f921b1a8-6eff-4729-a310-922fc4c6a742 · outbound

This paper cites Open-vocabulary animal keypoint detection with semantic-feature matching.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model Open-vocabulary animal keypoint detection with semantic-feature matching

Reference 95

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:22:17.934887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T17:22:15.020088Z digest=sha256:2cf8aaf2c5926d31e5006b6a4717606f521cb4711c6641e2790663d0aa98ab66

Observation 4447a533-086c-48e3-a756-4810421bc257 · outbound

This paper cites Adding conditional control to text-to-image diffusion models.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model Adding conditional control to text-to-image diffusion models

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:22:17.682866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T17:22:15.070625Z digest=sha256:908803563970d4b5f8b10c70f0525c304071b47a9b2982dd2381d89c12ade8b2

Observation 963fe659-f831-48ff-a10a-51ae5926795a · outbound

This paper cites GPT4RoI: Instruction Tuning Large Language Model on Region-of-Interest.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model GPT4RoI: Instruction Tuning Large Language Model on Region-of-Interest

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-06T17:22:15.142692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:22:15.142692Z digest=sha256:98a48bfb3d8fec8d4dbf1d03c5bab114cad6ef4fc6c8273211540d879bd052b2

Observation cb6a2be7-ccb1-4a1d-8b17-6225c4d86cb7 · outbound

This paper cites OPT: Open Pre-trained Transformer Language Models.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model OPT: Open Pre-trained Transformer Language Models

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-06T17:22:15.251884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:22:15.251884Z digest=sha256:f022fb8e18cf2b1bed2b35261987a53d80dd7c81477d5646af631a0233a58208

Observation 54596103-3b07-4237-954f-a2cad0f014e8 · outbound

This paper cites Clamp: Prompt-based contrastive learning for connecting language and animal pose.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model Clamp: Prompt-based contrastive learning for connecting language and animal pose

Reference 99

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:22:17.428420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T17:22:15.316629Z digest=sha256:3d76d3c22473171f4826093e3426a0769af8de06805a50c687d094765ea4de9d

Observation f0e8c4af-a326-435b-ba64-3a26311a833e · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-06T17:22:15.365213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:22:15.365213Z digest=sha256:5d421d0c307648f16ab310bad8e26797c694bdd6d21486300dc9f005335f684b

Pith citing papers

No inbound Pith citation observations are available.