Pith. sign in

Paper Citation Record · LEDGER

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model

As of 17 August 2026, this Paper Citation Record lists 100 of 108 outbound references and 0 inbound Pith citation observations for arXiv:2507.11102.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.11102 v1

Coverage vector

measured 100 of 108 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T17:22:15.365213Z

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

100 of 108 outbound references displayed

  • verified exact1
  • verified fuzzy33
  • unresolved66
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1e575dc4-462e-4c0d-8792-da60801519e8 · outbound

This paper cites GPT-4 Technical Report.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T17:22:08.397914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:22:08.397914Z digest=sha256:fcb63501f3553e25b2cebcdcedba2e05e3efe367bb40b5c6998e5e6cbb9ca6b8

Observation f21a9d29-8ac7-43af-8b26-d54d60991ba3 · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model Flamingo: a visual language model for few-shot learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T17:22:08.439737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:22:08.439737Z digest=sha256:1ea44a718948eba90f283e811519f3d67faccf7137775cf99fba30a6e58cc25b

Observation 7207cbbb-8cd1-42d0-b7aa-1c872101389e · outbound

This paper cites 2d human pose estimation: New benchmark and state of the art analysis.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model 2d human pose estimation: New benchmark and state of the art analysis

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T17:22:08.500514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:22:08.500514Z digest=sha256:3177edc8e6340cfdf82414973096f0cf077c12943a289b4649edfbca93e7d996

Observation 9af81807-c9b1-483e-a1f1-5bd556b02e07 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T17:22:08.549301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:22:08.549301Z digest=sha256:2d4aba799959ae853f5d66f92bde1b9152859d2210662639ef12ed51fd942036

Observation 52b6fe1e-6910-40d6-9465-debb48ae05f9 · outbound

This paper cites Language models are few-shot learners.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model Language models are few-shot learners

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T17:22:08.604576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:22:08.604576Z digest=sha256:67527a25a9b3a8b92c4c55e0464f430ff52f2c3159539cb17bc9438713f219a3

Observation 3d5dcce1-3c34-4caf-bfc3-bc718347085b · outbound

This paper cites Cross-domain adaptation for animal pose estimation.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model Cross-domain adaptation for animal pose estimation

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T17:22:08.657396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:22:08.657396Z digest=sha256:843dc52ed281742d0089de932e4f8318baca83945d4e28548986400e1b426c1b

Observation 722e2f4f-6d8a-4e55-bd06-43be25a962dc · outbound

This paper cites Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T17:22:08.708494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:22:08.708494Z digest=sha256:322dbaa45e041b02077ce00a7feed305984b7641324c048d315ade3ae00d19c2

Observation 41e14b63-2c7c-442a-a5df-bdb37a9e620a · outbound

This paper cites MotionLLM: Understanding Human Behaviors from Human Motions and Videos.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model MotionLLM: Understanding Human Behaviors from Human Motions and Videos

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T17:22:08.784420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:22:08.784420Z digest=sha256:8584f82c0465da21de0aafc487476e5bf21e4975656ad1e0eac8a20a9f49a102

Observation df60423c-1372-4f5a-a04d-352ff4cf9608 · outbound

This paper cites Cascaded pyramid network for multi-person pose estimation.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model Cascaded pyramid network for multi-person pose estimation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T17:22:08.873244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:22:08.873244Z digest=sha256:b7a71c95bf40a72887df5fafb1a21604833df57a8d7bbf2ea3af356cb738c8e8

Observation ad8c833b-17f8-450f-87cb-5f33596beff0 · outbound

This paper cites Higherhrnet: Scale-aware representation learning for bottom-up human pose estimation.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model Higherhrnet: Scale-aware representation learning for bottom-up human pose estimation

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T17:22:08.937648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:22:08.937648Z digest=sha256:04d5934bfe8421d2c06c499d78d826779d18b9803707813f37c2e7739709d4d1

Observation adab5577-926b-4db0-a236-928264f079de · outbound

This paper cites Palm: Scaling language modeling with pathways.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model Palm: Scaling language modeling with pathways

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T17:22:09.022376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:22:09.022376Z digest=sha256:ad74ab9e3f931ce6d0d6e449ee22307bebf8f7baa2c70924bdbb158146486fce

Observation 81eae217-1435-45f6-aa0d-867eeb2773ad · outbound

This paper cites Instructblip: Towards general-purpose vision-language models with instruction tuning.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model Instructblip: Towards general-purpose vision-language models with instruction tuning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T17:22:09.091426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:22:09.091426Z digest=sha256:d396f837a2b385ff06e23840a623eef6546ee14a45515db0b8428b1f30ed57fe

Observation c30815ce-5e4e-4f70-85a7-6f2e35045019 · outbound

This paper cites Model-agnostic meta-learning for fast adaptation of deep networks.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model Model-agnostic meta-learning for fast adaptation of deep networks

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T17:22:09.153335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:22:09.153335Z digest=sha256:5d8d30fa6e1e6619f1dc346c1bddd2c438282d6f99d0f7c8099beec49816be45

Observation 4273fe13-5bf9-4cbb-9c21-187d885d7cef · outbound

This paper cites Deepfashion2: A versatile benchmark for detection, pose estimation, segmentation and re-identification of clothing images.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model Deepfashion2: A versatile benchmark for detection, pose estimation, segmentation and re-identification of clothing images

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T17:22:09.221988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:22:09.221988Z digest=sha256:ce5894b343a19efc213b3d9557e67d407db66f17d19d2d124d81fff9a386e168

Observation 63dbaaf5-b091-4f5f-912c-d04fd5006345 · outbound

This paper cites LivePortrait: Efficient Portrait Animation with Stitching and Retargeting Control.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model LivePortrait: Efficient Portrait Animation with Stitching and Retargeting Control

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T17:22:09.273146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:22:09.273146Z digest=sha256:9317c7012e5ca8f8377fa37df930de89e6b865c844f8efa9039e6bfced24f2cc

Observation 66112164-54e7-471d-afcc-d52f7409b51f · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model LoRA: Low-Rank Adaptation of Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T17:22:09.347197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:22:09.347197Z digest=sha256:74c8bab0bbf292cac14fb838f963e0de416b7567342c37a559caa1e3e030263f

Observation e20c208d-7fd7-46dc-86ac-a750eb9ee0c6 · outbound

This paper cites Mistral 7B.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model Mistral 7B

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T17:22:09.421311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:22:09.421311Z digest=sha256:a8595b6dcb608c7bbe008ba523457cb54e50d3255a51e72aa5a50035fc07d4cd

Observation edc05f21-d471-49ec-9075-97a12da249c4 · outbound

This paper cites Multi-person articulated tracking with spatial and temporal embeddings.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model Multi-person articulated tracking with spatial and temporal embeddings

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T17:22:09.503131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:22:09.503131Z digest=sha256:9dfb627f7351ab573f6aed4ff6de5c105b4d01ee5116b20bed52a93984256f45

Observation 98051cca-7b82-4cd6-85fa-ce489709bbf6 · outbound

This paper cites Differentiable hierarchical graph grouping for multi-person pose estimation.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model Differentiable hierarchical graph grouping for multi-person pose estimation

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T17:22:09.546326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:22:09.546326Z digest=sha256:6e3d0bfdf2aa743acb78fa5fe110ba2538165cf5012145c90d252b0add526f36

Observation 28d9a720-b559-4054-85c2-507a171dcc03 · outbound

This paper cites Whole-body human pose estimation in the wild.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model Whole-body human pose estimation in the wild

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T17:22:09.607010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:22:09.607010Z digest=sha256:4eb434a4cfaff50e31b4fe8dfacb3616b1fc18088f1ee12e3d75caa1a50e21ba

Observation acb0fec6-1929-4787-8c39-1d07c9d5ca6f · outbound

This paper cites Human-art: A versatile human-centric dataset bridging natural and artificial scenes.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model Human-art: A versatile human-centric dataset bridging natural and artificial scenes

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T17:22:09.665516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:22:09.665516Z digest=sha256:c2c2082c5928a034fcd9b16b9caaba341e22296e5a878894f1a7f686da6f51d0

Observation aa15c8a4-e594-438d-89e3-6b7be211d618 · outbound

This paper cites Humansd: A native skeleton-guided diffusion model for human image generation.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model Humansd: A native skeleton-guided diffusion model for human image generation

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T17:22:09.732831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:22:09.732831Z digest=sha256:306fdb202e2ad058082169eeb6258cb3742bd271773ec2ff939270b104d09084

Observation 19cfd8aa-c4c9-4adf-aa5a-6c3e1a57a06f · outbound

This paper cites Animalweb: A large-scale hierarchical dataset of annotated animal faces.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model Animalweb: A large-scale hierarchical dataset of annotated animal faces

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T17:22:09.795152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:22:09.795152Z digest=sha256:6f5e4e1263da2ce2fa77cc5ec41d0e85ebc6b35c2804f8bb713392bb89862e90

Observation 965e575e-c9ea-4af7-b542-be3f8a7b3ca7 · outbound

This paper cites Segment anything.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model Segment anything

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T17:22:09.853259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:22:09.853259Z digest=sha256:f9abf6382b212687336482dc4d77e90ca972e732c806b3f0c3f8bc9f473de9b8

Observation 3e53591d-314e-4c55-a300-da9b080d313c · outbound

This paper cites in the wild.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model in the wild

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T17:22:09.915013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:22:09.915013Z digest=sha256:d4172b2d04a95787fd2c641212c30067c42c49b0b81c8c2bf16c8ccf94401e66

Observation 096df928-9178-48e5-8a28-324a82b4db17 · outbound

This paper cites LISA: Reasoning Segmentation via Large Language Model.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model LISA: Reasoning Segmentation via Large Language Model

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T17:22:09.963924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:22:09.963924Z digest=sha256:170bb0e6aeaa3ba4cd941f5c0e3cfe1b1002151112939b21e8df3f34d060779b

Observation 41980a60-03c0-424b-914f-2f2352626e52 · outbound

This paper cites What matters when building vision-language models?.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model What matters when building vision-language models?

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T17:22:10.040901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:22:10.040901Z digest=sha256:6a78d68ec31260fcd327eba32f5eb1fc2c545df1f36867ff19efe9e3969d66b4

Observation e5a426ab-7b91-4d0e-8765-35979addef97 · outbound

This paper cites MIMIC-IT: Multi-Modal In-Context Instruction Tuning.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model MIMIC-IT: Multi-Modal In-Context Instruction Tuning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T17:22:10.082980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:22:10.082980Z digest=sha256:d173afc1857ff4eb0863c84f157763c399ad48b2996189cf5ace00e178deb9e5

Observation 0831afbf-fc6c-4d6d-a7f9-89bacda92b36 · outbound

This paper cites Crowdpose: Efficient crowded scenes pose estimation and a new benchmark.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model Crowdpose: Efficient crowded scenes pose estimation and a new benchmark

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T17:22:10.140423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:22:10.140423Z digest=sha256:2b69acc68e7c13db879253fa0a8a5ce39d18431b547664f0715b51510adea75d

Observation 35498739-a94d-48f3-866a-72b9c207bf4d · outbound

This paper cites Human pose regression with residual log-likelihood estimation.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model Human pose regression with residual log-likelihood estimation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T17:22:10.212644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:22:10.212644Z digest=sha256:5eb06429843260372ab7c30972c895ce908c6da17588db133ba31687bf3e48c7

Observation 7becd0fc-4cd5-4589-a604-332897af47e1 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T17:22:10.277017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:22:10.277017Z digest=sha256:3adbaa86dde4af8398db22f6cbdf9b31054f7ce10f63a2e1165856950a9d33df

Observation aca1c1be-834d-41ae-aeb5-8c43a9a3e5ca · outbound

This paper cites TokenPose: Learning Keypoint Tokens for Human Pose Estimation.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model TokenPose: Learning Keypoint Tokens for Human Pose Estimation

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-08-06T17:22:16.549635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T17:22:10.347059Z digest=sha256:32596a5f805dcfb6328f62e55ff1a9e0a9b46c7e8cb17a8703248ff3a970d512

Observation 891c5c39-6db8-43bc-8402-b29843605d37 · outbound

This paper cites Simcc: A simple coordinate classification perspective for human pose estimation.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model Simcc: A simple coordinate classification perspective for human pose estimation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T17:22:10.392295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:22:10.392295Z digest=sha256:24e84aeeb898aed84d7cec6c8ff98491a1c5f17e6ea9d6e2c073028e73098fc4

Observation 6069bdca-5dab-4ef4-a61d-a758b62f7f20 · outbound

This paper cites Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T17:22:10.455266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:22:10.455266Z digest=sha256:98c7e54080b54823fde8fdd435fdf8d904031655190744e60a53cd9247a5f42e

Observation 08768bc4-df80-43f8-a3c7-0cdc7de55359 · outbound

This paper cites VILA: On Pre-training for Visual Language Models.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model VILA: On Pre-training for Visual Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T17:22:10.522038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:22:10.522038Z digest=sha256:63b0a3d8f369d5ba1b47456be56f5d52931a074f4c6b2cef594c78d3616e18d2

Observation eefea3c9-1154-4999-9ecf-ae97a6a6659e · outbound

This paper cites Microsoft coco: Common objects in context.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model Microsoft coco: Common objects in context

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T17:22:10.573970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:22:10.573970Z digest=sha256:fb5f74100f2e8975a4e2966be27c435d46f757e93840056abb627c599f9bf096

Observation 29dfe3b6-5bb1-4e61-a502-9567b7b7ff88 · outbound

This paper cites Improved baselines with visual instruction tuning, 2023 a.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model Improved baselines with visual instruction tuning, 2023 a

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T17:22:10.659686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:22:10.659686Z digest=sha256:81a90845f5e366fd0a5a5e59ef9cb0f2bf22d6234a143723d817c5f19a0c84ad

Observation ae8ffabf-0df9-44f0-8646-2f4000136edc · outbound

This paper cites Visual instruction tuning, 2023 b.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model Visual instruction tuning, 2023 b

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T17:22:10.712260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:22:10.712260Z digest=sha256:49260a4d7c9d4f9ef801194cb7b3e2b224022bf06de3e76ddec907c634ce0acc

Observation c9756c63-7999-4256-b83d-8bb282882a66 · outbound

This paper cites Llava-next: Improved reasoning, ocr, and world knowledge, January 2024.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model Llava-next: Improved reasoning, ocr, and world knowledge, January 2024

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T17:22:10.756326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:22:10.756326Z digest=sha256:f7a401270389a2e72a0f1ffc2dd286263e9df2d198614891f48c5bfdb6aebd36

Observation f30433d3-3016-47b7-90e4-eea774095811 · outbound

This paper cites Grounding dino: Marrying dino with grounded pre-training for open-set object detection.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model Grounding dino: Marrying dino with grounded pre-training for open-set object detection

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T17:22:10.818362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:22:10.818362Z digest=sha256:0dd904087c4eff05b82c49c687e146693eb976eeeac4df05ad9bd444c7ceb6a7

Observation daa8169b-1f48-4de4-ae48-fff91bc58e1a · outbound

This paper cites A convnet for the 2020s.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model A convnet for the 2020s

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T17:22:10.875282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:22:10.875282Z digest=sha256:1a76a6efc90e114f022b8aced0ecc6bb8ee4e63d40fcbd7cd8661fea184d4742

Observation 939d3dcc-b1d7-4e5e-b825-3378cd0e8548 · outbound

This paper cites Deepseek-vl: Towards real-world vision-language understanding, 2024.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model Deepseek-vl: Towards real-world vision-language understanding, 2024

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T17:22:10.920288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:22:10.920288Z digest=sha256:3d4bc86a6e87308d52ea29fa80eecf81c24800a0145bb702b793d9787d9adb42

Observation 6e3305b5-61ea-4571-a060-fea05536bf93 · outbound

This paper cites HumanTOMATO: Text-aligned Whole-body Motion Generation.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model HumanTOMATO: Text-aligned Whole-body Motion Generation

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T17:22:10.990900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:22:10.990900Z digest=sha256:dbda911e3fc88f923eb500e6f52b5016eec4ebafcc2dfdf15461e1fa4a528454

Observation b6d07928-854b-4ce8-931b-60c564c9f891 · outbound

This paper cites From keypoints to object landmarks via self-training correspondence: A novel approach to unsupervised landmark discovery.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model From keypoints to object landmarks via self-training correspondence: A novel approach to unsupervised landmark discovery

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:22:22.693855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T17:22:10.997048Z digest=sha256:578d435a8dd0241d8241a91484c86862fe447169a0905c60b4f9480d35c78aeb

Observation a5b53bb2-96bb-493c-b041-ddd1971f8bb5 · outbound

This paper cites MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T17:22:11.164532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:22:11.164532Z digest=sha256:722401e7e7cf666c82a3b6a73c8fd62e1827516f8af056bb09c1a43409012c0a

Observation 37e83a97-d2bc-4ebe-8f31-31181833f824 · outbound

This paper cites Gemma: Open Models Based on Gemini Research and Technology.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model Gemma: Open Models Based on Gemini Research and Technology

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T17:22:11.281720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:22:11.281720Z digest=sha256:07afce0f58301d0a243d2f783309ac9e4e03cd1a43a85da0587f74da350be91c

Observation f3bb7e68-4f45-46ec-bb30-0a496f696331 · outbound

This paper cites Interhand2.6m: A dataset and baseline for 3d interacting hand pose estimation from a single rgb image.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model Interhand2.6m: A dataset and baseline for 3d interacting hand pose estimation from a single rgb image

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:22:22.522833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T17:22:11.337039Z digest=sha256:d5860375a0e97e8916b01cdcc4d6740f70250da9c12181ea6a545049614cd94f

Observation 9dba3041-d485-46d1-ab8e-c9ffaa0219d1 · outbound

This paper cites Revisiting Fine-tuning for Few-shot Learning.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model Revisiting Fine-tuning for Few-shot Learning

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T17:22:11.496851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:22:11.496851Z digest=sha256:e08c680e77f3857fc03b47aac945cc0d73411fd615f5577ab17ff62d0e418ac6

Observation dc1ef1a2-29a6-48dd-92fa-dec6eceb1cdd · outbound

This paper cites Stacked hourglass networks for human pose estimation.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model Stacked hourglass networks for human pose estimation

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:22:22.402703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T17:22:11.582583Z digest=sha256:7ff895041c7207f6f7337e26f91c2313b268214ddc698d7cbae8c4fac234a156

Observation f6cdb89f-0137-4137-9b5e-dcc93f6750bb · outbound

This paper cites Animal kingdom: A large and diverse dataset for animal behavior understanding.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model Animal kingdom: A large and diverse dataset for animal behavior understanding

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:22:22.305728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T17:22:11.692051Z digest=sha256:e7cbdffbd6de79b89fcc5d5205ae2e4dbabad06b793694925a0f752190bbcba4

Observation 9bacbe1e-dc12-4098-b34a-28ed5de5a620 · outbound

This paper cites Single-stage multi-person pose machines.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model Single-stage multi-person pose machines

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:22:22.202969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T17:22:11.803010Z digest=sha256:357db75c9f3468b58df66e26e21784f5861e1cdca930d9d0fb937fd49fef7952

Observation cec821b3-8e61-44ef-94e5-0bc0296fc372 · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model DINOv2: Learning Robust Visual Features without Supervision

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T17:22:11.891931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:22:11.891931Z digest=sha256:c7152273ce7944728851b2147a9c887c895b6fba8bc859ab51052f789233ef11

Observation d5bacdeb-da28-496e-948a-7b8e26a9a10b · outbound

This paper cites Instruction Tuning with GPT-4.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model Instruction Tuning with GPT-4

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T17:22:12.010239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:22:12.010239Z digest=sha256:466b3bee5c40ddb1cf21ea23931df4dc9a66638b12ce3e657ece4dc352f528af

Observation 2f35e897-2827-487e-a2d1-71341e8baec7 · outbound

This paper cites Kosmos-2: Grounding Multimodal Large Language Models to the World.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model Kosmos-2: Grounding Multimodal Large Language Models to the World

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T17:22:12.091020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:22:12.091020Z digest=sha256:327d3b601916a9ee6ba69ae64d757cb0b28f73900bbe260829f95d23c39862d8

Observation f0717450-6257-4e36-8e3a-1b0e092292d2 · outbound

This paper cites DetGPT: Detect What You Need via Reasoning.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model DetGPT: Detect What You Need via Reasoning

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T17:22:12.156699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:22:12.156699Z digest=sha256:38eafddadbd09ac35a44ef1b5d68759663ed41fbdde68c5d7c25c28daf50235f

Observation ac7378b4-d987-40b7-aad6-4158d99f5e3e · outbound

This paper cites PerceptionGPT: Effectively Fusing Visual Perception into LLM.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model PerceptionGPT: Effectively Fusing Visual Perception into LLM

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T17:22:12.222475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:22:12.222475Z digest=sha256:9698725dab8015c650408eae45447897619d0cf663a0fbe0f5429e6424ea9ce4

Observation 00f6b4e1-e771-444b-8855-7ade893bbde3 · outbound

This paper cites Learning transferable visual models from natural language supervision.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model Learning transferable visual models from natural language supervision

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:22:22.125639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T17:22:12.301281Z digest=sha256:3eb23132d86d94414e70a84f32166a0d04454abc9512d97050ec347e1270284e

Observation ab397d07-183e-4853-b753-63ea015c68a6 · outbound

This paper cites Carfusion: Combining point tracking and part detection for dynamic 3d reconstruction of vehicles.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model Carfusion: Combining point tracking and part detection for dynamic 3d reconstruction of vehicles

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:22:22.053927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T17:22:12.360012Z digest=sha256:6118dd9e698928230dfb87e7e11dea850463d4fa1465a19a601c1952931cd87c

Observation c809e9d7-db83-40d6-82ee-37b313a06f21 · outbound

This paper cites Zafeiriou, and M.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model Zafeiriou, and M

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:22:21.902725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T17:22:12.443453Z digest=sha256:055c2efad4dabdf7ac8330fa9b77c682871d931f4e210ad8aef21d2a8b013b81

Observation 193d7a0f-d084-44a4-b50e-334e49b1ad21 · outbound

This paper cites Matching is not enough: A two-stage framework for category-agnostic pose estimation.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model Matching is not enough: A two-stage framework for category-agnostic pose estimation

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:22:21.697796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T17:22:12.518555Z digest=sha256:ab2402f509c0f7307479f742615f170adfa48084ca2f6f42c05637a7932a06ba

Observation 0044afed-fb6e-4f5d-92b4-7bd38c1191e3 · outbound

This paper cites Prototypical networks for few-shot learning.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model Prototypical networks for few-shot learning

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:22:21.562706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T17:22:12.586243Z digest=sha256:01ac443c67c1d2e6962d52c6ae11bcf2178750bdc6fa13eb946cb2391b595123

Observation d707b8cd-7738-4732-931a-b313d78d12ae · outbound

This paper cites Self-supervised keypoint discovery in behavioral videos.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model Self-supervised keypoint discovery in behavioral videos

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:22:21.473712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T17:22:12.645354Z digest=sha256:eb8cd455cd90b60a5e9c30537b1e54d9841879aeea2d0313c9790f9f5be875a2

Observation 4a5fb059-5cab-44c9-8db4-85a02e3febdd · outbound

This paper cites Deep high-resolution representation learning for human pose estimation.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model Deep high-resolution representation learning for human pose estimation

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:22:21.322446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T17:22:12.715825Z digest=sha256:016d81c7ce9d2b4dadc9c16392df4dc5b259561bf0f02cb8aba2378208a34a46

Observation 9d19d238-f1ea-4518-a951-d71afbf21c20 · outbound

This paper cites Compositional human pose regression.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model Compositional human pose regression

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:22:21.212839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T17:22:12.772097Z digest=sha256:3b5f8f3ab76fad708850d3540609c82c739fdb4c9f428a0d0c8af8f4f4f18136

Observation 020a2b09-2843-435d-96e9-b60b71384085 · outbound

This paper cites Deeppose: Human pose estimation via deep neural networks.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model Deeppose: Human pose estimation via deep neural networks

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:22:20.996690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T17:22:12.848054Z digest=sha256:15b0bf4e298723110869226cb9769b0535996c5acd88308e7a6a5c76551155c2

Observation 11fd6029-7466-4628-9867-b98b512596c8 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model LLaMA: Open and Efficient Foundation Language Models

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-06T17:22:12.913739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:22:12.913739Z digest=sha256:75ebd679c2db5699d727aaba875bb3a96a62273abff6f0e814cc54043be4f6e3

Observation ab97ce07-9584-4586-8107-d90835be4ba1 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-06T17:22:13.017689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:22:13.017689Z digest=sha256:d362d4286873360f405c5dffc1541c5af3a37d430da69cbec02207ed6ad83bff

Observation c6e17968-63d8-4058-99ed-d85d64764396 · outbound

This paper cites Attention is all you need.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model Attention is all you need

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:22:20.889936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T17:22:13.097863Z digest=sha256:4a678f5caf3478ff694f0e656a7d5763b7f8a77b20d172caaec3bd6c7de5ec67

Observation 1682b3c2-1b56-4d90-954f-4680153e4297 · outbound

This paper cites Locllm: Exploiting generalizable human keypoint localization via large language model.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model Locllm: Exploiting generalizable human keypoint localization via large language model

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:22:20.760112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T17:22:13.162173Z digest=sha256:406a2a30f5d671a2f74f247f4b9d4fec06110f8738cac67c6c1357f7cd0d09f6

Observation 1093694e-ff0d-42cc-b21b-0d6b8cc879ef · outbound

This paper cites Visionllm: Large language model is also an open-ended decoder for vision-centric tasks.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model Visionllm: Large language model is also an open-ended decoder for vision-centric tasks

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:22:20.495122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T17:22:13.224523Z digest=sha256:4551467c77b36d8c420f04650be7c89c809801785d092f9851044df2c31a349e

Observation 54aec589-a372-47f7-a2a5-8c701f4d0232 · outbound

This paper cites Convolutional pose machines.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model Convolutional pose machines

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:22:20.249219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T17:22:13.325920Z digest=sha256:e504326c15efefbce6813428c686515d48cfca944107367fccf9c2808ca71f92

Observation d43a7054-ee52-4ca0-8c0d-5cc729736004 · outbound

This paper cites VisionLLM v2: An End-to-End Generalist Multimodal Large Language Model for Hundreds of Vision-Language Tasks.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model VisionLLM v2: An End-to-End Generalist Multimodal Large Language Model for Hundreds of Vision-Language Tasks

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-06T17:22:13.366819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:22:13.366819Z digest=sha256:697e0a7acc7edf9450f6cdb35a239a521c5d9dfda50217c60297915b6f811e20

Observation 8029c7b6-d77d-49c4-a4dd-6d99ac6b047d · outbound

This paper cites F-LMM: Grounding Frozen Large Multimodal Models.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model F-LMM: Grounding Frozen Large Multimodal Models

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-06T17:22:13.424856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:22:13.424856Z digest=sha256:73988b8495138a9cd6ee9a68dd702b1493faa618776285a8e19006902fb9102e

Observation daa5e2e4-0e2d-460b-8b5b-bb93ef646c21 · outbound

This paper cites Look at boundary: A boundary-aware face alignment algorithm.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model Look at boundary: A boundary-aware face alignment algorithm

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:22:19.991106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T17:22:13.517966Z digest=sha256:264049a5f891d5d10779f905a589a2ea95722599ae5a0dbd08afd0e8b3ddcc26

Observation a8625e9d-9e14-4964-a116-79d40c55de5d · outbound

This paper cites Simple baselines for human pose estimation and tracking.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model Simple baselines for human pose estimation and tracking

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:22:19.816933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T17:22:13.609050Z digest=sha256:4ab081e6a6931a576f06d7483592016e472ba3b2e4e0c61d0f77955114d0262b

Observation 4af29d8b-5f40-4ccb-91ba-6ff19c44019e · outbound

This paper cites Pixel-aligned language model.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model Pixel-aligned language model

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:22:19.613550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T17:22:13.656756Z digest=sha256:dd4c336c5b3f1852641b04664c45e96c99125c9f08f6a27c7027834e086343b1

Observation 1b189f69-544b-4dbb-b977-e3fc681299bf · outbound

This paper cites Vipnas: Efficient video pose estimation via neural architecture search.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model Vipnas: Efficient video pose estimation via neural architecture search

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:22:19.315519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T17:22:13.687889Z digest=sha256:496b009820944280029aa82adacc5c02353319d529e036afc44160c98da8de26

Observation f51d56c7-31ae-4ef5-986b-2b22187b2ddc · outbound

This paper cites Pose for everything: Towards category-agnostic pose estimation.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model Pose for everything: Towards category-agnostic pose estimation

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:22:19.077852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T17:22:13.770149Z digest=sha256:75a77af48370b65df37c35999c84c540a1f228e3f84cd4effad9b647e0539827

Observation d665a2f6-f302-474a-9f64-27e72b9dfe03 · outbound

This paper cites Vitpose: Simple vision transformer baselines for human pose estimation.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model Vitpose: Simple vision transformer baselines for human pose estimation

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:22:18.920419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T17:22:13.820520Z digest=sha256:000ea9876a6dadecc54374f432404b23f01264f213f7ab1d919f830e553157af

Observation 623d16ab-ef14-4aeb-8c7f-ac9bdd74423a · outbound

This paper cites Boosting Human-Object Interaction Detection with Text-to-Image Diffusion Model.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model Boosting Human-Object Interaction Detection with Text-to-Image Diffusion Model

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-06T17:22:13.857000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:22:13.857000Z digest=sha256:3de80c8e84352c6af4a3dbac21c31b8e94776b4b31518731511ee90802babe4b

Observation e48fbddf-4a94-4d53-982f-5612c5b30cd0 · outbound

This paper cites Semantic human parsing via scalable semantic transfer over multiple label domains.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model Semantic human parsing via scalable semantic transfer over multiple label domains

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:22:18.797167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T17:22:13.935885Z digest=sha256:6363226bd9fc5590ba4650aa05bfb189556cb5c557cf603e33f0a1da322d217b

Observation d92d8528-f971-418f-befa-a03528242823 · outbound

This paper cites Neural interactive keypoint detection.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model Neural interactive keypoint detection

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:22:18.660229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T17:22:13.997849Z digest=sha256:91b404ae145720d63d4b5ff941c1c2fa703084516c423760e485949e07e1398c

Observation d6f0255e-3426-4a8b-b49c-dedc33e2547f · outbound

This paper cites Explicit Box Detection Unifies End-to-End Multi-Person Pose Estimation.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model Explicit Box Detection Unifies End-to-End Multi-Person Pose Estimation

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-06T17:22:14.093641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:22:14.093641Z digest=sha256:1431855550dded738de69e37684bcfdf3b39db41ecb52594ffbd2e3508b4126e

Observation 14404574-6dac-455e-868a-156d94e504ad · outbound

This paper cites X-Pose: Detecting Any Keypoints.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model X-Pose: Detecting Any Keypoints

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-06T17:22:14.151635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:22:14.151635Z digest=sha256:a7a003ace355332e0b1df573ca1a295ac8050e7aac5276edbd12bfe27e497cb1

Observation fbdff743-f45d-4789-a08a-b3958541ca59 · outbound

This paper cites F-HOI: Toward Fine-grained Semantic-Aligned 3D Human-Object Interactions.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model F-HOI: Toward Fine-grained Semantic-Aligned 3D Human-Object Interactions

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-06T17:22:14.216807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:22:14.216807Z digest=sha256:44d66e5af3c657c26506f7f16afd5b2aef11c17254610f8e22ced0c397770f1f

Observation 4ba0d7ec-2233-4007-aaac-5d034fe7b636 · outbound

This paper cites Kptllm: Unveiling the power of large language model for keypoint comprehension.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model Kptllm: Unveiling the power of large language model for keypoint comprehension

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:22:18.516527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T17:22:14.279793Z digest=sha256:5ac646e06d43232e7cb1cbe72e8412fd1e9c71e03ea86ca6c07cab771d1f0ab3

Observation 99992b3d-816a-4194-87e7-bafe0445bb99 · outbound

This paper cites Ed-pose++: Enhanced explicit box detection for conventional and interactive multi-object keypoint detection.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model Ed-pose++: Enhanced explicit box detection for conventional and interactive multi-object keypoint detection

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:22:18.361369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T17:22:14.319317Z digest=sha256:bfd6ba31e978cb9708595b442bed5a4b7178261e327bc33e26dc96a4577535ca

Observation c17f125e-98d2-4054-8b8b-f13350b71b81 · outbound

This paper cites Apt-36k: A large-scale benchmark for animal pose estimation and tracking.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model Apt-36k: A large-scale benchmark for animal pose estimation and tracking

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:22:18.199140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T17:22:14.386205Z digest=sha256:0932bc394c42b74ead7c18319ced8e36a522b11e9797133b944aeb4f632033f6

Observation ec2be9e1-6a25-49e9-92ff-b38306b2aec1 · outbound

This paper cites mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-06T17:22:14.478484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:22:14.478484Z digest=sha256:da5c4b9a740fe87ecf6b11efee89aa71bd8865646b1e22a97ce7c7111efc0c14

Observation 53a9ac90-37f5-488d-9b45-56ab876e9f8c · outbound

This paper cites Ferret: Refer and Ground Anything Anywhere at Any Granularity.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model Ferret: Refer and Ground Anything Anywhere at Any Granularity

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-06T17:22:14.559435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:22:14.559435Z digest=sha256:6c9d8d5ba96c2eb4452ea82e0422c1fc415da6a3db721db7cd57857a96cf8ce8

Observation 4bc9b194-deaa-45f1-95cb-ef08bbd1047d · outbound

This paper cites AP-10K: A Benchmark for Animal Pose Estimation in the Wild.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model AP-10K: A Benchmark for Animal Pose Estimation in the Wild

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-06T17:22:14.658414Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:22:14.658414Z digest=sha256:4341448da295774477e15108a8515830cc725df9884136eb362715d1bd3af3b6

Observation 2d8289d9-32c5-4689-b2d0-f2de7b9836a2 · outbound

This paper cites HRFormer: High-Resolution Transformer for Dense Prediction.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model HRFormer: High-Resolution Transformer for Dense Prediction

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-06T17:22:14.730911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:22:14.730911Z digest=sha256:cddd2633f4721c643ee019e88ef05d4391c7b90084d47641aa747fa28ed672c2

Observation 23390d0a-6e3d-4aa9-8eb2-366838c9140e · outbound

This paper cites Contextual Object Detection with Multimodal Large Language Models.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model Contextual Object Detection with Multimodal Large Language Models

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-06T17:22:14.824478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:22:14.824478Z digest=sha256:aea172587925a42f912479518f3729c2db10dbb3604491d27dd05e05baf50ced

Observation 1378ab1f-f0eb-418a-bf65-923ae2222a24 · outbound

This paper cites Sigmoid loss for language image pre-training.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model Sigmoid loss for language image pre-training

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:22:18.057396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T17:22:14.952673Z digest=sha256:0e7a9f9eece07080ad63f886122411fddeadc75b0d355043410e5d7c83aa71c9

Observation f921b1a8-6eff-4729-a310-922fc4c6a742 · outbound

This paper cites Open-vocabulary animal keypoint detection with semantic-feature matching.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model Open-vocabulary animal keypoint detection with semantic-feature matching

Reference 95

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:22:17.934887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T17:22:15.020088Z digest=sha256:c185dda4b598e03d1015a723efd8003e4426a913b7b6d5b8c850721fce111407

Observation 4447a533-086c-48e3-a756-4810421bc257 · outbound

This paper cites Adding conditional control to text-to-image diffusion models.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model Adding conditional control to text-to-image diffusion models

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:22:17.682866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T17:22:15.070625Z digest=sha256:8b6b2502cfc2dffa9fec3a451d19f2e4f906ed95e1e461257d2c36ed64fa0be6

Observation 963fe659-f831-48ff-a10a-51ae5926795a · outbound

This paper cites GPT4RoI: Instruction Tuning Large Language Model on Region-of-Interest.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model GPT4RoI: Instruction Tuning Large Language Model on Region-of-Interest

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-06T17:22:15.142692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:22:15.142692Z digest=sha256:3614e81df4469774669e720113a3d77400526232074913663513c3a98b7f10f3

Observation cb6a2be7-ccb1-4a1d-8b17-6225c4d86cb7 · outbound

This paper cites OPT: Open Pre-trained Transformer Language Models.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model OPT: Open Pre-trained Transformer Language Models

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-06T17:22:15.251884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:22:15.251884Z digest=sha256:5e2183db7657dcfd0b52dbd738316f6b3602f362d93ff9d77f473815131de072

Observation 54596103-3b07-4237-954f-a2cad0f014e8 · outbound

This paper cites Clamp: Prompt-based contrastive learning for connecting language and animal pose.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model Clamp: Prompt-based contrastive learning for connecting language and animal pose

Reference 99

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:22:17.428420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T17:22:15.316629Z digest=sha256:64661ae8cb6bb246af1b075a58c87bbfb401e4d16a68f4dfd7cd3b3053f4b855

Observation f0e8c4af-a326-435b-ba64-3a26311a833e · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-06T17:22:15.365213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:22:15.365213Z digest=sha256:d4526ffe94f4d62578ceee3a6dd02f46b9bcf1a43d8374532b5465a49e8ce1e6

Pith citing papers

No inbound Pith citation observations are available.