Pith. sign in

Paper Citation Record · LEDGER

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing

As of 13 August 2026, this Paper Citation Record lists 85 of 85 outbound references and 2 inbound Pith citation observations for arXiv:2411.16781.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.16781 v2

Coverage vector

measured 85 of 85 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T13:32:21.433343Z

measured 87 of 87 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T04:41:08.459981Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T21:55:03.803275Z

Reference resolution

85 of 85 outbound references displayed

  • verified exact0
  • verified fuzzy65
  • unresolved19
  • parse uncertain1
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d4a9d95a-cc54-4242-94f7-8010c3538f76 · outbound

This paper cites Gpt-4 technical report.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing Gpt-4 technical report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T13:32:20.912733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:32:20.912733Z digest=sha256:24e98537786be16861ea3e0bafd2af75894993fadb1547c9f4b8daf3985b2b54

Observation 7876657c-123a-4c00-bd8e-f9b94dd63012 · outbound

This paper cites Star-transformer: a spatio-temporal cross at- tention transformer for human action recognition.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing Star-transformer: a spatio-temporal cross at- tention transformer for human action recognition

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T13:32:20.917749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:32:20.917749Z digest=sha256:aa8bc59efe4848cf0a7f7589866a7d58017c1a1ac8342d1a43ffc04921338de8

Observation 64257746-f93a-4247-ae51-840318a12901 · outbound

This paper cites 2d human pose estimation: New benchmark and state of the art analysis.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing 2d human pose estimation: New benchmark and state of the art analysis

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T13:32:20.922135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:32:20.922135Z digest=sha256:d4c9d908b5528959ff77f3e89cecf95052ba18790e0f8793ee4043a8fc0a512e

Observation 0710852b-185f-46a2-bc17-5339ce4e5dac · outbound

This paper cites Qwen-vl: A frontier large vision-language model with versatile abilities.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing Qwen-vl: A frontier large vision-language model with versatile abilities

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T13:32:20.926753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:32:20.926753Z digest=sha256:37f241d7b77df0862e321a2fffd79c3c2da668ca78fbf89f3510aa7dd5a94f3a

Observation 618c2956-bb17-42a0-9955-61a8404d2a43 · outbound

This paper cites One token to seg them all: Language instructed rea- soning segmentation in videos.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing One token to seg them all: Language instructed rea- soning segmentation in videos

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T13:32:20.931493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:32:20.931493Z digest=sha256:772d567f076ed1b5328b01f2beddafb16db38b96116a4031779e169c3cf78f9e

Observation 84386f5b-e36d-4c59-87fd-b121bf6686b3 · outbound

This paper cites Meteor: An automatic metric for mt evaluation with improved correlation with hu- man judgments.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing Meteor: An automatic metric for mt evaluation with improved correlation with hu- man judgments

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:32:22.681723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:32:20.935657Z digest=sha256:89799b3a9856380e5041f1a849a39d797ef5a43110cfeaa11352c1fb2d677d4f

Observation 18720050-19d7-4b1c-b92d-ea0120947749 · outbound

This paper cites Keep it smpl: Automatic estimation of 3d human pose and shape from a single image.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing Keep it smpl: Automatic estimation of 3d human pose and shape from a single image

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:32:22.666850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:32:20.940623Z digest=sha256:ed352e51983cb9a2eca0666cc03a267ac16677fbf13fa2f2fc6509bd17b34633

Observation f12619a8-08b8-4fed-89fb-1fc90b4d2111 · outbound

This paper cites Towards bet- ter adversarial synthesis of human images from text.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing Towards bet- ter adversarial synthesis of human images from text

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:32:22.650345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:32:20.945599Z digest=sha256:3e0a996dbbe5684187217454cddd5bdc829eb269d5fd8e9413e6a295c1069f74

Observation eaba6592-26e3-45ca-a011-a1c59fa6334b · outbound

This paper cites Smpler-x: Scaling up expressive human pose and shape estimation.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing Smpler-x: Scaling up expressive human pose and shape estimation

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:32:22.632935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:32:20.949475Z digest=sha256:f7b820fb5fa7faf98539204cc16b750570d32a1c669f1d13bdae43a776882e27

Observation 85af158f-2e60-4438-960e-df3da4e3a9a9 · outbound

This paper cites Motionllm: Understanding human behaviors from human motions and videos.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing Motionllm: Understanding human behaviors from human motions and videos

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:32:22.614519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:32:20.953587Z digest=sha256:283178e0fcec718957f3fa490139ea7b4f79c23dc6cd6ea885d44799cd762687

Observation e130c720-0993-445a-8870-3d8ddb3088ec · outbound

This paper cites Channel-wise topology refinement graph convolution for skeleton-based action recognition.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing Channel-wise topology refinement graph convolution for skeleton-based action recognition

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:32:22.590230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:32:20.957532Z digest=sha256:77ca615f508c74348423a5abd39d1f07178bf1236b67dc78de6fbcb326a3d6dd

Observation 489fe822-fbc8-43e0-a267-2e8c147fb47c · outbound

This paper cites Learning phrase representations using rnn encoder-decoder for statistical machine translation.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing Learning phrase representations using rnn encoder-decoder for statistical machine translation

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:32:22.570041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:32:20.961555Z digest=sha256:65bcf02da36fa1f68d9edbe829351c40b644702e533ac3e2378f55261f73b54c

Observation b9136a30-9000-4a99-98bd-995a3bab6bac · outbound

This paper cites Posefix: correcting 3d human poses with natural language.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing Posefix: correcting 3d human poses with natural language

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:32:22.551363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:32:21.086359Z digest=sha256:7785db186b43d09d43c31fedd3fc1f6f9bed7515263442e525a5fe8452627f17

Observation a22aa2d8-6dde-493e-a01c-b73232649c71 · outbound

This paper cites Posescript: Linking 3d human poses and natural language.TPAMI, 2024.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing Posescript: Linking 3d human poses and natural language.TPAMI, 2024

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:32:22.528919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:32:21.095375Z digest=sha256:f9452a5e92a876ce638969c68b749c637279864bde59c186d209a6a0eafec70f

Observation 583e4259-9014-49d1-b768-3e52f335a8cd · outbound

This paper cites Poseembroider: Towards a 3d, visual, semantic-aware human pose representation.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing Poseembroider: Towards a 3d, visual, semantic-aware human pose representation

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:32:22.514957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:32:21.099673Z digest=sha256:cf1de41de7ca701330f429d385e379b0238f97534d1ca295e9fdc2d90aab836d

Observation 4309ce34-fc14-42a0-ace3-a3515811f555 · outbound

This paper cites Tokenhmr: Advancing human mesh recov- ery with a tokenized pose representation.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing Tokenhmr: Advancing human mesh recov- ery with a tokenized pose representation

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:32:22.493319Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:32:21.104976Z digest=sha256:42371f9313b4d9550476d481b0b8866c1acc4f140e515e4f376d747623414650

Observation 6f6d2052-187e-4cef-b672-5be79cf6357f · outbound

This paper cites Re- vitalizing optimization for 3d human pose and shape esti- mation: A sparse constrained formulation.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing Re- vitalizing optimization for 3d human pose and shape esti- mation: A sparse constrained formulation

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:32:22.481068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:32:21.108877Z digest=sha256:0b2e20a2e9f9471d3a6be4edce55574efbaf01f4c3315468c343f4487c5fc3eb

Observation 390a6615-27ba-466e-a2e6-5bcc94a7c95d · outbound

This paper cites Learning analytical posterior probabil- ity for human mesh recovery.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing Learning analytical posterior probabil- ity for human mesh recovery

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:32:22.468678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:32:21.112708Z digest=sha256:b0b53f5814015e6d29d1a9000121287800a0b547bcf3b6357dd3b3e4371fcf50

Observation ff1c9a1b-ec4d-4141-aec4-8e32ecf33157 · outbound

This paper cites Chatpose: Chatting about 3d human pose.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing Chatpose: Chatting about 3d human pose

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:32:22.455328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:32:21.119077Z digest=sha256:3d45a25b47b5331a5f5c3582352c1b1782913ba0712357da386fcf4581030af9

Observation 5bb87772-2c92-4eb2-9984-6751464dfd4e · outbound

This paper cites Vq-hps: Hu- man pose and shape estimation in a vector-quantized latent space.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing Vq-hps: Hu- man pose and shape estimation in a vector-quantized latent space

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:32:22.440633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:32:21.123544Z digest=sha256:ac7b9e4194bf3e9f552d4933ea6c120be6aa240b002cd908d19830468f5dfb20

Observation 41ac489c-5776-447f-b04d-1bacdd5c2c17 · outbound

This paper cites Mega: Masked generative autoencoder for human mesh recovery.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing Mega: Masked generative autoencoder for human mesh recovery

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:32:22.427046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:32:21.130396Z digest=sha256:10c7af318377bb564e1027760dd90c11da9b46c7532d8c562b6329bbe6065970

Observation e4769516-f9ef-41d0-af9c-9eeb653e25e7 · outbound

This paper cites Aifit: Automatic 3d human-interpretable feedback models for fitness training.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing Aifit: Automatic 3d human-interpretable feedback models for fitness training

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:32:22.413208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:32:21.134963Z digest=sha256:57324190aaaff7ccceb3954962ceb4255d9b1940c9acd476bcb35630f0a253f4

Observation 9466948d-567c-46a4-975c-c3917080cd9e · outbound

This paper cites Unified pose sequence modeling.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing Unified pose sequence modeling

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:32:22.399690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:32:21.139847Z digest=sha256:a59da766b9a986df3efb8a7bfb83b6880af29a990799efe97417a6798ef3b377

Observation 68089932-b864-4d82-8b19-118df3e87b86 · outbound

This paper cites Humans in 4d: Re- constructing and tracking humans with transformers.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing Humans in 4d: Re- constructing and tracking humans with transformers

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:32:22.385678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:32:21.143836Z digest=sha256:980fb2af0f047a2d20b756a12de6c8ea7d224c6cbfadc8852e35c1f2714584fe

Observation 33041c90-8215-474f-ac1d-bf82b637c59f · outbound

This paper cites Semantify: Simplifying the control of 3d morphable models using clip.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing Semantify: Simplifying the control of 3d morphable models using clip

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:32:22.365237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:32:21.147953Z digest=sha256:950c81740a4e034324accdd82a60e3d8314550d9d8c06a15538e4d07bc898684

Observation 00e4dfc9-7096-410a-8484-5145a9942887 · outbound

This paper cites Textbooks are all you need.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing Textbooks are all you need

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:32:22.348978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:32:21.154379Z digest=sha256:9ec036cba19341d511f7e48fd6cd09e29010a19fcf312f111f5c8e146b8ac73c

Observation 1252aec1-2562-4dd8-b177-1f11c6be110c · outbound

This paper cites Denoising dif- fusion probabilistic models.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing Denoising dif- fusion probabilistic models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T13:32:21.159567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:32:21.159567Z digest=sha256:46eb412380270b73472567f7880d4af52fbd4f3e5e149a6bb00d66c056a360a4

Observation 206c4be5-20dc-492f-9f47-ddb63b54324c · outbound

This paper cites Avatarclip: Zero-shot text- driven generation and animation of 3d avatars.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing Avatarclip: Zero-shot text- driven generation and animation of 3d avatars

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:32:22.327371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:32:21.164266Z digest=sha256:481b51b0a86bd17fbc0aa14f0e37cf7ec6796c4a3a903ad564c65cb00c4cbadf

Observation 0e1b1487-f2b4-494d-bbe7-cd694f061bef · outbound

This paper cites Lora: Low-rank adaptation of large language models.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing Lora: Low-rank adaptation of large language models

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:32:22.312089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:32:21.170376Z digest=sha256:a38889738d8eaa3da9ee1be5a92d5ca6b3b5fa0a34be85da5c34cf8f45d3bd54

Observation 2a406416-bd38-4005-b7ff-c9325f6a24cf · outbound

This paper cites an unresolved cited work.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-12T13:32:22.294645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:32:21.175172Z digest=sha256:35978f73298fa1e9cc150a8947efcbb0585940d335c6a3292160dd34a41c4ff3

Observation bfaa98e0-f5ae-4bc7-9858-20ee6442d96f · outbound

This paper cites Mistral 7b (2023).

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing Mistral 7b (2023)

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:32:22.277048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:32:21.179729Z digest=sha256:2ed2f2fc744b39c07818079110490d06e0438a3f894db7a7b64d59611df81610

Observation 125c29cc-af2e-4a43-8d5b-34660f115841 · outbound

This paper cites Motiongpt: Human motion as a foreign language.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing Motiongpt: Human motion as a foreign language

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T13:32:21.183957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:32:21.183957Z digest=sha256:644899ff4fb5376b53529ccfe114b2af89c63169d8308564bad9da08ee2aa38d

Observation daf88416-e854-4dfa-be74-387ffa80b8aa · outbound

This paper cites Learning effective hu- man pose estimation from inaccurate annotation.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing Learning effective hu- man pose estimation from inaccurate annotation

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:32:22.252169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:32:21.188573Z digest=sha256:9bebd06c1ca5b9996f6367bb7586043da1379cf161a1245ffe92d4d272202521

Observation 0d8d3e66-0a16-4507-865d-6e560fbd521d · outbound

This paper cites End-to-end recovery of human shape and pose.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing End-to-end recovery of human shape and pose

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:32:22.236153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:32:21.192875Z digest=sha256:e950541ad6fa434ad88cbf185886fece5e7045ed075727b7da443ccd2afad9f9

Observation e6904c24-4f28-4b1d-9023-75f697861773 · outbound

This paper cites Fixmypose: Pose correctional captioning and re- trieval.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing Fixmypose: Pose correctional captioning and re- trieval

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:32:22.221965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:32:21.197389Z digest=sha256:a55e5075a96b2e86f96dcd93796f9373eac8af8e96f38173f0be5d7c5524a50d

Observation 89d5e5ee-99e4-42b4-9bfd-8e4ab556e358 · outbound

This paper cites Segment any- thing.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing Segment any- thing

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T13:32:21.202172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:32:21.202172Z digest=sha256:53872ce289b049972920d0ff3ea0ab87e87e0ce76ec59cadeac9014d75f60c7f

Observation 3a5ef608-1c65-4fb3-b782-4b28855b6ad0 · outbound

This paper cites Lisa: Reasoning segmenta- tion via large language model.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing Lisa: Reasoning segmenta- tion via large language model

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T13:32:21.207670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:32:21.207670Z digest=sha256:27c978fcb452f687500ae569a6c1bd79cc72b0f424dc288a52c8db23ee64abad

Observation 658cc7c4-5e84-4895-9a1f-623ed5b87032 · outbound

This paper cites Llava-onevision: Easy visual task transfer.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing Llava-onevision: Easy visual task transfer

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:32:22.189486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:32:21.211655Z digest=sha256:5b41cb25280570f11c92590ed7189affd08b68f37dec297d69f17b1a330183cb

Observation 146ca529-ee33-4ac7-9f00-a43f1cc05f4f · outbound

This paper cites Video-llava: Learning united visual repre- sentation by alignment before projection.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing Video-llava: Learning united visual repre- sentation by alignment before projection

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:32:22.176369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:32:21.215654Z digest=sha256:c6ead0555022777308430ab9c1cdeb44785e8add6433c5a3f9bf9b626be7ff9c

Observation 9b343677-6b4c-4fce-a9c8-42aadf1ed9fd · outbound

This paper cites Rouge: A package for automatic evaluation of summaries.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing Rouge: A package for automatic evaluation of summaries

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T13:32:21.219603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:32:21.219603Z digest=sha256:1c1571680b95abf4d9ba5954e2a1c86a982ed7bac0ef55990120372330b2a150

Observation 39aff61a-637b-4c94-911f-692e15a205b6 · outbound

This paper cites Being comes from not-being: Open-vocabulary text-to-motion generation with wordless training.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing Being comes from not-being: Open-vocabulary text-to-motion generation with wordless training

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:32:22.155392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:32:21.223517Z digest=sha256:bf6de661700ff3c9b1ecb030a16b289dba21c54eb4f50b69f8c9a25dd87a3281

Observation 4d813e64-4013-42dd-b893-2af49af8e60f · outbound

This paper cites Chathuman: Language-driven 3d human understanding with retrieval-augmented tool reasoning.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing Chathuman: Language-driven 3d human understanding with retrieval-augmented tool reasoning

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:32:22.139661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:32:21.227355Z digest=sha256:f35dd1c02dadff29f872e9b26feeea81b1f80bd7453a0ef2cf8a16d4e14d38ea

Observation 1755b1c1-0985-4422-9f5b-0235735f5843 · outbound

This paper cites Microsoft coco: Common objects in context.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing Microsoft coco: Common objects in context

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:32:22.123783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:32:21.231695Z digest=sha256:52d808af3786b3d49aaa0c776528761038130f5df7c45ca007878c5abe7d6272

Observation 36489d87-307c-409c-898d-1a5b1d44fe09 · outbound

This paper cites Visual instruction tuning.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing Visual instruction tuning

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:32:22.108677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:32:21.235564Z digest=sha256:cdaf72f0f796c2f82fcf1b580c5c0f109671c82b4dceff618f5c2ee7edd6a02b

Observation 53c3bed6-6d39-4957-bc45-d127255eae69 · outbound

This paper cites Improved baselines with visual instruction tuning.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing Improved baselines with visual instruction tuning

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-12T13:32:21.239647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:32:21.239647Z digest=sha256:6bfea31152f2a59c3f8aa612603df01b642be2c3beeab9e8f6beba86292a3b00

Observation 6cd696fa-3315-4832-b608-1e46fc3064ba · outbound

This paper cites an unresolved cited work.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing Unresolved cited work

Reference 46

Resolution
unresolved
raw_fallback, observed 2026-08-12T13:32:22.085004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:32:21.243783Z digest=sha256:398f8389080a6dc51c64d7a108772872e5ff780675caf0e914db6e7d71aed27e

Observation 59e728d2-b0e5-45fc-83fb-cb7238f82bfe · outbound

This paper cites M 3gpt: An advanced multimodal, multitask framework for motion comprehension and generation.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing M 3gpt: An advanced multimodal, multitask framework for motion comprehension and generation

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:32:22.067985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:32:21.248502Z digest=sha256:e816ed217276d8137f01953713d68caf51ec5f9b672c7a446eb834043279aa1f

Observation ef6605ec-c997-43d6-ba72-277bed693146 · outbound

This paper cites Amass: Archive of mo- tion capture as surface shapes.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing Amass: Archive of mo- tion capture as surface shapes

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-12T13:32:21.252722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:32:21.252722Z digest=sha256:87b851fe8d2ef0dae736686f88d15edff1011ef5871fc8a7a2087d5cb9f1370e

Observation d1f61c97-5b83-4c01-9735-f83f1cd7693b · outbound

This paper cites Monocular 3d human pose estimation in the wild using improved cnn supervision.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing Monocular 3d human pose estimation in the wild using improved cnn supervision

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:32:22.040688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:32:21.257226Z digest=sha256:5aefd806284c65dda531be4f77a1a7a59fdafaa3e861f734a4fbbe155073834c

Observation 37c397a7-3b3d-4432-8676-b6451b05ad85 · outbound

This paper cites Hummuss: Human motion understanding using state space models.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing Hummuss: Human motion understanding using state space models

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:32:22.026305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:32:21.262874Z digest=sha256:7395f8a31039e612d2b04cf4f2f3abb2d205e4047a4fcefa912b7fe0241372f7

Observation 30ec4ef5-5d5b-4636-ba7b-6b76d87d6c93 · outbound

This paper cites Bleu: a method for automatic evaluation of machine translation.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing Bleu: a method for automatic evaluation of machine translation

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:32:22.007572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:32:21.267017Z digest=sha256:7f5f287c7e2386cc48e28d3252c148ab0556af8c4261847147921ad62bf8ca80

Observation c13ea94f-15d4-4a2b-ada4-ba3ebbed6392 · outbound

This paper cites an unresolved cited work.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing Unresolved cited work

Reference 52

Resolution
unresolved
raw_fallback, observed 2026-08-12T13:32:21.992983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:32:21.272160Z digest=sha256:bb3919d61e1ac36ede78129f88401947dd5f66e8b087606bad95a49cb47f2dca

Observation f7996632-5593-4697-9f46-c6da324c6749 · outbound

This paper cites Learn- ing transferable visual models from natural language super- vision.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing Learn- ing transferable visual models from natural language super- vision

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-12T13:32:21.280170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:32:21.280170Z digest=sha256:bac5faba00a76962c2e28205adcdf1e1bb32d6943bf2f6dc5c6dbae137eb1334

Observation 47101158-c16a-44a2-a8e3-3f223a612d7f · outbound

This paper cites Exploring the limits of transfer learning with a unified text-to-text transformer.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing Exploring the limits of transfer learning with a unified text-to-text transformer

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:32:21.967156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:32:21.285405Z digest=sha256:b5754b3ffe418773288e4e1177bd004599f4b5429161bf4c714ac0984f737146

Observation 242cf45e-6aec-4f84-b690-1fb76b55d11c · outbound

This paper cites Humor: 3d human motion model for robust pose estimation.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing Humor: 3d human motion model for robust pose estimation

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:32:21.951624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:32:21.291818Z digest=sha256:26bffa5d8769527a12dbb35f944f337342a75490265da7f551426f5c9ebbd30e

Observation d5e2eb03-b7e9-48f5-93e4-5b09f8689cd2 · outbound

This paper cites Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:32:21.937045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:32:21.298317Z digest=sha256:3389993a7555bc4792b58de1c85fb9b42242332dd0808ab700594fdc149aea06

Observation 555d9789-8f59-4a3c-9ca5-2b8d3b4e0166 · outbound

This paper cites Body talk: Crowdshaping realistic 3d avatars with words.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing Body talk: Crowdshaping realistic 3d avatars with words

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:32:21.922478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:32:21.304225Z digest=sha256:fadbf42b371ca3ac568eccd8800e4718104c86ebe978ede65c6c12bc68d41057

Observation dbdbca8d-4e47-481b-832b-277842654197 · outbound

This paper cites Aios: All-in-one-stage expressive human pose and shape estimation.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing Aios: All-in-one-stage expressive human pose and shape estimation

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:32:21.907398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:32:21.310717Z digest=sha256:d4073ff6d2ae93a7230c40b193f0218a1ff931a0d584cb016e1bbbfb0da0c262

Observation cd67ffb1-5e74-4d9a-9f17-a9363efe1e01 · outbound

This paper cites Llama: Open and efficient foundation language models.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing Llama: Open and efficient foundation language models

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:32:21.893392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:32:21.315582Z digest=sha256:01b71f7ed00a770f1b2c9930d5f7e954575e2f976c3d51d63739c9d45b078a21

Observation 80ceb2ec-20b9-4709-b8dc-1b993f1baf9c · outbound

This paper cites 3d human pose estimation via intuitive physics.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing 3d human pose estimation via intuitive physics

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:32:21.878143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:32:21.320093Z digest=sha256:079b7aa06cb7b9be8bf542a117e34718086c6633a5acdbc88c55d30aef6f67c6

Observation bbe40c2d-a3bf-4269-bda7-413dee71c8f5 · outbound

This paper cites Neural discrete representation learning.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing Neural discrete representation learning

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-12T13:32:21.324642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:32:21.324642Z digest=sha256:907d3000bd064b24f8f46590dd0c741f461a5cd104f57d94a2a5a8a6a77ba2f4

Observation ca28bed9-137a-4303-8f1b-2a5c094f7e46 · outbound

This paper cites Recovering ac- curate 3d human pose in the wild using imus and a moving camera.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing Recovering ac- curate 3d human pose in the wild using imus and a moving camera

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:32:21.854282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:32:21.329306Z digest=sha256:b41507fd16b960628af9dd285a2801cf4c0cd9365945a5ddb64848824c708328

Observation 0fe78897-2607-4d76-b83f-a916817a0ad6 · outbound

This paper cites Videomae v2: Scaling video masked autoencoders with dual masking.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing Videomae v2: Scaling video masked autoencoders with dual masking

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:32:21.839383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:32:21.333362Z digest=sha256:dc1fd8c2f7b5d0f69d1e925903f88e60a31560489d17d14704443d212bb91c20

Observation 5f8fef19-1172-49b2-8240-48eb37437015 · outbound

This paper cites Masked video distillation: Rethinking masked feature mod- eling for self-supervised video representation learning.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing Masked video distillation: Rethinking masked feature mod- eling for self-supervised video representation learning

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:32:21.823214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:32:21.337724Z digest=sha256:6d01cce3412da6bab1a9cd00db2435838d1ffacc47ac8dcd1b40e93fd1f8a847

Observation c3fde15c-4bdc-42a4-aee5-ab04ef616675 · outbound

This paper cites Zolly: Zoom focal length correctly for perspective- distorted human mesh reconstruction.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing Zolly: Zoom focal length correctly for perspective- distorted human mesh reconstruction

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:32:21.804549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:32:21.342598Z digest=sha256:486655a675aeb38f9f1da32b20b7225888ccf1b5c02fdd98282d7e235c56c7a4

Observation bf1c9eab-8d85-4415-92b1-89ab2d78b0d3 · outbound

This paper cites Cogvlm: Visual expert for pretrained language models.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing Cogvlm: Visual expert for pretrained language models

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:32:21.790484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:32:21.348023Z digest=sha256:ed97a9ae9aee9d4506632da4fc0412f5b94e553679d57953817be5c0f83f7797

Observation d9ddb6f0-be64-45a7-8e6e-63d335d8da32 · outbound

This paper cites Internvideo2: Scaling video foundation models for multimodal video understanding.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing Internvideo2: Scaling video foundation models for multimodal video understanding

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:32:21.771542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:32:21.352958Z digest=sha256:f6eb0bf6df024619ae77448f67683592e432e86a5c79fecc215f45f98c79da14

Observation 95df5f52-b8bc-43b2-8ed4-e786d8a3cf95 · outbound

This paper cites Occllama: An occupancy-language-action generative world model for au- tonomous driving.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing Occllama: An occupancy-language-action generative world model for au- tonomous driving

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:32:21.754573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:32:21.357567Z digest=sha256:2c4ea22cc1e57e0ea25ec5ffe28ddfb6b6fe5aae805d0861966071b49a277d9e

Observation a3791995-a88d-48ba-a198-cd649437d5a4 · outbound

This paper cites Motionllm: Multimodal motion-language learning with large language models.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing Motionllm: Multimodal motion-language learning with large language models

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:32:21.740719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:32:21.361735Z digest=sha256:acf61222f0de81b45438046dccbaee03fb8d29b16690f3f19977acf2b02a138d

Observation 58fcb0a1-17b4-4742-ba50-e2491ff90c62 · outbound

This paper cites Show-o: One single transformer to unify multimodal understanding and generation.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing Show-o: One single transformer to unify multimodal understanding and generation

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:32:21.726148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:32:21.365581Z digest=sha256:9ce8b56dddcdfeb3171704004b5d40124f7cdcb7c51a015def877f7b94b0f8bc

Observation dee4e429-7000-48bd-a48d-6304b05284c6 · outbound

This paper cites Smpler: Tam- ing transformers for monocular 3d human shape and pose estimation.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing Smpler: Tam- ing transformers for monocular 3d human shape and pose estimation

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:32:21.710679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:32:21.369509Z digest=sha256:e9b0a26eb941e9d1233c8578916f0143c1432d29c5e6278843acfcf58532ab09

Observation fc1f439a-454b-4263-9c70-d6f06b30ebaf · outbound

This paper cites Qwen2 technical report.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing Qwen2 technical report

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:32:21.690567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:32:21.373384Z digest=sha256:37b2b0507ba3c34bef10281208cadb756480a84e037137ce2a572aef8f86e1de

Observation 6c73b9cf-f7c0-45a3-8d48-94795c21b4dc · outbound

This paper cites Uni- audio 1.5: Large language model-driven audio codec is a few-shot audio task learner.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing Uni- audio 1.5: Large language model-driven audio codec is a few-shot audio task learner

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:32:21.666633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:32:21.377256Z digest=sha256:281250bbb22092f580a3ad3b43d15a8dec9d483faeb892be85051951a3a617c4

Observation 9ab6d6af-98f5-412f-a463-789852b3f402 · outbound

This paper cites mplug- owl3: Towards long image-sequence understanding in multi- modal large language models, 2024.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing mplug- owl3: Towards long image-sequence understanding in multi- modal large language models, 2024

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:32:21.651711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:32:21.382032Z digest=sha256:6f24e6f14a52466e5a40e70aff5bb751a0dc683e03e78d7299bb66984c43a368

Observation 070da214-fd14-4611-a9e6-035bbb3c628b · outbound

This paper cites Speechgpt: Empow- ering large language models with intrinsic cross-modal con- versational abilities.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing Speechgpt: Empow- ering large language models with intrinsic cross-modal con- versational abilities

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:32:21.635603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:32:21.388321Z digest=sha256:6ce8eb891581f8a9d6a71780d6a5cdc1f160adedc3fecf821452eb592d423b13

Observation 80b6d2f3-338a-4e27-b3d2-389936a221b6 · outbound

This paper cites Pymaf: 3d human pose and shape regression with pyramidal mesh alignment feedback loop.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing Pymaf: 3d human pose and shape regression with pyramidal mesh alignment feedback loop

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:32:21.608653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:32:21.393026Z digest=sha256:cf8dbf1662ce1470c24013a53f789dd57ad6eabf9c9dcb0fbac6fd10ccf5e1e0

Observation 0923ea77-9991-4f7d-b6ed-0e7059ff899e · outbound

This paper cites Mo- tiongpt: Finetuned llms are general-purpose motion genera- tors.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing Mo- tiongpt: Finetuned llms are general-purpose motion genera- tors

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-12T13:32:21.397086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:32:21.397086Z digest=sha256:afaa2d206f3eb5dc34986d5bb72e76c39dd9110edfd48b6a1c113d24ad0d4ea5

Observation 10617ccd-78e5-4739-b49e-bbac5808ac5b · outbound

This paper cites Single im- age action recognition using semantic body part actions.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing Single im- age action recognition using semantic body part actions

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:32:21.576250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:32:21.401677Z digest=sha256:cc1ff469b0fbe858f3b991059688c7b003dee4a065df0d03337b3160daca45db

Observation bd19217a-22b0-4553-a834-422408928f10 · outbound

This paper cites Transfusion: Pre- dict the next token and diffuse images with one multi-modal model.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing Transfusion: Pre- dict the next token and diffuse images with one multi-modal model

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:32:21.557106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:32:21.406441Z digest=sha256:643ba867ac371ca1b374aea242f156b6d03f725ea8224fd9a12c52c4a1d1a280

Observation 2747de09-6f07-4641-8364-a8e97054a666 · outbound

This paper cites Minigpt-4: Enhancing vision-language understanding with advanced large language models.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing Minigpt-4: Enhancing vision-language understanding with advanced large language models

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:32:21.541729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:32:21.411596Z digest=sha256:cb3858413257c7a43bc99bd4a50d61fd2be8de1222857071a0f5c4c778925b06

Observation 2fbf5d57-adee-4b25-9015-171dd9c2f7f7 · outbound

This paper cites Therightkneeandthigharestraight,andtheleftlegisalsostraight.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing Therightkneeandthigharestraight,andtheleftlegisalsostraight

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:32:21.526223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:32:21.416075Z digest=sha256:59b25024575a6e4f4adea373687ca04f886b4e654a08428fb5f73c3308fba2ec

Observation a8cc1aff-ce1c-4d1c-8c34-94e38ef00ca9 · outbound

This paper cites This extensive data effectively facilitates the alignment of pose and text modalities.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing This extensive data effectively facilitates the alignment of pose and text modalities

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:32:21.510767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:32:21.421879Z digest=sha256:0ac54b2f8421ef0a24dd3b8279e02c4ea94166864ebea4acf57b9587ce9737aa

Observation d480161a-2b97-4133-a588-a36d37ded064 · outbound

This paper cites an unresolved cited work.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing Unresolved cited work

Reference 84

Resolution
unresolved
raw_fallback, observed 2026-08-12T13:32:21.495015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:32:21.427240Z digest=sha256:a10a514e6e94d9eecd764e61e279dacaeea634bbd23cfb37fa77d7868884fbd7

Observation 8ae9533b-a6b3-4252-b69c-ed44c7bf8e5f · outbound

This paper cites The results show that our approach more accurately estimates human poses, even in scenarios with complex limb articulations.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing The results show that our approach more accurately estimates human poses, even in scenarios with complex limb articulations

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:32:21.476877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:32:21.433343Z digest=sha256:45984a73f1ff404096294f846225ba22d91c4b50864bf60a2db9f64df9e2e922

Observation ef29c048-4adc-4374-ad36-65e4d83caaf2 · outbound

This paper cites an unresolved cited work.

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing Unresolved cited work

Reference 2023

Resolution
parse uncertain
no resolver link, observed 2026-08-12T13:32:21.090809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:32:21.090809Z digest=sha256:f450ed2d23efe8823185d0c89953a5f6027d398fa2531599b8cf22ceca108b33

Pith citing papers

Observation deb65451-07b8-4f89-b9a5-20ee598e4a3e · inbound

Human-Centric Foundation Models: Perception, Generation and Agentic Modeling cites this paper.

Human-Centric Foundation Models: Perception, Generation and Agentic Modeling UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-08T04:41:08.459981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:41:08.459981Z digest=sha256:3d737b7e005b117408efa57b4f83de120197bb68d95306c11aac5d6956079ed0

Observation f031606a-9685-4d35-967b-9da7383f2d0e · inbound

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model cites this paper.

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-08-05T21:55:03.873453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-05T21:55:02.001344Z digest=sha256:033745e18f1f1f4fdb3d08b77f7e47bef5841afb4187c88e7c6bd41cc01a8c1b