Pith. sign in

Paper Citation Record · LEDGER

SmartAvatar: Text- and Image-Guided Human Avatar Generation with VLM AI Agents

As of 18 August 2026, this Paper Citation Record lists 26 of 26 outbound references and 0 inbound Pith citation observations for arXiv:2506.04606.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.04606 v1

Coverage vector

measured 26 of 26 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:41:39.824941Z

measured 26 of 26 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

26 of 26 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved22
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch4

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ee0b2cfa-35cf-4910-93c5-1836c154d872 · outbound

This paper cites GPT-4 Technical Report.

SmartAvatar: Text- and Image-Guided Human Avatar Generation with VLM AI Agents GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T10:41:39.714323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:41:39.714323Z digest=sha256:447c6cf4ab6a30b50bec25d2b1cbbc0b92c4b10d788ca4ec55b7eddc4ee79ebf

Observation e68f483f-94a4-47ae-9c1d-88bd9c6501d6 · outbound

This paper cites AvatarCraft: Transforming Text into Neural Human Avatars with Parameterized Shape and Pose Control.

SmartAvatar: Text- and Image-Guided Human Avatar Generation with VLM AI Agents AvatarCraft: Transforming Text into Neural Human Avatars with Parameterized Shape and Pose Control

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T10:41:39.736092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:41:39.736092Z digest=sha256:b327f3008416579e11ea15250c05485cedbaba79fea5130b70af18c03b13d0d8

Observation c37a48f8-18fc-446c-a3df-fe125ff6eb97 · outbound

This paper cites BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models.

SmartAvatar: Text- and Image-Guided Human Avatar Generation with VLM AI Agents BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T10:41:39.740007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:41:39.740007Z digest=sha256:2536cc19a56e9e8a900e39d19a9b967b66db64aecc784067095c03823c499a69

Observation ac184a21-3bb7-450a-a9d8-33504e110371 · outbound

This paper cites TADA! Text to Animatable Digital Avatars.

SmartAvatar: Text- and Image-Guided Human Avatar Generation with VLM AI Agents TADA! Text to Animatable Digital Avatars

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T10:41:39.744076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:41:39.744076Z digest=sha256:2b0518edd02faada8b986425946076b963c5660ad12714d7ec15537585310515

Observation 795b7840-cf2a-4cf9-9c5b-0025bdffbbe3 · outbound

This paper cites Visual Instruction Tuning.

SmartAvatar: Text- and Image-Guided Human Avatar Generation with VLM AI Agents Visual Instruction Tuning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T10:41:39.747720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:41:39.747720Z digest=sha256:b36dc521794d8ce0df5aefd7c7265754f6c091a590f556b4edab6e0bf509bd2b

Observation 464db041-1c9c-472b-9cc4-221915909147 · outbound

This paper cites AgentBench: Evaluating LLMs as Agents.

SmartAvatar: Text- and Image-Guided Human Avatar Generation with VLM AI Agents AgentBench: Evaluating LLMs as Agents

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T10:41:39.751445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:41:39.751445Z digest=sha256:280b4dd6c4963b4286bc628df576e141e39512eb6c04665e254fbcb750ab851d

Observation 9d43d8c9-e1b8-403a-ad63-96d2acb49759 · outbound

This paper cites Self-Refine: Iterative Refinement with Self-Feedback.

SmartAvatar: Text- and Image-Guided Human Avatar Generation with VLM AI Agents Self-Refine: Iterative Refinement with Self-Feedback

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T10:41:39.759301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:41:39.759301Z digest=sha256:511c0c0122baaf285aa780f27a08a26ba266479168c9a78b6f053a073c8f139d

Observation 1353627b-702d-4fee-9e35-aedf2b9feafa · outbound

This paper cites URL https://arxiv.org/abs/ 2304.03442.

SmartAvatar: Text- and Image-Guided Human Avatar Generation with VLM AI Agents URL https://arxiv.org/abs/ 2304.03442

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T10:41:39.762937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:41:39.762937Z digest=sha256:5734630518a0f7a802b03cd39ce1329bcd7a5c284e61eccb7f43ce5e56982e0c

Observation f6aec695-2dc4-42fe-9694-2ba047b30b1c · outbound

This paper cites CharacterGen: Efficient 3D Character Generation from Single Images with Multi-View Pose Canonicalization.

SmartAvatar: Text- and Image-Guided Human Avatar Generation with VLM AI Agents CharacterGen: Efficient 3D Character Generation from Single Images with Multi-View Pose Canonicalization

Reference 15

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T10:41:40.013907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T10:41:39.771357Z digest=sha256:84cec43ccb0b02e644312cef6183d453c91cfd6440c996334f7a070881095906

Observation 0ae27a19-bb7d-45c8-a90a-163c9f79c48d · outbound

This paper cites Iteration of Thought: Leveraging Inner Dialogue for Autonomous Large Language Model Reasoning.

SmartAvatar: Text- and Image-Guided Human Avatar Generation with VLM AI Agents Iteration of Thought: Leveraging Inner Dialogue for Autonomous Large Language Model Reasoning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T10:41:39.775381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:41:39.775381Z digest=sha256:70f702e0ba65d20c071c67e01de524be4924b8555c5f3d8bca161b20aab315d1

Observation 2440ee96-9bca-442c-bc12-f83a7e3550d0 · outbound

This paper cites Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding.

SmartAvatar: Text- and Image-Guided Human Avatar Generation with VLM AI Agents Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T10:41:39.784487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:41:39.784487Z digest=sha256:0bbdc0d0f96593347ac1901378145095d2ae9e3155b8da64a6c25f5168a99e34

Observation 109e1a66-33b1-4ffd-80fb-7bf8336653d3 · outbound

This paper cites DreamGaussian: Generative Gaussian Splatting for Efficient 3D Content Creation.

SmartAvatar: Text- and Image-Guided Human Avatar Generation with VLM AI Agents DreamGaussian: Generative Gaussian Splatting for Efficient 3D Content Creation

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T10:41:39.788222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:41:39.788222Z digest=sha256:b98aaba021bdec069cae541dc2bce233c2a0cb7217117043a0b633637d82dacb

Observation 672510bf-4af3-49a9-8d90-f75f67a417b1 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

SmartAvatar: Text- and Image-Guided Human Avatar Generation with VLM AI Agents Gemini: A Family of Highly Capable Multimodal Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T10:41:39.792243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:41:39.792243Z digest=sha256:ef176f03fcd8917d97240691ce2fd94e1fb5424fc6041f881d57fd48b2cb66ef

Observation 179e2aa8-ddb8-47f0-8e3e-2f727b916229 · outbound

This paper cites Voyager: An Open-Ended Embodied Agent with Large Language Models.

SmartAvatar: Text- and Image-Guided Human Avatar Generation with VLM AI Agents Voyager: An Open-Ended Embodied Agent with Large Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T10:41:39.799981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:41:39.799981Z digest=sha256:d2889161ca258a49f2c4e9518ec6cd9a81628279530ad38a003dd28dd439dfbf

Observation ab2b0cbd-497b-4bef-aee8-4226b831beeb · outbound

This paper cites Structured 3D Latents for Scalable and Versatile 3D Generation.

SmartAvatar: Text- and Image-Guided Human Avatar Generation with VLM AI Agents Structured 3D Latents for Scalable and Versatile 3D Generation

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T10:41:39.804402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:41:39.804402Z digest=sha256:a086fa19156764a83255920f1a71eff1ac3f722966593c7938174d4f874c534d

Observation 71778dda-07f4-4aaa-a986-80a7c5997c52 · outbound

This paper cites WorldGPT: A Sora-Inspired Video AI Agent as Rich World Models from Text and Image Inputs.

SmartAvatar: Text- and Image-Guided Human Avatar Generation with VLM AI Agents WorldGPT: A Sora-Inspired Video AI Agent as Rich World Models from Text and Image Inputs

Reference 24

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T10:41:39.915452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T10:41:39.808081Z digest=sha256:3fa111278342f898254b693dc7fdce998856e70995c2679ad67cd11cfa1f91be

Observation 46d7bc0c-c24a-429d-b1c0-e129e2f8af17 · outbound

This paper cites AvatarGen: a 3D Generative Model for Animatable Human Avatars.

SmartAvatar: Text- and Image-Guided Human Avatar Generation with VLM AI Agents AvatarGen: a 3D Generative Model for Animatable Human Avatars

Reference 25

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T10:41:39.897848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T10:41:39.811748Z digest=sha256:cb9e85188f678b57eddef99b9457915bad389748852e8ccb67176ecce27eaa1d

Observation 4c3688f1-8754-4aa2-a93a-1081b07e0e2b · outbound

This paper cites Adding Conditional Control to Text-to-Image Diffusion Models.

SmartAvatar: Text- and Image-Guided Human Avatar Generation with VLM AI Agents Adding Conditional Control to Text-to-Image Diffusion Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T10:41:39.816391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:41:39.816391Z digest=sha256:ad7325a22869692c72dc1589fe3c231e0dd2f837f3d2d998891f6bd1cf80b38b

Observation dfa004f7-bd2c-4913-a1c4-b38a39225106 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

SmartAvatar: Text- and Image-Guided Human Avatar Generation with VLM AI Agents MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T10:41:39.824941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:41:39.824941Z digest=sha256:e8dd6602abf4f464075927be07c03dd8f1104eaeb098406620388b3b957d935c

Observation 9d6d6703-8fe7-407a-9bd1-6bbbc0a90cac · outbound

This paper cites doi: 10.1145/2816795.2818013.

SmartAvatar: Text- and Image-Guided Human Avatar Generation with VLM AI Agents doi: 10.1145/2816795.2818013

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-07T10:41:39.755671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:41:39.755671Z digest=sha256:5c9c6d1e0eae4280edeeaff729274a6d26381baafe50ea950694d20db67ac258

Observation 3ed96b5b-db96-46d3-ae2e-966c36cebbd1 · outbound

This paper cites Expressive Body Capture: 3D Hands, Face, and Body from a Single Image.

SmartAvatar: Text- and Image-Guided Human Avatar Generation with VLM AI Agents Expressive Body Capture: 3D Hands, Face, and Body from a Single Image

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-07T10:41:39.766934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:41:39.766934Z digest=sha256:ed4eb7bf9eea84787542c103b848492e01ef57c2f6ec3c2452e5b716df442ac0

Observation 96cc0822-947f-497b-bf70-8c2d52e32e5b · outbound

This paper cites Zero-Shot Text-to-Image Generation.

SmartAvatar: Text- and Image-Guided Human Avatar Generation with VLM AI Agents Zero-Shot Text-to-Image Generation

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-07T10:41:39.779709Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:41:39.779709Z digest=sha256:99219874664f5eb51121057bc7aa887cc49a3123c1a341f19f84d865c96cc2cb

Observation ba95ddda-170e-419e-b4f0-e00ec34ef932 · outbound

This paper cites EVA3D: Compositional 3D Human Generation from 2D Image Collections.

SmartAvatar: Text- and Image-Guided Human Avatar Generation with VLM AI Agents EVA3D: Compositional 3D Human Generation from 2D Image Collections

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-07T10:41:39.727380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:41:39.727380Z digest=sha256:1458209c26da60d0723d771be64d1b1630032b283b65a2d42869513968a47dad

Observation 5a9a7a47-5443-4d02-b457-2dcc451d77ac · outbound

This paper cites AG3D: Learning to Generate 3D Avatars from 2D Image Collections.

SmartAvatar: Text- and Image-Guided Human Avatar Generation with VLM AI Agents AG3D: Learning to Generate 3D Avatars from 2D Image Collections

Reference 2023

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T10:41:40.269231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T10:41:39.723413Z digest=sha256:05e1737722bc70265c59badebbae92904d64cbda48a02c150b5404164ab746b5

Observation 0326ba3a-d26e-4e6e-90e5-1cc9b3fe4256 · outbound

This paper cites StructLDM: Structured Latent Diffusion for 3D Human Generation.

SmartAvatar: Text- and Image-Guided Human Avatar Generation with VLM AI Agents StructLDM: Structured Latent Diffusion for 3D Human Generation

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T10:41:39.732131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:41:39.732131Z digest=sha256:57ed934450412c6abf5f63c73fb3661f69c62797a5a8a6580347e9c1f42e1f9b

Observation 07f8e3c3-0c4b-49da-9c4b-4fc59bc741a6 · outbound

This paper cites LLMs can see and hear without any training.

SmartAvatar: Text- and Image-Guided Human Avatar Generation with VLM AI Agents LLMs can see and hear without any training

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T10:41:39.719320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:41:39.719320Z digest=sha256:438b3c31f03c1feff5cc43346902ed128531b12ae989d26758758605d74af562

Pith citing papers

No inbound Pith citation observations are available.