Pith. sign in

Paper Citation Record · LEDGER

SmartAvatar: Text- and Image-Guided Human Avatar Generation with VLM AI Agents

As of 10 August 2026, this Paper Citation Record lists 26 of 26 outbound references and 0 inbound Pith citation observations for arXiv:2506.04606.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.04606 v1

Coverage vector

measured 26 of 26 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:41:39.824941Z

measured 26 of 26 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

26 of 26 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved22
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch4

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ee0b2cfa-35cf-4910-93c5-1836c154d872 · outbound

This paper cites GPT-4 Technical Report.

SmartAvatar: Text- and Image-Guided Human Avatar Generation with VLM AI Agents GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T10:41:39.714323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:41:39.714323Z digest=sha256:9373e093763162bc0fba0ed8c41bd6d074cbe97b0146ea93ed957997f3355ddd

Observation e68f483f-94a4-47ae-9c1d-88bd9c6501d6 · outbound

This paper cites AvatarCraft: Transforming Text into Neural Human Avatars with Parameterized Shape and Pose Control.

SmartAvatar: Text- and Image-Guided Human Avatar Generation with VLM AI Agents AvatarCraft: Transforming Text into Neural Human Avatars with Parameterized Shape and Pose Control

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T10:41:39.736092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:41:39.736092Z digest=sha256:6c2d2571adb688d0a23f40e3db354a62720c5e1e2118aa63e5cb11683dfcb549

Observation c37a48f8-18fc-446c-a3df-fe125ff6eb97 · outbound

This paper cites BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models.

SmartAvatar: Text- and Image-Guided Human Avatar Generation with VLM AI Agents BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T10:41:39.740007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:41:39.740007Z digest=sha256:471d870d397ca8fd6d713734cc5ca39c060b73b0739f94054ce00434b21fe140

Observation ac184a21-3bb7-450a-a9d8-33504e110371 · outbound

This paper cites TADA! Text to Animatable Digital Avatars.

SmartAvatar: Text- and Image-Guided Human Avatar Generation with VLM AI Agents TADA! Text to Animatable Digital Avatars

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T10:41:39.744076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:41:39.744076Z digest=sha256:f9fb238705647ab58b26fc159dac593365a20d5af4556c7e8bcede05de53a194

Observation 795b7840-cf2a-4cf9-9c5b-0025bdffbbe3 · outbound

This paper cites Visual Instruction Tuning.

SmartAvatar: Text- and Image-Guided Human Avatar Generation with VLM AI Agents Visual Instruction Tuning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T10:41:39.747720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:41:39.747720Z digest=sha256:84598146176b324bd09e4c23fe809b51f83f56d053fefdfac8a0315d65b713f4

Observation 464db041-1c9c-472b-9cc4-221915909147 · outbound

This paper cites AgentBench: Evaluating LLMs as Agents.

SmartAvatar: Text- and Image-Guided Human Avatar Generation with VLM AI Agents AgentBench: Evaluating LLMs as Agents

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T10:41:39.751445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:41:39.751445Z digest=sha256:319e09edce4289cadb62632a46aa484b33a241a544496c253753a2ec36e72f25

Observation 9d43d8c9-e1b8-403a-ad63-96d2acb49759 · outbound

This paper cites Self-Refine: Iterative Refinement with Self-Feedback.

SmartAvatar: Text- and Image-Guided Human Avatar Generation with VLM AI Agents Self-Refine: Iterative Refinement with Self-Feedback

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T10:41:39.759301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:41:39.759301Z digest=sha256:a8d0b76ab1c28d2d9a7bb67723c59d2214f49fb28aa62a4937fd5906840b1a5a

Observation 1353627b-702d-4fee-9e35-aedf2b9feafa · outbound

This paper cites URL https://arxiv.org/abs/ 2304.03442.

SmartAvatar: Text- and Image-Guided Human Avatar Generation with VLM AI Agents URL https://arxiv.org/abs/ 2304.03442

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T10:41:39.762937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:41:39.762937Z digest=sha256:d8862b764110b9e0868f00d3a9e0e57b1da58edd86b9953bca571925d0eb298a

Observation f6aec695-2dc4-42fe-9694-2ba047b30b1c · outbound

This paper cites CharacterGen: Efficient 3D Character Generation from Single Images with Multi-View Pose Canonicalization.

SmartAvatar: Text- and Image-Guided Human Avatar Generation with VLM AI Agents CharacterGen: Efficient 3D Character Generation from Single Images with Multi-View Pose Canonicalization

Reference 15

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T10:41:40.013907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T10:41:39.771357Z digest=sha256:e0e076d4c30cbdf0ccaa3633f20c3f6f8f243e47548f51903ac70e2e8f9ddd86

Observation 0ae27a19-bb7d-45c8-a90a-163c9f79c48d · outbound

This paper cites Iteration of Thought: Leveraging Inner Dialogue for Autonomous Large Language Model Reasoning.

SmartAvatar: Text- and Image-Guided Human Avatar Generation with VLM AI Agents Iteration of Thought: Leveraging Inner Dialogue for Autonomous Large Language Model Reasoning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T10:41:39.775381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:41:39.775381Z digest=sha256:fedcb124f330201bddec0893f4a21439e713d41a59ff41d1b521b21e4f25d2c1

Observation 2440ee96-9bca-442c-bc12-f83a7e3550d0 · outbound

This paper cites Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding.

SmartAvatar: Text- and Image-Guided Human Avatar Generation with VLM AI Agents Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T10:41:39.784487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:41:39.784487Z digest=sha256:1579247bd5b1c6f81f39126e629dafe43f9e7b64ff6de860eeefc33d6e242556

Observation 109e1a66-33b1-4ffd-80fb-7bf8336653d3 · outbound

This paper cites DreamGaussian: Generative Gaussian Splatting for Efficient 3D Content Creation.

SmartAvatar: Text- and Image-Guided Human Avatar Generation with VLM AI Agents DreamGaussian: Generative Gaussian Splatting for Efficient 3D Content Creation

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T10:41:39.788222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:41:39.788222Z digest=sha256:e454ccfb2ac01c51f8d81a412fe7f452f0d534f2b96924c1f0278a7148d7312e

Observation 672510bf-4af3-49a9-8d90-f75f67a417b1 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

SmartAvatar: Text- and Image-Guided Human Avatar Generation with VLM AI Agents Gemini: A Family of Highly Capable Multimodal Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T10:41:39.792243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:41:39.792243Z digest=sha256:cdd98b8bbc64b80d024d5668b0b6f039c4e537dbf974aa6416e3b9252e20ff82

Observation 179e2aa8-ddb8-47f0-8e3e-2f727b916229 · outbound

This paper cites Voyager: An Open-Ended Embodied Agent with Large Language Models.

SmartAvatar: Text- and Image-Guided Human Avatar Generation with VLM AI Agents Voyager: An Open-Ended Embodied Agent with Large Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T10:41:39.799981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:41:39.799981Z digest=sha256:51599f249cbf19770a009450dcd6e80fde1d3d7b6d5ee770b8600a48357268ef

Observation ab2b0cbd-497b-4bef-aee8-4226b831beeb · outbound

This paper cites Structured 3D Latents for Scalable and Versatile 3D Generation.

SmartAvatar: Text- and Image-Guided Human Avatar Generation with VLM AI Agents Structured 3D Latents for Scalable and Versatile 3D Generation

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T10:41:39.804402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:41:39.804402Z digest=sha256:c5e101223e3f6dbb960c64b90513358547e83b62cc4b852dd53259c708609bbe

Observation 71778dda-07f4-4aaa-a986-80a7c5997c52 · outbound

This paper cites WorldGPT: A Sora-Inspired Video AI Agent as Rich World Models from Text and Image Inputs.

SmartAvatar: Text- and Image-Guided Human Avatar Generation with VLM AI Agents WorldGPT: A Sora-Inspired Video AI Agent as Rich World Models from Text and Image Inputs

Reference 24

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T10:41:39.915452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T10:41:39.808081Z digest=sha256:111dd8966a0c3af12d397fe461dd9145b8596884efc4f9772bf1038c7d2bf93c

Observation 46d7bc0c-c24a-429d-b1c0-e129e2f8af17 · outbound

This paper cites AvatarGen: a 3D Generative Model for Animatable Human Avatars.

SmartAvatar: Text- and Image-Guided Human Avatar Generation with VLM AI Agents AvatarGen: a 3D Generative Model for Animatable Human Avatars

Reference 25

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T10:41:39.897848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T10:41:39.811748Z digest=sha256:ba55ec739b42add383df07d7ceb375571e9583d007add5e3a0d1403beddcdd73

Observation 4c3688f1-8754-4aa2-a93a-1081b07e0e2b · outbound

This paper cites Adding Conditional Control to Text-to-Image Diffusion Models.

SmartAvatar: Text- and Image-Guided Human Avatar Generation with VLM AI Agents Adding Conditional Control to Text-to-Image Diffusion Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T10:41:39.816391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:41:39.816391Z digest=sha256:2bcb76f62c03b03091a2bad3d8b62d76ff206bd48b9c217cddb8c9598c2ae426

Observation dfa004f7-bd2c-4913-a1c4-b38a39225106 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

SmartAvatar: Text- and Image-Guided Human Avatar Generation with VLM AI Agents MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T10:41:39.824941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:41:39.824941Z digest=sha256:9f4123ac17ba4fa66a03796d0e59e34887196946e24c0afc4809bdc346a2d8ec

Observation 9d6d6703-8fe7-407a-9bd1-6bbbc0a90cac · outbound

This paper cites doi: 10.1145/2816795.2818013.

SmartAvatar: Text- and Image-Guided Human Avatar Generation with VLM AI Agents doi: 10.1145/2816795.2818013

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-07T10:41:39.755671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:41:39.755671Z digest=sha256:87412a61f2de3060a17c327c88a985e2f3e1ef592d5c075093bbfe5fc4eeb282

Observation 3ed96b5b-db96-46d3-ae2e-966c36cebbd1 · outbound

This paper cites Expressive Body Capture: 3D Hands, Face, and Body from a Single Image.

SmartAvatar: Text- and Image-Guided Human Avatar Generation with VLM AI Agents Expressive Body Capture: 3D Hands, Face, and Body from a Single Image

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-07T10:41:39.766934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:41:39.766934Z digest=sha256:4f00b357b1732f57517f04eeec8f8b830ff50c62749501ad37309748cce35d6b

Observation 96cc0822-947f-497b-bf70-8c2d52e32e5b · outbound

This paper cites Zero-Shot Text-to-Image Generation.

SmartAvatar: Text- and Image-Guided Human Avatar Generation with VLM AI Agents Zero-Shot Text-to-Image Generation

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-07T10:41:39.779709Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:41:39.779709Z digest=sha256:6768cdc36e1e9ff67a96386e5518d3817a4345f9b084a576296cb1087660a86c

Observation ba95ddda-170e-419e-b4f0-e00ec34ef932 · outbound

This paper cites EVA3D: Compositional 3D Human Generation from 2D Image Collections.

SmartAvatar: Text- and Image-Guided Human Avatar Generation with VLM AI Agents EVA3D: Compositional 3D Human Generation from 2D Image Collections

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-07T10:41:39.727380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:41:39.727380Z digest=sha256:78eaf44cd5ea29001ad18c4a11007a0835eb58ccabe90a9b12a6ccd74fd803d5

Observation 5a9a7a47-5443-4d02-b457-2dcc451d77ac · outbound

This paper cites AG3D: Learning to Generate 3D Avatars from 2D Image Collections.

SmartAvatar: Text- and Image-Guided Human Avatar Generation with VLM AI Agents AG3D: Learning to Generate 3D Avatars from 2D Image Collections

Reference 2023

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T10:41:40.269231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T10:41:39.723413Z digest=sha256:3f14865a1de3bfccf18cc9184f4ad4ce90fdba91474ad7f4998519e893c2d1dd

Observation 0326ba3a-d26e-4e6e-90e5-1cc9b3fe4256 · outbound

This paper cites StructLDM: Structured Latent Diffusion for 3D Human Generation.

SmartAvatar: Text- and Image-Guided Human Avatar Generation with VLM AI Agents StructLDM: Structured Latent Diffusion for 3D Human Generation

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T10:41:39.732131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:41:39.732131Z digest=sha256:478873a07cded1d82794638b3950d0741c5e14794c10794e49ccf78f674135e0

Observation 07f8e3c3-0c4b-49da-9c4b-4fc59bc741a6 · outbound

This paper cites LLMs can see and hear without any training.

SmartAvatar: Text- and Image-Guided Human Avatar Generation with VLM AI Agents LLMs can see and hear without any training

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T10:41:39.719320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:41:39.719320Z digest=sha256:a766851a20732e389a2fb0aded91223a9266ae08c774089d181183188a0b0fad

Pith citing papers

No inbound Pith citation observations are available.