Pith. sign in

Paper Citation Record · LEDGER

VLOGGER: Multimodal Diffusion for Embodied Avatar Synthesis

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 7 inbound Pith citation observations for arXiv:2403.08764.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2403.08764 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 7 of 7 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:07:15.011660Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T02:16:26.572604Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 477d88a1-bf97-4697-8fbd-6fd9f95d32fd · inbound

JoyVASA: Portrait and Animal Image Animation with Diffusion-Based Audio-Driven Facial Dynamics and Head Motion Generation cites this paper.

JoyVASA: Portrait and Animal Image Animation with Diffusion-Based Audio-Driven Facial Dynamics and Head Motion Generation VLOGGER: Multimodal Diffusion for Embodied Avatar Synthesis

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-23T17:03:12.574515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-23T16:59:57.727035Z digest=sha256:1bb347084f0d93e312cf9546008dbba1c80671311ab63a73a13558ebc6c002f7

Observation b1686fa9-01dc-4ba4-91c5-e233b32bcb2f · inbound

HunyuanVideo-HOMA: Generic Human-Object Interaction in Multimodal Driven Human Animation cites this paper.

HunyuanVideo-HOMA: Generic Human-Object Interaction in Multimodal Driven Human Animation VLOGGER: Multimodal Diffusion for Embodied Avatar Synthesis

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T05:07:15.011660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:07:15.011660Z digest=sha256:3489afde41d1eea4000652d04dbc7caf8deb5449ef92af7be69f8f86f6444137

Observation 2a41ba8e-4dd0-466a-969b-6a826e4cb14e · inbound

MirrorMe: Towards Realtime and High Fidelity Audio-Driven Halfbody Animation cites this paper.

MirrorMe: Towards Realtime and High Fidelity Audio-Driven Halfbody Animation VLOGGER: Multimodal Diffusion for Embodied Avatar Synthesis

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T22:15:06.874990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:15:06.874990Z digest=sha256:cab7cc7f032372da130394155e603c98cd7812fb2d13cc2b293a0bf60fbbd314

Observation 55ba466d-8195-47bc-870d-cab19cd2a9c5 · inbound

Democratizing High-Fidelity Co-Speech Gesture Video Generation cites this paper.

Democratizing High-Fidelity Co-Speech Gesture Video Generation VLOGGER: Multimodal Diffusion for Embodied Avatar Synthesis

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T19:00:03.193055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:00:03.193055Z digest=sha256:b148f5fae674bc1c0266d4777796a286761be670e2b1933e4733a33ce709e522

Observation 75aff3a9-89ae-4ee7-82fe-024be2d77f6e · inbound

Human Motion Video Generation: A Survey cites this paper.

Human Motion Video Generation: A Survey VLOGGER: Multimodal Diffusion for Embodied Avatar Synthesis

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-05T10:36:52.868747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:36:52.868747Z digest=sha256:83280352ad1c65882209f3b65ba5f1f02337b4432eb6a5853b247b45673fe803

Observation 87fa9fe3-c486-4f36-99f3-de57c8131ef1 · inbound

Mamba-Enhanced Implicit Motion Learning for Audio-Driven Portrait Animation cites this paper.

Mamba-Enhanced Implicit Motion Learning for Audio-Driven Portrait Animation VLOGGER: Multimodal Diffusion for Embodied Avatar Synthesis

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:16:26.574247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T11:07:20.577007Z digest=sha256:35adfa43f3b243069b076de53742730b1b274525e9e15b6fa50d42c6e8c91528

Observation 05028948-7764-44fc-8935-07aa5a4baf34 · inbound

AgentHOI: Multi-Agent Reasoning for Human-Object-Interaction Video Generation via Implicit Representation Alignment cites this paper.

AgentHOI: Multi-Agent Reasoning for Human-Object-Interaction Video Generation via Implicit Representation Alignment VLOGGER: Multimodal Diffusion for Embodied Avatar Synthesis

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-01T05:27:16.036956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T05:27:16.036956Z digest=sha256:4ed9e48404e8c7bd1ca1c59e0e1469545490f83eb67e25a52fb745f5e6154f35