Pith. sign in

Paper Citation Record · LEDGER

TextHOI-3D: Text-to-3D Hand-Object Interaction via Discrete Multi-View Generation and Joint Mesh Optimization

As of 18 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 0 inbound Pith citation observations for arXiv:2606.11805.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.11805 v1

Coverage vector

measured 31 of 31 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-27T09:59:43.874631Z

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

31 of 31 outbound references displayed

  • verified exact9
  • verified fuzzy0
  • unresolved22
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e74a529a-0816-4df2-ab39-c63dca8eea03 · outbound

This paper cites Reconstructing hand-object interactions in the wild.

TextHOI-3D: Text-to-3D Hand-Object Interaction via Discrete Multi-View Generation and Joint Mesh Optimization Reconstructing hand-object interactions in the wild

Reference 1

Resolution
unresolved
no resolver link, observed 2026-06-27T09:59:43.874631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T09:59:43.874631Z digest=sha256:57bc867707d83b6931421a0b2d691a121f95c479e84ee12546968760af4c0530

Observation 5bcaacf8-9861-422e-93b8-aa8e73cfcaf7 · outbound

This paper cites Text2hoi: Text-guided 3d motion generation for hand-object interaction.

TextHOI-3D: Text-to-3D Hand-Object Interaction via Discrete Multi-View Generation and Joint Mesh Optimization Text2hoi: Text-guided 3d motion generation for hand-object interaction

Reference 2

Resolution
unresolved
no resolver link, observed 2026-06-27T09:59:43.874631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T09:59:43.874631Z digest=sha256:061d862b141f29ec1fe0887d4a6d87c4058187b14cb4dc73177430e463f69906

Observation 75454c51-d163-43a7-ba1a-9310d0fd6eaa · outbound

This paper cites Generative pretraining from pixels.

TextHOI-3D: Text-to-3D Hand-Object Interaction via Discrete Multi-View Generation and Joint Mesh Optimization Generative pretraining from pixels

Reference 3

Resolution
unresolved
no resolver link, observed 2026-06-27T09:59:43.874631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T09:59:43.874631Z digest=sha256:b872b581a6aa8b5302741a682df15bde8915a5838260e2410a6fd27a4f260227

Observation 0aef0014-7ad4-4c98-b7ad-fbebc5a56183 · outbound

This paper cites Alignsdf: Pose-aligned signed distance fields for hand-object reconstruction.

TextHOI-3D: Text-to-3D Hand-Object Interaction via Discrete Multi-View Generation and Joint Mesh Optimization Alignsdf: Pose-aligned signed distance fields for hand-object reconstruction

Reference 4

Resolution
unresolved
no resolver link, observed 2026-06-27T09:59:43.874631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T09:59:43.874631Z digest=sha256:24cf1857a2c4780951aa77103020b88a04f22559ba4d78c8ebd80bf23d5f4b13

Observation afcd23be-6a81-4667-988a-5a621e790523 · outbound

This paper cites gsdf: Geometry-driven signed distance functions for 3d hand-object reconstruction.

TextHOI-3D: Text-to-3D Hand-Object Interaction via Discrete Multi-View Generation and Joint Mesh Optimization gsdf: Geometry-driven signed distance functions for 3d hand-object reconstruction

Reference 5

Resolution
unresolved
no resolver link, observed 2026-06-27T09:59:43.874631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T09:59:43.874631Z digest=sha256:055abff8a0effa435124105a07955b0d35e671f2c434cb816fa152f12d1f41a8

Observation 3a973239-dea6-43c9-962a-9db66658ceb5 · outbound

This paper cites Taming transformers for high-resolution image synthesis.

TextHOI-3D: Text-to-3D Hand-Object Interaction via Discrete Multi-View Generation and Joint Mesh Optimization Taming transformers for high-resolution image synthesis

Reference 6

Resolution
unresolved
no resolver link, observed 2026-06-27T09:59:43.874631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T09:59:43.874631Z digest=sha256:f2dbd433480cdc73cce53cdda5b600a874ed2a979a9d0810a42f1263a99ca9ec

Observation 6633bb25-b669-4d83-b0a2-b952bc4646d4 · outbound

This paper cites Honnotate: A method for 3d annotation of hand and object poses.

TextHOI-3D: Text-to-3D Hand-Object Interaction via Discrete Multi-View Generation and Joint Mesh Optimization Honnotate: A method for 3d annotation of hand and object poses

Reference 7

Resolution
unresolved
no resolver link, observed 2026-06-27T09:59:43.874631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T09:59:43.874631Z digest=sha256:481450c7d133de8163f3bd3687f44cbcb5fb3afa32ae7d34d8c36e65bacb4f0a

Observation f3adef3a-cafc-47fa-b417-a7c9ecf0726d · outbound

This paper cites Learning joint reconstruction of hands and manipulated objects.

TextHOI-3D: Text-to-3D Hand-Object Interaction via Discrete Multi-View Generation and Joint Mesh Optimization Learning joint reconstruction of hands and manipulated objects

Reference 8

Resolution
unresolved
no resolver link, observed 2026-06-27T09:59:43.874631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T09:59:43.874631Z digest=sha256:d455943a34ce0d3a3ec729f3716fbed8659a6e6cd49728be2685f93211101143

Observation 47514864-d4a8-4745-970d-489b0c7f065b · outbound

This paper cites Leveraging photometric consistency over time for sparsely supervised hand-object reconstruction.

TextHOI-3D: Text-to-3D Hand-Object Interaction via Discrete Multi-View Generation and Joint Mesh Optimization Leveraging photometric consistency over time for sparsely supervised hand-object reconstruction

Reference 9

Resolution
unresolved
no resolver link, observed 2026-06-27T09:59:43.874631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T09:59:43.874631Z digest=sha256:df4b17fcfbe5a47eb672eed67bdd5733cd608fa64a7432d1015cdc003a7f597a

Observation 9f5c9d72-17ba-47d0-9e17-8ae72eb7e814 · outbound

This paper cites Classifier-Free Diffusion Guidance.

TextHOI-3D: Text-to-3D Hand-Object Interaction via Discrete Multi-View Generation and Joint Mesh Optimization Classifier-Free Diffusion Guidance

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-07-03T10:27:56.855035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-27T09:59:43.874631Z digest=sha256:31e33fce14ba37a867e41272c690fbd75643399a288656733db4c47b169b2448

Observation 39f722dd-d2e9-4b4e-bb3a-7162f60a8462 · outbound

This paper cites The Curious Case of Neural Text Degeneration.

TextHOI-3D: Text-to-3D Hand-Object Interaction via Discrete Multi-View Generation and Joint Mesh Optimization The Curious Case of Neural Text Degeneration

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-07-03T10:27:56.857582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-27T09:59:43.874631Z digest=sha256:0878607bc654c635663f829cb614da36ac7a347328b21b5f5756ecdc7615eea8

Observation 3dba5774-d574-4c82-9c7a-783b8395959e · outbound

This paper cites LRM: Large Reconstruction Model for Single Image to 3D.

TextHOI-3D: Text-to-3D Hand-Object Interaction via Discrete Multi-View Generation and Joint Mesh Optimization LRM: Large Reconstruction Model for Single Image to 3D

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-07-03T10:27:56.840287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-27T09:59:43.874631Z digest=sha256:4b36660c8fd67a0584f87ddb7823fcda647bac9526c2fda4fed4d127c4f786fd

Observation 4e8b580a-8a3f-430e-a8b1-dc741115c386 · outbound

This paper cites Zero-1-to-3: Zero-shot one image to 3d object.

TextHOI-3D: Text-to-3D Hand-Object Interaction via Discrete Multi-View Generation and Joint Mesh Optimization Zero-1-to-3: Zero-shot one image to 3d object

Reference 13

Resolution
unresolved
no resolver link, observed 2026-06-27T09:59:43.874631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T09:59:43.874631Z digest=sha256:726ac4cc695bac46a8b8dd18dfdc611cbf096f25e26bee5e48c052538bd98c16

Observation 4862b1c9-c01a-4ca7-9ef9-df1d0c8e38e2 · outbound

This paper cites Decoupled Weight Decay Regularization.

TextHOI-3D: Text-to-3D Hand-Object Interaction via Discrete Multi-View Generation and Joint Mesh Optimization Decoupled Weight Decay Regularization

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-07-03T10:27:56.846020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-27T09:59:43.874631Z digest=sha256:9ad815a9e9facede85076f3515720f016d3137a45d2312afe88da29f40cb0e26

Observation e761997e-92a6-4dd9-adef-a56f5ac62f10 · outbound

This paper cites Reconstructing hands in 3d with transformers.

TextHOI-3D: Text-to-3D Hand-Object Interaction via Discrete Multi-View Generation and Joint Mesh Optimization Reconstructing hands in 3d with transformers

Reference 15

Resolution
unresolved
no resolver link, observed 2026-06-27T09:59:43.874631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T09:59:43.874631Z digest=sha256:e8974421f8671d90233998e79da75b5a2a3a7b9d573718e79466873ab472229f

Observation 21740990-6cc8-4f06-8e97-1bd63a19891f · outbound

This paper cites DreamFusion: Text-to-3D using 2D Diffusion.

TextHOI-3D: Text-to-3D Hand-Object Interaction via Discrete Multi-View Generation and Joint Mesh Optimization DreamFusion: Text-to-3D using 2D Diffusion

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-07-03T10:27:56.849412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-27T09:59:43.874631Z digest=sha256:9ea8801abfbf3656f41da11811177944710af6f24a0e986fd01560ebb95e2302

Observation a1b906c0-8d1d-48b8-abe8-cd4cff78cad8 · outbound

This paper cites Learning transferable visual models from natural language supervision.

TextHOI-3D: Text-to-3D Hand-Object Interaction via Discrete Multi-View Generation and Joint Mesh Optimization Learning transferable visual models from natural language supervision

Reference 17

Resolution
unresolved
no resolver link, observed 2026-06-27T09:59:43.874631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T09:59:43.874631Z digest=sha256:5f96b922fded23a2dc9f6e9b4e963876eb9a8132451c2011dd5c6b36476dd1db

Observation c9793f5c-5e6b-4fb5-8bec-574f23a04e8f · outbound

This paper cites Accelerating 3D Deep Learning with PyTorch3D.

TextHOI-3D: Text-to-3D Hand-Object Interaction via Discrete Multi-View Generation and Joint Mesh Optimization Accelerating 3D Deep Learning with PyTorch3D

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-07-03T10:27:56.843276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-27T09:59:43.874631Z digest=sha256:93a0622ab486637158a3faeef23497affe57d6a6cfd1ab38ae84b962ea0c0833

Observation 944895e8-a486-498b-98f9-64143b1dea53 · outbound

This paper cites Generating diverse high-fidelity images with vq-vae-2.

TextHOI-3D: Text-to-3D Hand-Object Interaction via Discrete Multi-View Generation and Joint Mesh Optimization Generating diverse high-fidelity images with vq-vae-2

Reference 19

Resolution
unresolved
no resolver link, observed 2026-06-27T09:59:43.874631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T09:59:43.874631Z digest=sha256:100db3e6357c517cc337c613427a4dc0e3545147dbae8d22dda80cf5cc6d005d

Observation 45e4b46a-af01-4d77-af92-b2e33f728b67 · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

TextHOI-3D: Text-to-3D Hand-Object Interaction via Discrete Multi-View Generation and Joint Mesh Optimization High-resolution image synthesis with latent diffusion models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-06-27T09:59:43.874631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T09:59:43.874631Z digest=sha256:f9481dafe47c0a4bf0de083168691c07824f19629dd4f9c3d633ba18115b5a4c

Observation d70d405b-88e8-4ff6-b92b-ef2a7f8dc4f9 · outbound

This paper cites an unresolved cited work.

TextHOI-3D: Text-to-3D Hand-Object Interaction via Discrete Multi-View Generation and Joint Mesh Optimization Unresolved cited work

Reference 21

Resolution
unresolved
no resolver link, observed 2026-06-27T09:59:43.874631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T09:59:43.874631Z digest=sha256:63ceb20ad279f4c6de4ad3c241891e820d8063d9f1ea709890ba9cc50511067b

Observation d15bd7d1-bd28-426e-8c9d-1ee2c80a21af · outbound

This paper cites MVDream: Multi-view Diffusion for 3D Generation.

TextHOI-3D: Text-to-3D Hand-Object Interaction via Discrete Multi-View Generation and Joint Mesh Optimization MVDream: Multi-view Diffusion for 3D Generation

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-07-03T10:27:56.837575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-27T09:59:43.874631Z digest=sha256:6912d99ee0b94505e621f6aa3f47561be677c9b6d647b34fbe5e200f176ee133

Observation fa8219fd-72b9-4af5-90bd-d952167abea5 · outbound

This paper cites MVDiffusion: Enabling Holistic Multi-view Image Generation with Correspondence-Aware Diffusion.

TextHOI-3D: Text-to-3D Hand-Object Interaction via Discrete Multi-View Generation and Joint Mesh Optimization MVDiffusion: Enabling Holistic Multi-view Image Generation with Correspondence-Aware Diffusion

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-07-03T10:27:56.852339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-27T09:59:43.874631Z digest=sha256:9a36362faaf8d4ad03df8a0dc621dc4f4eafb2f041ebd0f859c35ce32bdc0aea

Observation 08085502-c884-4ad7-922f-7df00d222544 · outbound

This paper cites Visual autoregressive modeling: Scalable image generation via next-scale prediction.Advances in neural information processing systems, 37:84839–84865, 2024.

TextHOI-3D: Text-to-3D Hand-Object Interaction via Discrete Multi-View Generation and Joint Mesh Optimization Visual autoregressive modeling: Scalable image generation via next-scale prediction.Advances in neural information processing systems, 37:84839–84865, 2024

Reference 24

Resolution
unresolved
no resolver link, observed 2026-06-27T09:59:43.874631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T09:59:43.874631Z digest=sha256:3655fbe8f19424d069f76bbbf852d8d50d130825d18a3058e4cd978cacd38d32

Observation 38f560e2-aba6-4985-8b1b-cd86488c0722 · outbound

This paper cites Pixel recurrent neural networks.

TextHOI-3D: Text-to-3D Hand-Object Interaction via Discrete Multi-View Generation and Joint Mesh Optimization Pixel recurrent neural networks

Reference 25

Resolution
unresolved
no resolver link, observed 2026-06-27T09:59:43.874631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T09:59:43.874631Z digest=sha256:7868a5126c1c78d342d313ed7cfec5b8a347a2b068a45f9d87d17555a84f76a9

Observation 43b01606-f003-463b-b93a-3b6624e04fb3 · outbound

This paper cites Neural discrete representation learning.Advances in neural information processing systems, 30, 2017.

TextHOI-3D: Text-to-3D Hand-Object Interaction via Discrete Multi-View Generation and Joint Mesh Optimization Neural discrete representation learning.Advances in neural information processing systems, 30, 2017

Reference 26

Resolution
unresolved
no resolver link, observed 2026-06-27T09:59:43.874631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T09:59:43.874631Z digest=sha256:c80028fdc17952548487b62435c0b614c5ba70cbdd4aa943e08eb99b3ad07e64

Observation c72e84f7-61f3-4142-bda9-7fd213c20923 · outbound

This paper cites Image quality assessment: from error visibility to structural similarity.IEEE transactions on image processing, 13(4):600–612, 2004.

TextHOI-3D: Text-to-3D Hand-Object Interaction via Discrete Multi-View Generation and Joint Mesh Optimization Image quality assessment: from error visibility to structural similarity.IEEE transactions on image processing, 13(4):600–612, 2004

Reference 27

Resolution
unresolved
no resolver link, observed 2026-06-27T09:59:43.874631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T09:59:43.874631Z digest=sha256:33c93d92ef4e05559546dac2e106684deee98abfeba0c62ecc97110ad35bd522

Observation c7dfc73a-ad96-469e-97bd-aa1a5ed5fc19 · outbound

This paper cites InstantMesh: Efficient 3D Mesh Generation from a Single Image with Sparse-view Large Reconstruction Models.

TextHOI-3D: Text-to-3D Hand-Object Interaction via Discrete Multi-View Generation and Joint Mesh Optimization InstantMesh: Efficient 3D Mesh Generation from a Single Image with Sparse-view Large Reconstruction Models

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-07-03T10:27:56.860266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-27T09:59:43.874631Z digest=sha256:76d8ce47ad4f6820eb37b1b53441289884a193aad10335c8109002d18aa2a169

Observation 6a08f6c2-9900-4a80-bd15-3acfe315393c · outbound

This paper cites What’s in your hands? 3d reconstruction of generic objects in hands.

TextHOI-3D: Text-to-3D Hand-Object Interaction via Discrete Multi-View Generation and Joint Mesh Optimization What’s in your hands? 3d reconstruction of generic objects in hands

Reference 29

Resolution
unresolved
no resolver link, observed 2026-06-27T09:59:43.874631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T09:59:43.874631Z digest=sha256:0234c3263cfd77525ce2486522f5638ddc1141ec2530d004964fb3a62a0c897a

Observation 70b19a03-c4de-4cb5-9edc-75d74104b24b · outbound

This paper cites Moho: Learning single-view hand-held object reconstruction with multi-view occlusion-aware supervision.

TextHOI-3D: Text-to-3D Hand-Object Interaction via Discrete Multi-View Generation and Joint Mesh Optimization Moho: Learning single-view hand-held object reconstruction with multi-view occlusion-aware supervision

Reference 30

Resolution
unresolved
no resolver link, observed 2026-06-27T09:59:43.874631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T09:59:43.874631Z digest=sha256:9b501dab4b630b8e0e5c14ec68e9cc4950d2b66e829a93478d41b5cd049b0926

Observation 3298a7e5-138d-4a78-a72d-ad8119e15c72 · outbound

This paper cites The unreasonable effectiveness of deep features as a perceptual metric.

TextHOI-3D: Text-to-3D Hand-Object Interaction via Discrete Multi-View Generation and Joint Mesh Optimization The unreasonable effectiveness of deep features as a perceptual metric

Reference 31

Resolution
unresolved
no resolver link, observed 2026-06-27T09:59:43.874631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T09:59:43.874631Z digest=sha256:23693fce4558689d91c29ff25179701be84cd5d3525ea3b6b9cc4cff4f64b698

Pith citing papers

No inbound Pith citation observations are available.