Pith. sign in

Paper Citation Record · LEDGER

SViMo: Synchronized Diffusion for Video and Motion Generation in Hand-object Interaction Scenarios

As of 9 August 2026, this Paper Citation Record lists 72 of 72 outbound references and 5 inbound Pith citation observations for arXiv:2506.02444.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.02444 v3

Coverage vector

measured 72 of 72 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:29:06.878705Z

measured 77 of 77 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T04:10:14.672688Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-14T00:18:30.107372Z

Reference resolution

72 of 72 outbound references displayed

  • verified exact1
  • verified fuzzy30
  • unresolved41
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation fc5e5238-9f2a-4b5e-ad10-7e160adb5df5 · outbound

This paper cites Body of Her: A Preliminary Study on End-to-End Humanoid Agent.

SViMo: Synchronized Diffusion for Video and Motion Generation in Hand-object Interaction Scenarios Body of Her: A Preliminary Study on End-to-End Humanoid Agent

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:59.252913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:59.252913Z digest=sha256:198718ed6450210f3cb49193cfd599aa4b355906097a480642c9d1d30717d2b7

Observation f7f07dd0-afa8-4423-9915-89916cf3086f · outbound

This paper cites Qwen2.5-VL Technical Report.

SViMo: Synchronized Diffusion for Video and Motion Generation in Hand-object Interaction Scenarios Qwen2.5-VL Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:59.438395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:59.438395Z digest=sha256:db88fd91b8556c71241b40f3358d70efffaf1f77a73bf7b0f8a17258b0f04b9f

Observation f7cb1d75-dffb-4694-8c49-a23ef5ff2ab0 · outbound

This paper cites Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets.

SViMo: Synchronized Diffusion for Video and Motion Generation in Hand-object Interaction Scenarios Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:59.510279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:59.510279Z digest=sha256:eacd4f193838058ea19bd6a8b441e5452396678754959bcf60f5d388c3ec1757

Observation 7ddaac18-9d8d-4a40-85d2-b7c3891e1941 · outbound

This paper cites Physically plausible full-body hand-object interaction synthesis.

SViMo: Synchronized Diffusion for Video and Motion Generation in Hand-object Interaction Scenarios Physically plausible full-body hand-object interaction synthesis

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:29:15.278550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:28:59.562898Z digest=sha256:f7eeca6eed48e4b06eeb027c6b955f2264b5514213ac268d5d46266f6c141bca

Observation 88f12b54-76ff-4f82-99a5-afcfbbf7ba93 · outbound

This paper cites Video generation models as world simulators.

SViMo: Synchronized Diffusion for Video and Motion Generation in Hand-object Interaction Scenarios Video generation models as world simulators

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:59.634748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:59.634748Z digest=sha256:51d44eac15c1d761d5f2ff89cfdb6a63ada24947df2325df7a3c3336a206f805

Observation 376c7fd6-a8d1-4fe7-8652-8b4a68d62048 · outbound

This paper cites Emerging properties in self-supervised vision transformers.

SViMo: Synchronized Diffusion for Video and Motion Generation in Hand-object Interaction Scenarios Emerging properties in self-supervised vision transformers

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:59.770914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:59.770914Z digest=sha256:8eb72f7afd752cb208ebdb2da59f9b4490933808480d9555c7bdae440b27d73b

Observation 6068e9ad-7fe3-423c-b148-398a55b5d045 · outbound

This paper cites Text2hoi: Text-guided 3d motion generation for hand-object interaction.

SViMo: Synchronized Diffusion for Video and Motion Generation in Hand-object Interaction Scenarios Text2hoi: Text-guided 3d motion generation for hand-object interaction

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:59.890753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:59.890753Z digest=sha256:9132ae376c4193ed7ec5bf37f598e5db5f1c0e0cfb1983d83a20eb2ce8c37867

Observation 1499def3-c35a-4461-80af-35ccc4d7ead7 · outbound

This paper cites Dexycb: A benchmark for capturing hand grasping of objects.

SViMo: Synchronized Diffusion for Video and Motion Generation in Hand-object Interaction Scenarios Dexycb: A benchmark for capturing hand grasping of objects

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:59.947785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:59.947785Z digest=sha256:ebcd8f19a0eba9c9605090bf5f86aa9dd766bc480f0804163a398c5807a1ad81

Observation 3c290931-61bb-4ed0-853e-ddf3935892c2 · outbound

This paper cites Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks.

SViMo: Synchronized Diffusion for Video and Motion Generation in Hand-object Interaction Scenarios Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T11:29:00.054337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:29:00.054337Z digest=sha256:abbbaa763b863c9b41cd2062c998d75bafa91e95652731d3b992e94eb1defd88

Observation 0e3d241b-db13-453f-a9ba-7ba44c20972f · outbound

This paper cites Diverse human motion prediction via gumbel-softmax sampling from an auxiliary space.

SViMo: Synchronized Diffusion for Video and Motion Generation in Hand-object Interaction Scenarios Diverse human motion prediction via gumbel-softmax sampling from an auxiliary space

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:29:14.906162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:29:00.205296Z digest=sha256:ef35fb5bbf00d7a5705b383d7d585c8da6a94ff5c6c3b530099ab1087a88cf3f

Observation e6b4d1d3-6fa3-4107-ad08-2a7ae6231d7e · outbound

This paper cites Cg-hoi: Contact-guided 3d human-object interaction generation.

SViMo: Synchronized Diffusion for Video and Motion Generation in Hand-object Interaction Scenarios Cg-hoi: Contact-guided 3d human-object interaction generation

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:29:14.670755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:29:00.320954Z digest=sha256:5f73308a3a16dbf43a1fee480b1b1808cd4101f706e60f92e4b5e9046fe40f78

Observation 6bbf3627-d530-4fca-b477-e617644b11f4 · outbound

This paper cites Arctic: A dataset for dexterous bimanual hand-object manipulation.

SViMo: Synchronized Diffusion for Video and Motion Generation in Hand-object Interaction Scenarios Arctic: A dataset for dexterous bimanual hand-object manipulation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T11:29:00.418663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:29:00.418663Z digest=sha256:6aee9647bc3d29c5c4859fcfe734fd42acbca364293c815564a8138086cae6f3

Observation b7e3a644-57d3-451e-9ec5-c6e17944cbb1 · outbound

This paper cites Coohoi: Learning cooperative human-object interaction with manipulated object dynamics.

SViMo: Synchronized Diffusion for Video and Motion Generation in Hand-object Interaction Scenarios Coohoi: Learning cooperative human-object interaction with manipulated object dynamics

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:29:14.456528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:29:00.565115Z digest=sha256:6a553390603524028a335b7ae67a9bb117a51093c13c44bf80fa6512bbc61b88

Observation 15868f4e-08f4-4729-b0e6-44e50208efcf · outbound

This paper cites Prediction with action: Visual policy learning via joint denoising process.

SViMo: Synchronized Diffusion for Video and Motion Generation in Hand-object Interaction Scenarios Prediction with action: Visual policy learning via joint denoising process

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:29:14.221492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:29:00.671840Z digest=sha256:8e4ebd54f7ccd528c8eb2490766a26e560bc1852be5165490d814fe5ee94c1dd

Observation a43d552e-a8ea-46b4-8de5-27ee47bc0eaf · outbound

This paper cites Stochastic scene-aware motion prediction.

SViMo: Synchronized Diffusion for Video and Motion Generation in Hand-object Interaction Scenarios Stochastic scene-aware motion prediction

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T11:29:00.783005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:29:00.783005Z digest=sha256:a9533c5640d49ed638c71267cfedd0899bb892f9c742c605cf67070d06b265bc

Observation 90dfb212-eea2-47ec-8c03-fd8cfe0030ff · outbound

This paper cites Denoising diffusion probabilistic models.

SViMo: Synchronized Diffusion for Video and Motion Generation in Hand-object Interaction Scenarios Denoising diffusion probabilistic models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T11:29:00.916542Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:29:00.916542Z digest=sha256:8c77d1f7b31de652416d19accd0cbe55b6e91d3526833f047f9ac2d247b8e572

Observation d37851e6-0b83-45df-92d5-90ef879c4e3b · outbound

This paper cites Cogvideo: Large-scale pretraining for text-to-video generation via transformers.

SViMo: Synchronized Diffusion for Video and Motion Generation in Hand-object Interaction Scenarios Cogvideo: Large-scale pretraining for text-to-video generation via transformers

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:29:13.932132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:29:01.016743Z digest=sha256:63e43b74130cc3c9201b7c75715f68e35f1f2c87538e7f7b2d894a3a8aa152f1

Observation af3ea639-e121-414f-a1c8-8489ba011bc3 · outbound

This paper cites Animate anyone: Consistent and controllable image-to-video synthesis for character animation.

SViMo: Synchronized Diffusion for Video and Motion Generation in Hand-object Interaction Scenarios Animate anyone: Consistent and controllable image-to-video synthesis for character animation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T11:29:01.156559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:29:01.156559Z digest=sha256:cda62c3b2959a3b71f7ecc92eb1f4a640101e65e0be30fba235a9337999b51ee

Observation a3d023e0-872a-472c-af94-869ef263f45a · outbound

This paper cites Vbench: Comprehensive benchmark suite for video generative models.

SViMo: Synchronized Diffusion for Video and Motion Generation in Hand-object Interaction Scenarios Vbench: Comprehensive benchmark suite for video generative models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T11:29:01.232642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:29:01.232642Z digest=sha256:f0468bc8ca272b1e8e7860a268b6bda5531a040d77fad3cee50e02c64127223d

Observation 1f83025c-a539-4be8-b4d9-e1537f72974d · outbound

This paper cites The power of sound (tpos): Audio reactive video generation with stable diffusion.

SViMo: Synchronized Diffusion for Video and Motion Generation in Hand-object Interaction Scenarios The power of sound (tpos): Audio reactive video generation with stable diffusion

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:29:13.679724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:29:01.343728Z digest=sha256:ad1e5f8c4b8a3daa4b8b018484d27bb5ad336f1595cac352febae120429915a4

Observation 862781df-f76c-4a6c-8594-e3c06f35902c · outbound

This paper cites VACE: All-in-One Video Creation and Editing.

SViMo: Synchronized Diffusion for Video and Motion Generation in Hand-object Interaction Scenarios VACE: All-in-One Video Creation and Editing

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T11:29:01.468907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:29:01.468907Z digest=sha256:da23c4be4fa276a1de4945f3e64a31c97b962525e35716996aba2fde28b66fe7

Observation f1fa39dd-4b4b-4d98-aba5-f0d28ec6b820 · outbound

This paper cites HunyuanVideo: A Systematic Framework For Large Video Generative Models.

SViMo: Synchronized Diffusion for Video and Motion Generation in Hand-object Interaction Scenarios HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T11:29:01.608522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:29:01.608522Z digest=sha256:eb0650365830db179cd043dd1ea378fbe8cb5ceec68879c57d27570d40854764

Observation 6509e18e-422a-40d1-978e-9d2b3bddf12f · outbound

This paper cites Nifty: Neural object interaction fields for guided human motion synthesis.

SViMo: Synchronized Diffusion for Video and Motion Generation in Hand-object Interaction Scenarios Nifty: Neural object interaction fields for guided human motion synthesis

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:29:13.461294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:29:01.712935Z digest=sha256:520d88a99f8fef7fb14a6c0c30d15883d1fd1849835975ebc76d5eee97a67d95

Observation 704b2776-f350-4a1e-8336-33cfe9d43e40 · outbound

This paper cites Interhandgen: Two-hand interaction generation via cascaded reverse diffusion.

SViMo: Synchronized Diffusion for Video and Motion Generation in Hand-object Interaction Scenarios Interhandgen: Two-hand interaction generation via cascaded reverse diffusion

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:29:13.229870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:29:01.841671Z digest=sha256:abebef1afe8e6d1ff0f312dd73fcebc0c5d6ac8dc463a08e234df159d8e6cdbd

Observation 6061c08c-5794-4aa5-b5ad-23aa0503fb3b · outbound

This paper cites Controllable human-object interaction synthesis.

SViMo: Synchronized Diffusion for Video and Motion Generation in Hand-object Interaction Scenarios Controllable human-object interaction synthesis

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T11:29:01.973487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:29:01.973487Z digest=sha256:6a8f1c9bc8c22cd0d67e1a2f8bf73875cb0f841989b95314d4d725863214db85

Observation 4ad78e16-7d2f-42f1-8162-63c01103766e · outbound

This paper cites Object motion guided human motion synthesis.

SViMo: Synchronized Diffusion for Video and Motion Generation in Hand-object Interaction Scenarios Object motion guided human motion synthesis

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T11:29:02.099180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:29:02.099180Z digest=sha256:b8dc7f980a7d6973938fcb06daab1e42f9f903328c04860e949571efe106ab9b

Observation 1a37028e-a14d-45fc-a6c3-28b8ecf7f0fd · outbound

This paper cites Task-oriented human-object interactions generation with implicit neural representations.

SViMo: Synchronized Diffusion for Video and Motion Generation in Hand-object Interaction Scenarios Task-oriented human-object interactions generation with implicit neural representations

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:29:13.077736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:29:02.159259Z digest=sha256:925fb9074dfa565a425a0e4cbf105061b35d54a101c2dd650cf739d5f02c798e

Observation 778cd270-11c7-4e00-9a06-cdefb38374af · outbound

This paper cites Vision-language foundation models as effective robot imitators.

SViMo: Synchronized Diffusion for Video and Motion Generation in Hand-object Interaction Scenarios Vision-language foundation models as effective robot imitators

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:29:12.879061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:29:02.272037Z digest=sha256:deccb320ca9f2a112c4fa7bdae93d203c6d24e5f58be03f84e5a5eb7b05b47ff

Observation 7841240a-2b58-4190-a8ce-a0cebf37014c · outbound

This paper cites Amt: All-pairs multi-field transforms for efficient frame interpolation.

SViMo: Synchronized Diffusion for Video and Motion Generation in Hand-object Interaction Scenarios Amt: All-pairs multi-field transforms for efficient frame interpolation

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:29:12.671992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:29:02.412088Z digest=sha256:1b0b63d21ce347d93c2b62651b79b289dddfcbb40d241558e092237aacf58ad9

Observation c44dbbad-f78c-496b-8e5c-8b778324d395 · outbound

This paper cites Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023.

SViMo: Synchronized Diffusion for Video and Motion Generation in Hand-object Interaction Scenarios Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T11:29:02.518157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:29:02.518157Z digest=sha256:e7138d03eae1ae0ad7342fd174a3263805b9e222a0038f92f23038c32b662664

Observation 28d60e2c-6b2c-4f67-b468-3f19292ebae9 · outbound

This paper cites Javisdit: Joint audio-video diffusion transformer with hierarchical spatio-temporal prior synchronization.

SViMo: Synchronized Diffusion for Video and Motion Generation in Hand-object Interaction Scenarios Javisdit: Joint audio-video diffusion transformer with hierarchical spatio-temporal prior synchronization

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T11:29:02.653794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:29:02.653794Z digest=sha256:90a0f19715105b7c87f6f24dd937085fd15e818203770eec2c1bfbc4e2caf77c

Observation ba4503df-60bb-47f6-a857-5313a4c1294a · outbound

This paper cites Primitive-based 3d human-object interaction modelling and programming.

SViMo: Synchronized Diffusion for Video and Motion Generation in Hand-object Interaction Scenarios Primitive-based 3d human-object interaction modelling and programming

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:29:12.407256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:29:02.756572Z digest=sha256:b8a945b50a4152c2dd58bac04bdab1c68977a7659736feb075d71b090a3016ea

Observation f0c0aa2d-1aeb-4973-8e75-5fa41bda1d29 · outbound

This paper cites Geneoh diffusion: Towards generalizable hand-object interaction denoising via denoising diffusion.

SViMo: Synchronized Diffusion for Video and Motion Generation in Hand-object Interaction Scenarios Geneoh diffusion: Towards generalizable hand-object interaction denoising via denoising diffusion

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:29:12.179360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:29:02.865643Z digest=sha256:40a5d5688f2b68734a6fbcafc0fcfb1af2af1b1b09f3e14008f6fd3eb3a861d6

Observation 7bc42f5b-ea4c-4053-8ee7-2cdedb949707 · outbound

This paper cites Taco: Benchmarking generalizable bimanual tool-action-object understanding.

SViMo: Synchronized Diffusion for Video and Motion Generation in Hand-object Interaction Scenarios Taco: Benchmarking generalizable bimanual tool-action-object understanding

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T11:29:02.945310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:29:02.945310Z digest=sha256:572bc4fcea926930a9d2d80ec4546e67677cd37605c3ca8e78bbd3899dc0d634

Observation 15c7421d-9dfc-4fcf-ac89-347c671da035 · outbound

This paper cites Hoi4d: A 4d egocentric dataset for category-level human-object interaction.

SViMo: Synchronized Diffusion for Video and Motion Generation in Hand-object Interaction Scenarios Hoi4d: A 4d egocentric dataset for category-level human-object interaction

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:29:11.934837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:29:03.014251Z digest=sha256:613ea270d4ba563d2bc67302200d6fb1d26736af4a3f791d312251d5024d7130

Observation 0d7158d8-3810-49d5-b17b-621d14f5f181 · outbound

This paper cites Visual-RFT: Visual Reinforcement Fine-Tuning.

SViMo: Synchronized Diffusion for Video and Motion Generation in Hand-object Interaction Scenarios Visual-RFT: Visual Reinforcement Fine-Tuning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T11:29:03.126140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:29:03.126140Z digest=sha256:9f3a37d38afe430741d85716a72b4606c42c4b4aea9ef9299dc9fdd86d7f5ccf

Observation 025c8fac-bb0a-4fe7-a7b7-65b6a93e7be7 · outbound

This paper cites Omnigrasp: Grasping diverse objects with simulated humanoids.

SViMo: Synchronized Diffusion for Video and Motion Generation in Hand-object Interaction Scenarios Omnigrasp: Grasping diverse objects with simulated humanoids

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:29:11.673914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:29:03.223803Z digest=sha256:db6d852ccb3d87e6552ff594b537ad1485aa54e84049e265828b8238a5a7bdcf

Observation f42461c3-a015-461e-af72-4fcd54a1e1c8 · outbound

This paper cites Nerf: Representing scenes as neural radiance fields for view synthesis.

SViMo: Synchronized Diffusion for Video and Motion Generation in Hand-object Interaction Scenarios Nerf: Representing scenes as neural radiance fields for view synthesis

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T11:29:03.376595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:29:03.376595Z digest=sha256:bca512590b2b7191e318f6bf22a6cb9589a53e06499aca7112042cef41ee4c7a

Observation b662a0da-3f5e-4073-81af-8693d78100ae · outbound

This paper cites ManiVideo: Generating Hand-Object Manipulation Video with Dexterous and Generalizable Grasping.

SViMo: Synchronized Diffusion for Video and Motion Generation in Hand-object Interaction Scenarios ManiVideo: Generating Hand-Object Manipulation Video with Dexterous and Generalizable Grasping

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T11:29:03.505575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:29:03.505575Z digest=sha256:a3677c314614ea4d67ba649f5b56e04df43c48fe50dbcb7d4b3d81c6daff09c3

Observation a2198627-a368-4d79-9c2b-c0ef489fbc62 · outbound

This paper cites Scalable diffusion models with transformers.

SViMo: Synchronized Diffusion for Video and Motion Generation in Hand-object Interaction Scenarios Scalable diffusion models with transformers

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T11:29:03.652566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:29:03.652566Z digest=sha256:ab4d3db4dbcb51b978c80fcebccca5e6fa7d03fdf9970adaa6022d8ae5cc26d9

Observation 93224f88-b4bc-4486-91b4-98431024c3ed · outbound

This paper cites HOI-Diff: Text-Driven Synthesis of 3D Human-Object Interactions using Diffusion Models.

SViMo: Synchronized Diffusion for Video and Motion Generation in Hand-object Interaction Scenarios HOI-Diff: Text-Driven Synthesis of 3D Human-Object Interactions using Diffusion Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T11:29:03.743604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:29:03.743604Z digest=sha256:9cd61b0a6a4117da802b82845cd83625978bb37356426201dce4af65d3dac1fe

Observation 1ff6d839-d05f-4233-b385-999e7d9e8697 · outbound

This paper cites Hierarchical generation of human- object interactions with diffusion probabilistic models.

SViMo: Synchronized Diffusion for Video and Motion Generation in Hand-object Interaction Scenarios Hierarchical generation of human- object interactions with diffusion probabilistic models

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:29:11.441585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:29:03.835908Z digest=sha256:2b7ecab87ef8d09435f67d83c0e1020aa35d02f9cb8768b86a49f6849b0ceec3

Observation 6b567976-c336-4f54-aa2c-799edab2aa0c · outbound

This paper cites Learning transferable visual models from natural language supervision.

SViMo: Synchronized Diffusion for Video and Motion Generation in Hand-object Interaction Scenarios Learning transferable visual models from natural language supervision

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T11:29:03.951478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:29:03.951478Z digest=sha256:9d207c2ab134560440ec56382a606c0d53b360c28fd6b6d1d8b58e84ae471677

Observation 2bb6fe21-95fe-478b-a244-617248d39e83 · outbound

This paper cites Exploring the limits of transfer learning with a unified text-to-text transformer.

SViMo: Synchronized Diffusion for Video and Motion Generation in Hand-object Interaction Scenarios Exploring the limits of transfer learning with a unified text-to-text transformer

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T11:29:04.037964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:29:04.037964Z digest=sha256:c2ccf93e82cd2441772d9175d2cc3bd7535a009b0b4cc58c2effe229a170c5d4

Observation 38c69bab-f3bb-4ad7-872f-d312d59ab737 · outbound

This paper cites Zero: Memory optimizations toward training trillion parameter models.

SViMo: Synchronized Diffusion for Video and Motion Generation in Hand-object Interaction Scenarios Zero: Memory optimizations toward training trillion parameter models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T11:29:04.208113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:29:04.208113Z digest=sha256:aceb42f7b8badabcc60ec06435a15e9f36ed3ee6bcb83075813a7dba81998f21

Observation 270da79e-0c15-4a3c-96af-7fa997ca26ee · outbound

This paper cites Zero-shot text-to-image generation.

SViMo: Synchronized Diffusion for Video and Motion Generation in Hand-object Interaction Scenarios Zero-shot text-to-image generation

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T11:29:04.351168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:29:04.351168Z digest=sha256:d1fbc175f07cd9c0bd62dae17381b95e3e5d53146dd575f0fe5b43f47a90a3dc

Observation 5499978a-9eff-4e8f-9a75-3232ac3ca477 · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

SViMo: Synchronized Diffusion for Video and Motion Generation in Hand-object Interaction Scenarios High-resolution image synthesis with latent diffusion models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T11:29:04.461811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:29:04.461811Z digest=sha256:80bde203d3a463521ae2d7d19290e122bef4b77ddd68e3f95bfcf0cb215f2a02

Observation 600a4446-509d-4136-b0ea-d0fc9cdc6c4d · outbound

This paper cites Photorealistic text-to- image diffusion models with deep language understanding.

SViMo: Synchronized Diffusion for Video and Motion Generation in Hand-object Interaction Scenarios Photorealistic text-to- image diffusion models with deep language understanding

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T11:29:04.570786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:29:04.570786Z digest=sha256:f655510c0e384a7b557781163261cdbb95d70c89c306c5bfbdad772d1b2afc04

Observation 5ed1133c-a3bf-4f78-b182-ae50846e14a6 · outbound

This paper cites Hand-Object Interaction Pretraining from Videos.

SViMo: Synchronized Diffusion for Video and Motion Generation in Hand-object Interaction Scenarios Hand-Object Interaction Pretraining from Videos

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T11:29:04.683935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:29:04.683935Z digest=sha256:753d16003017fdc0ebe935a3727e8699c3ae0eb9f0b8af4ac5eeff73ae4d4bea

Observation 50def3c1-627d-4f21-9cba-998e4202f6ef · outbound

This paper cites Grab: A dataset of whole-body human grasping of objects.

SViMo: Synchronized Diffusion for Video and Motion Generation in Hand-object Interaction Scenarios Grab: A dataset of whole-body human grasping of objects

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:29:11.084079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:29:04.800452Z digest=sha256:f01d9f7060115aa87396938648156b473584340aeebf182d0f75051492f4bd24

Observation c3803749-a5dd-449b-bc5f-ada3713a056b · outbound

This paper cites Any-to-any generation via composable diffusion.

SViMo: Synchronized Diffusion for Video and Motion Generation in Hand-object Interaction Scenarios Any-to-any generation via composable diffusion

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:29:10.833308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:29:04.885671Z digest=sha256:8e2f81c013fee52a54b1989b91a0f5221cb0909d93b0d06c8fca4e3a7c3d53e7

Observation b29a7c5c-0251-4a74-b1b3-e00ba1d0f11e · outbound

This paper cites Raft: Recurrent all-pairs field transforms for optical flow.

SViMo: Synchronized Diffusion for Video and Motion Generation in Hand-object Interaction Scenarios Raft: Recurrent all-pairs field transforms for optical flow

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T11:29:04.953168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:29:04.953168Z digest=sha256:7f17dccb35dba5717452cbeef73d1a2f7c6f724d3cdfb7b0fdd830206b809eeb

Observation c2afb283-976e-4948-bcd5-32f5e6efde19 · outbound

This paper cites Human motion diffusion model.

SViMo: Synchronized Diffusion for Video and Motion Generation in Hand-object Interaction Scenarios Human motion diffusion model

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T11:29:05.028705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:29:05.028705Z digest=sha256:a53412f0143fe9b6860ab913acb6844862fc467df69cd059a202c44b4142927e

Observation 08143ec8-ca17-4dfb-8927-5e4083f58fb1 · outbound

This paper cites Deepsimho: Stable pose estimation for hand-object interaction via physics simulation.

SViMo: Synchronized Diffusion for Video and Motion Generation in Hand-object Interaction Scenarios Deepsimho: Stable pose estimation for hand-object interaction via physics simulation

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:29:10.543054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:29:05.104217Z digest=sha256:ae35e47a933b17bd6b05f01cdf0d05f3e4098aaacef864d7e5be718232b5d2c3

Observation d19071c2-fc6a-4069-b907-8c35f059d613 · outbound

This paper cites Cogvlm: Visual expert for pretrained language models.

SViMo: Synchronized Diffusion for Video and Motion Generation in Hand-object Interaction Scenarios Cogvlm: Visual expert for pretrained language models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T11:29:05.143860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:29:05.143860Z digest=sha256:8adf403a4a2c7e287ce717dfdddf856c8e6e5374527dbd8bd689d5070591b5a6

Observation 7d32ac8f-6ed1-4526-9d2e-d7aa85ee4bad · outbound

This paper cites PhysHOI: Physics-Based Imitation of Dynamic Human-Object Interaction.

SViMo: Synchronized Diffusion for Video and Motion Generation in Hand-object Interaction Scenarios PhysHOI: Physics-Based Imitation of Dynamic Human-Object Interaction

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T11:29:05.220366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:29:05.220366Z digest=sha256:fc84d97c4c4b68c1e079f00e2b7fc3cc60a8ca532aaafc4b7b0a1fc30d61b46b

Observation 50bdc760-77cb-44a8-8dbf-ef1cc691b94a · outbound

This paper cites Easyanimate: A high-performance long video generation method based on transformer architecture.

SViMo: Synchronized Diffusion for Video and Motion Generation in Hand-object Interaction Scenarios Easyanimate: A high-performance long video generation method based on transformer architecture

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T11:29:05.314134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:29:05.314134Z digest=sha256:a8324cbe1dbc90c5a3b178fb15b522ea946b6e9a25060471785d005db1446876

Observation 5a401baa-84e2-423e-9cf6-f8262337e588 · outbound

This paper cites Interdiff: Generating 3d human-object interactions with physics-informed diffusion.

SViMo: Synchronized Diffusion for Video and Motion Generation in Hand-object Interaction Scenarios Interdiff: Generating 3d human-object interactions with physics-informed diffusion

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T11:29:05.380924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:29:05.380924Z digest=sha256:607381b9ae1fee28a3221315a2ad85d0db57ea33d5803a8497590fc50e22528f

Observation e62e22f1-2c2e-47be-99ec-8931538b9667 · outbound

This paper cites Intermimic: Towards universal whole-body control for physics-based human-object interactions.

SViMo: Synchronized Diffusion for Video and Motion Generation in Hand-object Interaction Scenarios Intermimic: Towards universal whole-body control for physics-based human-object interactions

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:29:10.325426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:29:05.471978Z digest=sha256:3432f039c065a7fdba4e8729d5ca40c1bba5ebf62418fa2d054f9fb657aa73da

Observation def8529f-e32e-4ab7-8cb7-e31f92890363 · outbound

This paper cites Interdreamer: Zero-shot text to 3d dynamic human-object interaction.

SViMo: Synchronized Diffusion for Video and Motion Generation in Hand-object Interaction Scenarios Interdreamer: Zero-shot text to 3d dynamic human-object interaction

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:29:10.117411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:29:05.546487Z digest=sha256:7089f95c8dd93cc82ccbf82f3bac2476973066afad1463b86ac8f658d74fce88

Observation 772d192a-2603-4302-838c-3c549f6d46c5 · outbound

This paper cites Magicanimate: Temporally consistent human image animation using diffusion model.

SViMo: Synchronized Diffusion for Video and Motion Generation in Hand-object Interaction Scenarios Magicanimate: Temporally consistent human image animation using diffusion model

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T11:29:05.632537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:29:05.632537Z digest=sha256:a8ffd076c1d77db87eba5d8bdc8473eab2ba3b66693f2f2139bfae3838aae046

Observation 96c659a6-e3e8-4831-98cc-1275780fb212 · outbound

This paper cites AnchorCrafter: Animate Cyber-Anchors Selling Your Products via Human-Object Interacting Video Generation.

SViMo: Synchronized Diffusion for Video and Motion Generation in Hand-object Interaction Scenarios AnchorCrafter: Animate Cyber-Anchors Selling Your Products via Human-Object Interacting Video Generation

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T11:29:05.702331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:29:05.702331Z digest=sha256:676975ec99f966535674bd6967b829bb2d4fa9a33b745aa0aa9b8e48cd452777

Observation 84744637-242e-4531-bfcb-479fa12ea8cf · outbound

This paper cites Oakink: A large- scale knowledge repository for understanding hand-object interaction.

SViMo: Synchronized Diffusion for Video and Motion Generation in Hand-object Interaction Scenarios Oakink: A large- scale knowledge repository for understanding hand-object interaction

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:29:09.810360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:29:05.797285Z digest=sha256:4cab1419dc5cf0e92d6b974007b2374c7124c880faac1bb26f16f5c2560530f9

Observation fc7a88ea-b638-4a6f-ab1d-2ad03c9da93a · outbound

This paper cites The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision).

SViMo: Synchronized Diffusion for Video and Motion Generation in Hand-object Interaction Scenarios The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision)

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T11:29:05.840015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:29:05.840015Z digest=sha256:a1d20d707304eb49fd66f2d1ca7369e7778ed753cb816c8f9b0de3c724a89d5a

Observation 3a7b274a-0c18-4a97-b69e-09c546b768a9 · outbound

This paper cites Cogvideox: Text-to-video diffusion models with an expert transformer.

SViMo: Synchronized Diffusion for Video and Motion Generation in Hand-object Interaction Scenarios Cogvideox: Text-to-video diffusion models with an expert transformer

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:29:09.542934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:29:05.870422Z digest=sha256:1f55cc66e2d1ea71fa95dc066ad8448f73bc5e51bff859b7a389a571d91bb263

Observation 864b99ea-f983-4e63-bbd3-0784000f2e8a · outbound

This paper cites Diverse and aligned audio-to- video generation via text-to-video model adaptation.

SViMo: Synchronized Diffusion for Video and Motion Generation in Hand-object Interaction Scenarios Diverse and aligned audio-to- video generation via text-to-video model adaptation

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:29:09.302029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:29:06.010511Z digest=sha256:3909b0abc263ab3d64c7197f007fc20670581b631ecdcbc09ac1bf824d1af824

Observation f2d41a80-24d2-4dac-a3ff-2acb967e47d9 · outbound

This paper cites Oakink2: A dataset of bimanual hands-object manipulation in complex task completion.

SViMo: Synchronized Diffusion for Video and Motion Generation in Hand-object Interaction Scenarios Oakink2: A dataset of bimanual hands-object manipulation in complex task completion

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:29:08.986003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:29:06.121308Z digest=sha256:bfa9c4853da900b589ae686f87c5d89f23877d2df185bf8c43c37f2dfe2f8955

Observation 34b26231-8c05-45e5-9c00-b99e66179bbb · outbound

This paper cites ManiDext: Hand-Object Manipulation Synthesis via Continuous Correspondence Embeddings and Residual-Guided Diffusion.

SViMo: Synchronized Diffusion for Video and Motion Generation in Hand-object Interaction Scenarios ManiDext: Hand-Object Manipulation Synthesis via Continuous Correspondence Embeddings and Residual-Guided Diffusion

Reference 68

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:29:07.077298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:29:06.314273Z digest=sha256:c280c34a0ff599bbe31fd7f3effce150f03367165613356a4a3ae794ff32a851

Observation f2db1453-ab85-47f9-97a0-0bb598b0199e · outbound

This paper cites Couch: Towards controllable human-chair interactions.

SViMo: Synchronized Diffusion for Video and Motion Generation in Hand-object Interaction Scenarios Couch: Towards controllable human-chair interactions

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:29:08.711736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:29:06.492979Z digest=sha256:34daf711ac31abca68cbeea420152617b86830741822fd6aee96399f97550448

Observation 8ff90525-bead-41b8-b02f-8533fb05c192 · outbound

This paper cites Emdm: Efficient motion diffusion model for fast and high-quality motion generation.

SViMo: Synchronized Diffusion for Video and Motion Generation in Hand-object Interaction Scenarios Emdm: Efficient motion diffusion model for fast and high-quality motion generation

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:29:08.378340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:29:06.643675Z digest=sha256:448ad3f71c0f03e3b4bdff7540815dd935f115217aa67f59ffdc86617ff3fc83

Observation 3b4a5d85-8cb1-42a5-98ed-9a6c23b3703f · outbound

This paper cites Champ: Controllable and consistent human image animation with 3d parametric guidance.

SViMo: Synchronized Diffusion for Video and Motion Generation in Hand-object Interaction Scenarios Champ: Controllable and consistent human image animation with 3d parametric guidance

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:29:08.084227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:29:06.800490Z digest=sha256:caa9fd6ab1ed5a2f566bede4e6ea48772d9003f0b0b47d98ca4b559fe279c4ce

Observation 06640dcd-bcb3-4c4b-92ab-e3a65c68a1ad · outbound

This paper cites Rt-2: Vision-language-action models transfer web knowledge to robotic control.

SViMo: Synchronized Diffusion for Video and Motion Generation in Hand-object Interaction Scenarios Rt-2: Vision-language-action models transfer web knowledge to robotic control

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:29:07.738463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:29:06.878705Z digest=sha256:5f2c9f808e20ed0ac9af358bca2cd1ca4810ab270ca0ced30b466fb5a21f5ecd

Pith citing papers

Observation 3e3cca65-46dd-480f-ade2-e5256b5c6423 · inbound

HVG-3D: Bridging Real and Simulation Domains for 3D-Conditional Hand-Object Interaction Video Synthesis cites this paper.

HVG-3D: Bridging Real and Simulation Domains for 3D-Conditional Hand-Object Interaction Video Synthesis SViMo: Synchronized Diffusion for Video and Motion Generation in Hand-object Interaction Scenarios

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-14T00:18:30.110150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-14T00:14:36.537688Z digest=sha256:11aab7ec641194b80c21af21c44673c45d70a89ff9cf1d91e2614c4674ac0982

Observation f8beea9c-8d04-4da1-b448-79dfe4f7dd5b · inbound

StreamHOI: Interaction-aware Temporal Memory Adaptation for Streaming HOI Video Generation cites this paper.

StreamHOI: Interaction-aware Temporal Memory Adaptation for Streaming HOI Video Generation SViMo: Synchronized Diffusion for Video and Motion Generation in Hand-object Interaction Scenarios

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-01T10:37:55.063859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:37:55.063859Z digest=sha256:ca52ad63525fad37ed98c8863c7265083b1fdac9aa1f9615c4c584b3611e3c19

Observation a81105ac-ac46-4737-99dc-576b15490232 · inbound

AgentHOI: Multi-Agent Reasoning for Human-Object-Interaction Video Generation via Implicit Representation Alignment cites this paper.

AgentHOI: Multi-Agent Reasoning for Human-Object-Interaction Video Generation via Implicit Representation Alignment SViMo: Synchronized Diffusion for Video and Motion Generation in Hand-object Interaction Scenarios

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-01T05:27:16.175479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T05:27:16.175479Z digest=sha256:2d1d005a98019f61a0e49852d4015f9f9657d71371d97d426613b8368840a467

Observation c76c5d80-b73d-4dfc-8da3-743d5c59f131 · inbound

PhotoHOI: Synthesizing 3D Hand-Object Interactions from a Single RGB Photograph cites this paper.

PhotoHOI: Synthesizing 3D Hand-Object Interactions from a Single RGB Photograph SViMo: Synchronized Diffusion for Video and Motion Generation in Hand-object Interaction Scenarios

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-04T18:35:26.468228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T18:35:26.468228Z digest=sha256:ed5faa49433c5b1a110426ecaf6a12b5e64e945ed02a8bfb63928b977ed56b38

Observation 0cd1448c-0687-4bd5-9198-157da069b78e · inbound

PhotoHOI: Synthesizing 3D Hand-Object Interactions from a Single RGB Photograph cites this paper.

PhotoHOI: Synthesizing 3D Hand-Object Interactions from a Single RGB Photograph SViMo: Synchronized Diffusion for Video and Motion Generation in Hand-object Interaction Scenarios

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T04:10:14.672688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:10:14.672688Z digest=sha256:0102c79901309f5400e93b53d95a94896531a88cf8e575d406faaf59a960e1de