Pith. sign in

Paper Citation Record · LEDGER

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation

As of 7 August 2026, this Paper Citation Record lists 100 of 130 outbound references and 22 inbound Pith citation observations for arXiv:2505.20292.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.20292 v4

Coverage vector

measured 100 of 130 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:59:35.677086Z

measured 122 of 122 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 22 of 22 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T12:42:42.185508Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T21:10:09.694521Z

Reference resolution

100 of 130 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved99
  • parse uncertain1
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f622355a-fc11-46d9-8882-071363473815 · outbound

This paper cites GPT-4 Technical Report.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:27.582566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:27.582566Z digest=sha256:200e7d81f86819d70db017489e9f9dcc02976babf9e8d73d3066fc89d00e154d

Observation 674e39e3-f51d-4ec1-b250-274525a49cf0 · outbound

This paper cites Detecting ai-generated images using vision transformers: A robust approach for safeguarding visual media integrity.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation Detecting ai-generated images using vision transformers: A robust approach for safeguarding visual media integrity

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:27.677003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:27.677003Z digest=sha256:257efc0a5bdb1a48007a7a332384e5589f2c040187d91547305ae481d0396469

Observation b31a7f79-6b63-43d9-bd3b-b1ec684f5ca5 · outbound

This paper cites Qwen2.5-VL Technical Report.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation Qwen2.5-VL Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:27.789237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:27.789237Z digest=sha256:fb2c5c3d0435cf9f7fd23620872631332f5693a06fa1207db6da2955a4113f09

Observation d4d2a682-d6fc-4019-872e-3d68462c74af · outbound

This paper cites Frozen in time: A joint video and image encoder for end-to-end retrieval.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation Frozen in time: A joint video and image encoder for end-to-end retrieval

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:27.939740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:27.939740Z digest=sha256:8d053e29bb6aef713b13e3154ddd170e38836d3a2239d9ac33fc99843053547e

Observation e211f830-c275-4c09-812e-117d517cad8a · outbound

This paper cites Vidu: a Highly Consistent, Dynamic and Skilled Text-to-Video Generator with Diffusion Models.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation Vidu: a Highly Consistent, Dynamic and Skilled Text-to-Video Generator with Diffusion Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:28.048350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:28.048350Z digest=sha256:242c1f1df16a180f03625dc42f5343c8f7a4b432c2315f5d2838366fb481e3b0

Observation c83fa217-46b4-4b85-a1b9-10cffa1791f9 · outbound

This paper cites an unresolved cited work.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation Unresolved cited work

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:28.178541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:28.178541Z digest=sha256:b8b22ed17c538be037799cefb2b2214dd3f7e457a9c03ce5460d19651f9472ee

Observation eba92d87-f4a5-49c3-8088-45ad15c42153 · outbound

This paper cites DiTCtrl: Exploring Attention Control in Multi-Modal Diffusion Transformer for Tuning-Free Multi-Prompt Longer Video Generation.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation DiTCtrl: Exploring Attention Control in Multi-Modal Diffusion Transformer for Tuning-Free Multi-Prompt Longer Video Generation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:28.322361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:28.322361Z digest=sha256:d272e8d03ddae357d89aa16cf6f17491fb43d2ea06f2a0641d1ed66cd2566398

Observation 2b2e282a-64fc-4df7-8bcc-652b8e8d3538 · outbound

This paper cites MagicPose: Realistic Human Poses and Facial Expressions Retargeting with Identity-aware Diffusion.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation MagicPose: Realistic Human Poses and Facial Expressions Retargeting with Identity-aware Diffusion

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:28.549879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:28.549879Z digest=sha256:91e635685aa8ea837e1892e567d4eb2bb6ca1c36d11033aef5422bb0b7b809eb

Observation 5041d6bd-3208-4592-b253-afb119d6efc5 · outbound

This paper cites PhotoVerse: Tuning-Free Image Customization with Text-to-Image Diffusion Models.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation PhotoVerse: Tuning-Free Image Customization with Text-to-Image Diffusion Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:28.605625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:28.605625Z digest=sha256:980d6d7b9453d9bc907677f1f5d4613cac955cd18103b831358a64d987593eb6

Observation a83ac1d4-5098-4588-8898-770acbd2eee6 · outbound

This paper cites OD-VAE: An Omni-dimensional Video Compressor for Improving Latent Video Diffusion Model.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation OD-VAE: An Omni-dimensional Video Compressor for Improving Latent Video Diffusion Model

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:28.683440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:28.683440Z digest=sha256:9c59302ff35d4a25370377d908446e2f0233fed736f66ff0db2ffe6e8ba2c560

Observation 9a4bf5ea-8a1e-4037-9f18-93bd8d8d0591 · outbound

This paper cites Panda-70M: Captioning 70M Videos with Multiple Cross-Modality Teachers.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation Panda-70M: Captioning 70M Videos with Multiple Cross-Modality Teachers

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:28.739467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:28.739467Z digest=sha256:bb68d5f2159823190373df2bdab81274e9d4127e236169d6264f307e1b5c82ce

Observation 08a4a1b2-29ee-4e32-b9c5-a14526531a4f · outbound

This paper cites Multi-subject Open-set Personalization in Video Generation.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation Multi-subject Open-set Personalization in Video Generation

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:28.849584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:28.849584Z digest=sha256:ead5b98040cf948538f1fa99056d3bc3bc302bae1495830bae140dd7e1204127

Observation 32a1c487-0f0f-45d1-b6cb-035ad5f57224 · outbound

This paper cites UniReal: Universal Image Generation and Editing via Learning Real-world Dynamics.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation UniReal: Universal Image Generation and Editing via Learning Real-world Dynamics

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:28.964637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:28.964637Z digest=sha256:961e3cec7f7e0fcba65008ba37a5fa8a51eb721621de7d5b7005606befaa6139

Observation 9b8bc10a-163f-4858-ab5a-41425d925ea7 · outbound

This paper cites Yolo- world: Real-time open-vocabulary object detection.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation Yolo- world: Real-time open-vocabulary object detection

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:29.040098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:29.040098Z digest=sha256:c33fcd129494a04b517ea88a841d3213f232f6b84246256c14f9ae5818ed4b20

Observation a6948c17-7919-4f42-9b3a-095e838483d6 · outbound

This paper cites improved-aesthetic-predictor.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation improved-aesthetic-predictor

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:29.159161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:29.159161Z digest=sha256:1e15d4c802dab7363e56977f58c2fd8f4d56fd8ab0ca2021c385575941e41de0

Observation 8ea5b247-88ce-431e-86f4-3e07a6732a56 · outbound

This paper cites Arcface: Additive angular margin loss for deep face recognition.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation Arcface: Additive angular margin loss for deep face recognition

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:29.222747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:29.222747Z digest=sha256:f24197f95995a3b6a6619f086f45960d72fb5edc5a6fa92d77ea86ffadd031c7

Observation 7344fd96-2a37-4a56-9bff-dc864b12df8a · outbound

This paper cites CINEMA: Coherent Multi-Subject Video Generation via MLLM-Based Guidance.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation CINEMA: Coherent Multi-Subject Video Generation via MLLM-Based Guidance

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:29.343261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:29.343261Z digest=sha256:faf790b2ea6948d1701d60901d5ebae52a5593095ddd8cf63938536a769189ef

Observation 5f763547-6771-4408-af08-5709abd1f81d · outbound

This paper cites Worldscore: A unified evaluation benchmark for world generation.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation Worldscore: A unified evaluation benchmark for world generation

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:29.388166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:29.388166Z digest=sha256:a123671816f1c0a1d044c6327bf913935a65f9025776ad5dbe09b80823a6d7f4

Observation 78244fcb-96a4-4777-9e21-b732843a7d56 · outbound

This paper cites Vchitect-2.0: Parallel Transformer for Scaling Up Video Diffusion Models.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation Vchitect-2.0: Parallel Transformer for Scaling Up Video Diffusion Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:29.452382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:29.452382Z digest=sha256:4543646199606f7393e66928c97cc964dc123946aebcee2ba50f1363735e6253

Observation 87b78b0f-fefc-47bf-a394-e5f5bcd6680b · outbound

This paper cites SkyReels-A2: Compose Anything in Video Diffusion Transformers.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation SkyReels-A2: Compose Anything in Video Diffusion Transformers

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:29.507056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:29.507056Z digest=sha256:34b57f2733ec9047953884324a7eb0ef119a56a4a62a53d3748cb29f7ace2956

Observation 4038376a-8f8a-411c-83bd-87e66b361c56 · outbound

This paper cites Ingredients: Blending Custom Photos with Video Diffusion Transformers.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation Ingredients: Blending Custom Photos with Video Diffusion Transformers

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:29.585694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:29.585694Z digest=sha256:cca3ff0effb4e2303438a3d990c62ed2f1bbb63e32be887e5137bd7f96c01239

Observation 7e1227c4-7833-4794-8ce6-f823e278d27e · outbound

This paper cites AE-NeRF: Augmenting Event-Based Neural Radiance Fields for Non-ideal Conditions and Larger Scene.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation AE-NeRF: Augmenting Event-Based Neural Radiance Fields for Non-ideal Conditions and Larger Scene

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:29.672932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:29.672932Z digest=sha256:1f71e3c7ea69a2c9472b2396aa76d9d4eaaf1b0f1bc309936f19502598558c7d

Observation 0bcb0d86-c873-465d-9046-9c71c1b4ca33 · outbound

This paper cites An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual Inversion.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual Inversion

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:29.775085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:29.775085Z digest=sha256:8847526e31a8c8e58d6685d091e15631461e53484a0403b6a681408a75fa0e76

Observation c460b892-36fd-4c44-9c27-2e96a2b65ab3 · outbound

This paper cites AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:29.829842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:29.829842Z digest=sha256:b1bcf3fc0818cb21513303f7aa92cee07b81a1ff60a93776d878807a228e3923

Observation cef193aa-9fac-4a73-9318-e231dbffa46f · outbound

This paper cites PuLID: Pure and Lightning ID Customization via Contrastive Alignment.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation PuLID: Pure and Lightning ID Customization via Contrastive Alignment

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:29.869968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:29.869968Z digest=sha256:08c9ca7a5bedd0a349e681a181cda6e2a19ccb291daf718a390bcb1c46558f5e

Observation b9ca7979-c9c2-4e8a-8da3-5d1cf9e33c7f · outbound

This paper cites UniPortrait: A Unified Framework for Identity-Preserving Single- and Multi-Human Image Personalization.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation UniPortrait: A Unified Framework for Identity-Preserving Single- and Multi-Human Image Personalization

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:29.959239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:29.959239Z digest=sha256:358ce7f9cbff3217ef131777bc159a01b760aba55742ab600d258b6c1d50e1e5

Observation 2a96e897-15fa-4f12-9e15-faf72dbb4d0b · outbound

This paper cites VideoScore: Building Automatic Metrics to Simulate Fine-grained Human Feedback for Video Generation.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation VideoScore: Building Automatic Metrics to Simulate Fine-grained Human Feedback for Video Generation

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:30.049247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:30.049247Z digest=sha256:1f289a145cb38e2e777a22341127dd3a80513215e930f28ddde5557acd075c10

Observation 7b23639f-4312-4f64-a82f-555e6e29f70c · outbound

This paper cites ID-Animator: Zero-Shot Identity-Preserving Human Video Generation.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation ID-Animator: Zero-Shot Identity-Preserving Human Video Generation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:30.179985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:30.179985Z digest=sha256:0ec3b4cd429f6f52fe1f68f971b6e4e8f8dc555af1fa7b6abd0232d9435af58a

Observation 79771ac8-067f-45f3-b446-1b6261545c69 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation LoRA: Low-Rank Adaptation of Large Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:30.262848Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:30.262848Z digest=sha256:a4d0d9f0d1c85370290b0da3c641dc66c852eb970e34c2b3c7b11ba621d333b0

Observation b14917ee-acdd-488e-a78f-8696687ca2b8 · outbound

This paper cites Animate anyone: Consistent and controllable image-to-video synthesis for character animation.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation Animate anyone: Consistent and controllable image-to-video synthesis for character animation

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:30.325203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:30.325203Z digest=sha256:f4e1a4cfbd8de7d01b69418ce2f247ebd39dbf847af82883078a9c796684cf04

Observation 3f861033-8931-486c-b7df-ae3fe9cb8860 · outbound

This paper cites Animate Anyone 2: High-Fidelity Character Image Animation with Environment Affordance.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation Animate Anyone 2: High-Fidelity Character Image Animation with Environment Affordance

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:30.416393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:30.416393Z digest=sha256:f898509597415c308f2dfc0b4a2393d0dd8aea060bcd1e9b3cfbaa622c291abf

Observation 63e9d5a2-120e-4bf2-8687-2cb8e04ecf70 · outbound

This paper cites HunyuanCustom: A Multimodal-Driven Architecture for Customized Video Generation.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation HunyuanCustom: A Multimodal-Driven Architecture for Customized Video Generation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:30.485928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:30.485928Z digest=sha256:92294c0569f0892d1f3d1cfe5529a2beafafc50f575dc7c5d5efb7714b68df60

Observation d2626035-6181-445e-b482-e4885d7f8b11 · outbound

This paper cites Curricularface: adaptive curriculum learning loss for deep face recognition.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation Curricularface: adaptive curriculum learning loss for deep face recognition

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:30.553705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:30.553705Z digest=sha256:48f9c7f378586040104a3ce418a53681242d9d4b3967ed3000535f292a1d7df6

Observation ed245bbb-b229-4733-8b8c-c399a21ac318 · outbound

This paper cites ConceptMaster: Multi-Concept Video Customization on Diffusion Transformer Models Without Test-Time Tuning.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation ConceptMaster: Multi-Concept Video Customization on Diffusion Transformer Models Without Test-Time Tuning

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:30.608834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:30.608834Z digest=sha256:9faf55d685d1a1d4e247419e399713f5e36961c11eeeab8bd5601550bc07a661

Observation 65e3c255-602b-4982-bb4b-57471c45dd2c · outbound

This paper cites VBench: Comprehensive Benchmark Suite for Video Generative Models.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation VBench: Comprehensive Benchmark Suite for Video Generative Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:30.690138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:30.690138Z digest=sha256:a32421089a74c022c695e178387a079c2b60de1b8f5c12a0f46c450c4a4d6831

Observation 288b44fb-26ab-4086-8dd5-595bdb0df7fe · outbound

This paper cites VBench++: Comprehensive and Versatile Benchmark Suite for Video Generative Models.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation VBench++: Comprehensive and Versatile Benchmark Suite for Video Generative Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:30.771224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:30.771224Z digest=sha256:94f7eedc303eaeef6a9ac698d611eb2e80b4964e3f5cc14454ef24154dc5ec7b

Observation 0baf8290-2faa-4be9-bc22-e5438a041216 · outbound

This paper cites InfiniteYou: Flexible Photo Recrafting While Preserving Your Identity.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation InfiniteYou: Flexible Photo Recrafting While Preserving Your Identity

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:30.860050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:30.860050Z digest=sha256:180f0e96f1fa5b9087d571f8b2f56467d80090c9a9841db6db02450a4ac8ee33

Observation 7d0df9c7-c0c8-473b-a301-0bf36b3a1813 · outbound

This paper cites Videobooth: Diffusion-based video generation with image prompts.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation Videobooth: Diffusion-based video generation with image prompts

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:30.937726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:30.937726Z digest=sha256:6848e17f651e23b2e7e1088bd6497e4e8c85948aac5438351311d71494d83494

Observation 94ca3e3a-605d-4741-ace1-aa094a4bcc71 · outbound

This paper cites VACE: All-in-One Video Creation and Editing.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation VACE: All-in-One Video Creation and Editing

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:30.975711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:30.975711Z digest=sha256:885a57417cbd69a3b34cb8459a5179bc153131cb4d708c9e55ee3237e379c1c1

Observation d40dcba6-b12e-407f-8fee-4f445c5d4eda · outbound

This paper cites HunyuanVideo: A Systematic Framework For Large Video Generative Models.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:31.009125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:31.009125Z digest=sha256:7853b8b491a9a38aba79682ae87efe3c34af23d52742b87c2a9e68054c8c37e5

Observation a9f50fd2-8410-4887-94a1-fddec84a7c47 · outbound

This paper cites Subjective-Aligned Dataset and Metric for Text-to-Video Quality Assessment.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation Subjective-Aligned Dataset and Metric for Text-to-Video Quality Assessment

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:31.052527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:31.052527Z digest=sha256:37e1f22829404ab6cc714d48556207ec4473a5027a601aaf8def0e54d6345864

Observation 32bfaf9c-8b5c-4e53-b2cc-80377defbc32 · outbound

This paper cites an unresolved cited work.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation Unresolved cited work

Reference 43

Resolution
parse uncertain
no resolver link, observed 2026-08-07T13:59:31.099118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:31.099118Z digest=sha256:2a9a69ffdd89070be00cca88618657ee69668bd920f944e39fcbabde16961f71

Observation 75367925-d66c-4aca-915c-9a298cb79360 · outbound

This paper cites Pika-2.0 lab discord server.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation Pika-2.0 lab discord server

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:31.162597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:31.162597Z digest=sha256:2b6acb0cfdd8b8a9680447aad28ee459aedd84aea78d52448772ed52f4b8a718

Observation 6233e01f-9820-43fc-bfbf-9768baece134 · outbound

This paper cites OpenHumanVid: A Large-Scale High-Quality Dataset for Enhancing Human-Centric Video Generation.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation OpenHumanVid: A Large-Scale High-Quality Dataset for Enhancing Human-Centric Video Generation

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:31.215680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:31.215680Z digest=sha256:8383704bff0cba2a773ac4954060e85fe2b89794311356b7b9ff446f9bc2b33c

Observation c8fa7d41-df9b-48aa-b5ff-6f05e963ecee · outbound

This paper cites Improving Synthetic Image Detection Towards Generalization: An Image Transformation Perspective.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation Improving Synthetic Image Detection Towards Generalization: An Image Transformation Perspective

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:31.251492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:31.251492Z digest=sha256:4e29bd84fd699732b8868e81d70f16496cfdb3f3cdd5929b37b105b819fb3f30

Observation a965af12-e251-4d03-a94f-103a3ec054b3 · outbound

This paper cites Photomaker: Customizing realistic human photos via stacked id embedding.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation Photomaker: Customizing realistic human photos via stacked id embedding

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:31.325770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:31.325770Z digest=sha256:206c67cb12c02fc4adb70efedcd6a5e8dccfd2ed46a7ff12ac45ccac453f2323

Observation f70ac60e-5534-4725-a8dd-f8827e1945cd · outbound

This paper cites WF-VAE: Enhancing Video VAE by Wavelet-Driven Energy Flow for Latent Video Diffusion Model.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation WF-VAE: Enhancing Video VAE by Wavelet-Driven Energy Flow for Latent Video Diffusion Model

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:31.446588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:31.446588Z digest=sha256:5aec08393909e9d467b959c46ef001c66ea2eaebfd4b6828ecf26993d09722bb

Observation cd2fdf07-8d2a-4136-9e97-f530416fa399 · outbound

This paper cites Movie Weaver: Tuning-Free Multi-Concept Video Personalization with Anchored Prompts.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation Movie Weaver: Tuning-Free Multi-Concept Video Personalization with Anchored Prompts

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:31.555321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:31.555321Z digest=sha256:2deab8ff3ee25ce267b8474b07aa6907e33e14ae84a41ac669f43c681d2d5135

Observation afbf14cf-615b-453b-a715-eba9e447d2ff · outbound

This paper cites Open-Sora Plan: Open-Source Large Video Generation Model.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation Open-Sora Plan: Open-Source Large Video Generation Model

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:31.643439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:31.643439Z digest=sha256:535b594b32ab746cc903a19300fe27d9a8ec190a068226d0c578ce55bfd94728

Observation 646f6a09-7bce-4ea8-92fc-7faecb7daab8 · outbound

This paper cites Video-llava: Learning united visual representation by alignment before projection.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation Video-llava: Learning united visual representation by alignment before projection

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:31.756572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:31.756572Z digest=sha256:6c0275887e30bb2b78a7d227a3d5051ca5d1f4daf8ec17b43586421445862267

Observation 35763072-1ad3-4e7f-b05d-d4b2202915bb · outbound

This paper cites Stiv: Scalable text and image condi- tioned video generation.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation Stiv: Scalable text and image condi- tioned video generation

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:31.827870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:31.827870Z digest=sha256:e0a1437e5a66b8addf59afcb86a447dc31cce6ea7e5777137bbf6a798c1d7147

Observation af6623f3-3f69-48f3-a4b0-9ef802ea327f · outbound

This paper cites DeepSeek-V3 Technical Report.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation DeepSeek-V3 Technical Report

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:31.890570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:31.890570Z digest=sha256:939d79c008945071edc8cc4c4bfb8cb27aaad3b5126db88f67fb29eaf22c126c

Observation 0f02fd0d-b56c-4bce-8c2d-c8eaff2fd744 · outbound

This paper cites Lumina-Video: Efficient and Flexible Video Generation with Multi-scale Next-DiT.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation Lumina-Video: Efficient and Flexible Video Generation with Multi-scale Next-DiT

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:31.958999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:31.958999Z digest=sha256:0d42dc4b5c2582d009937fbcc6a9ea943f7d69403099d16ed7b9eb2c4adf111f

Observation 8b4584ab-6c87-4bf6-a062-9600001f2d39 · outbound

This paper cites Phantom: Subject-consistent video generation via cross-modal alignment.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation Phantom: Subject-consistent video generation via cross-modal alignment

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:32.071589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:32.071589Z digest=sha256:7adc94d84e549a309b2dfbd3305fb4ee8925cad6c63728ee2a2a225791b8f655

Observation 22e8865e-7ae2-4d68-8efe-989ccfa59e3b · outbound

This paper cites Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:32.143783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:32.143783Z digest=sha256:6b8b2714663ad115cdf1e14b4fd8113044a686a1c8342a6084dec0ad51bb3192

Observation 13596194-ebe1-4953-8804-bda4c428597d · outbound

This paper cites Evalcrafter: Benchmarking and evaluating large video generation models.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation Evalcrafter: Benchmarking and evaluating large video generation models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:32.195883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:32.195883Z digest=sha256:35f562cc5f59d2b559f0072e248954f61392a0183fac72d6b8e152f6bfac44a5

Observation f69d7047-7c83-44c7-b1eb-3f415ea0b1a7 · outbound

This paper cites Fetv: A benchmark for fine-grained evaluation of open-domain text-to-video genera- tion.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation Fetv: A benchmark for fine-grained evaluation of open-domain text-to-video genera- tion

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:32.278537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:32.278537Z digest=sha256:d9f58e1b2d2d0f4eb24b212865f483d42ea632a148f58de708a19e75e27f6294

Observation e992ba3e-adde-423e-a52a-b972d2a2508f · outbound

This paper cites Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:32.327532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:32.327532Z digest=sha256:379599914a6b125966d83acf88d071485ee0a6a3a64d72b4a83c518cf1f84bf2

Observation 985742aa-2d65-40e4-9bbc-f7c3a5bba31b · outbound

This paper cites Latte: Latent Diffusion Transformer for Video Generation.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation Latte: Latent Diffusion Transformer for Video Generation

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:32.388980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:32.388980Z digest=sha256:9c691c688e21ed4da16d09b458717acc01ed19c1e232af2be97351f157995592

Observation 79dc3038-4365-4037-a938-17f2ca0da233 · outbound

This paper cites Model Reveals What to Cache: Profiling-Based Feature Reuse for Video Diffusion Models.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation Model Reveals What to Cache: Profiling-Based Feature Reuse for Video Diffusion Models

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:32.433119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:32.433119Z digest=sha256:ec9a9d1e22b586b08fa40800c376f1eda404c4586eaadcd3549895fcb3d1a373

Observation ddf5a28d-2ff5-4c6f-b57f-cd36ecb9604c · outbound

This paper cites Follow your pose: Pose-guided text-to-video generation using pose-free videos.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation Follow your pose: Pose-guided text-to-video generation using pose-free videos

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:32.518451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:32.518451Z digest=sha256:6bde66d7511ff60dd33756a183082efbc5cfca2a09a874a41130e6496dfed548

Observation 59b261e3-a9e1-42ac-98d1-9da35f05c4fe · outbound

This paper cites Follow-Your-Click: Open-domain Regional Image Animation via Short Prompts.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation Follow-Your-Click: Open-domain Regional Image Animation via Short Prompts

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:32.625729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:32.625729Z digest=sha256:07721e6f4878345967906e2a281824c1a03686431ebd89149ed751ce6e00fbda

Observation 085f1cfa-a5dc-4b84-9ce5-37a9cab5b220 · outbound

This paper cites Follow-Your-Emoji: Fine-Controllable and Expressive Freestyle Portrait Animation.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation Follow-Your-Emoji: Fine-Controllable and Expressive Freestyle Portrait Animation

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:32.721019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:32.721019Z digest=sha256:ea6bf5be171776883ba8fba13b6198e3b388d1e9a0b5b5e45d2578f4d468c7dd

Observation c40ac82b-7a53-4239-94a5-0a2437246661 · outbound

This paper cites Magic-Me: Identity-Specific Video Customized Diffusion.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation Magic-Me: Identity-Specific Video Customized Diffusion

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:32.790081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:32.790081Z digest=sha256:5a60e11f379cf5e9a50ebc4903fd2cd63d49ad7cb9587f5d7de239b66895e70f

Observation ac1eba25-0a75-4758-af1e-bdaaab115766 · outbound

This paper cites Multi-task image classifier.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation Multi-task image classifier

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:32.885473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:32.885473Z digest=sha256:2645fb82d3f84e8319df2dd89a524c26e306f639410b705d9aca7bed9ab85126

Observation 37a1e504-5fc6-457a-8d8c-cc8d9be935db · outbound

This paper cites OpenVid-1M: A Large-Scale High-Quality Dataset for Text-to-video Generation.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation OpenVid-1M: A Large-Scale High-Quality Dataset for Text-to-video Generation

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:32.952192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:32.952192Z digest=sha256:39da0fb7d207461d8ef33aaf40b68fd8bede5124848a68731cf3727337a21f9d

Observation 0cedfbba-bee5-45cd-9865-a6bc31c75e8a · outbound

This paper cites Nyuad ai generated images detector.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation Nyuad ai generated images detector

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:33.006913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:33.006913Z digest=sha256:161cacd4622ac5656deca51f4eb34f244d7c188eca76b01d6afd7c44455fc5d7

Observation 1ff826d8-9563-40d7-af7f-753049c77cbf · outbound

This paper cites DreamDance: Animating Human Images by Enriching 3D Geometry Cues from 2D Poses.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation DreamDance: Animating Human Images by Enriching 3D Geometry Cues from 2D Poses

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:33.042872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:33.042872Z digest=sha256:48b3bf11e0fe1876b709d41f205e8a541f84a92852016f531866578c3fd7e829

Observation 9f51e9bb-64fb-4fe8-bbe2-17849e2b1c54 · outbound

This paper cites Open-Sora 2.0: Training a Commercial-Level Video Generation Model in $200k.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation Open-Sora 2.0: Training a Commercial-Level Video Generation Model in $200k

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:33.158196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:33.158196Z digest=sha256:b93dde0f550641077e94b0c8111be623dd3217138c54b50a6aa64d4eaed3744f

Observation 629342a2-e5a4-4261-bff7-42711df93a9c · outbound

This paper cites DreamBench++: A Human-Aligned Benchmark for Personalized Image Generation.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation DreamBench++: A Human-Aligned Benchmark for Personalized Image Generation

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:33.247451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:33.247451Z digest=sha256:b899ca92b7179a9d061336d01d64fea67757abd754922f304091e3a5d4348124

Observation 4ed2f581-5ca5-4541-af1a-e79b4c95c62f · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:33.328258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:33.328258Z digest=sha256:8352e62f3f6c13963c2c5e47850c45b7148d02cadef91b68fb7f52a5e49e6bfa

Observation 39695eb0-95dd-43df-a774-b98b2a9a1723 · outbound

This paper cites Learning transferable visual models from natural language supervision.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation Learning transferable visual models from natural language supervision

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:33.385897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:33.385897Z digest=sha256:dc1c07603cb41ccf61b2c39cf51972f9ad808c79c93fcdc34c592257ce6e4a87

Observation 80e7fd13-6dc1-40d0-8d2f-1a716c9df315 · outbound

This paper cites Hierarchical Text-Conditional Image Generation with CLIP Latents.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation Hierarchical Text-Conditional Image Generation with CLIP Latents

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:33.449820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:33.449820Z digest=sha256:c1c13b4e0abeeab62106dc088d094dd5e5bbd08bf3180843e5cf4b6f857ab942

Observation 33f3b505-4ad7-4a06-b3fd-bcf031b21a24 · outbound

This paper cites Zero-shot text-to-image generation.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation Zero-shot text-to-image generation

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:33.541902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:33.541902Z digest=sha256:bdf4628098785b46df3be0c01e6650825f233a66e5a3603b134ef66f48a0e02e

Observation 1281ad9c-c073-4665-95da-e2ca2f3f52e2 · outbound

This paper cites SAM 2: Segment Anything in Images and Videos.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation SAM 2: Segment Anything in Images and Videos

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:33.625427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:33.625427Z digest=sha256:f6fc221e9af884b4f2080195223ed2dcfa8780bafd777ae37550b0928d6fce4c

Observation fc7c93b7-3ee7-48ad-8d46-69a0dcb28789 · outbound

This paper cites Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:33.744844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:33.744844Z digest=sha256:79764bb97289e938aa9dcd2b746286efb131e7944798657e59be15311ddc8fcc

Observation a2501434-08d2-465e-b527-0a1c29caccfe · outbound

This paper cites Magi-1: Autoregressive video generation at scale, 2025.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation Magi-1: Autoregressive video generation at scale, 2025

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:33.801156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:33.801156Z digest=sha256:91fbefa82bc837608dd5dfaf2ba9e76397b359ebc1c520db3b747b023da4b9e5

Observation 88a9b364-495f-4981-a93e-fb0067626203 · outbound

This paper cites Seaweed-7B: Cost-Effective Training of Video Generation Foundation Model.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation Seaweed-7B: Cost-Effective Training of Video Generation Foundation Model

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:33.869824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:33.869824Z digest=sha256:3ee13457105ef51ec85382f8399c165833898dc96f4352dc0142086018faa068

Observation a22d91a2-71e3-43d3-a74f-9fde8ace6b8c · outbound

This paper cites Dragdiffusion: Harnessing diffusion models for interactive point-based image editing.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation Dragdiffusion: Harnessing diffusion models for interactive point-based image editing

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:33.939322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:33.939322Z digest=sha256:458b5364cc11757e818495c0fb118187c338f988105ce5b0a7236de655cc5338

Observation c6fdb347-8410-49d7-a3f4-87eebfbe09c0 · outbound

This paper cites Make-A-Video: Text-to-Video Generation without Text-Video Data.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation Make-A-Video: Text-to-Video Generation without Text-Video Data

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:34.017512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:34.017512Z digest=sha256:d5e03a4b7c898732a1e5d478bf5651c33867340ef50e59d1afe97ffa57032db1

Observation d490f859-6ccc-4e52-afd1-bf1382323b92 · outbound

This paper cites UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:34.089977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:34.089977Z digest=sha256:bd7237f58b68dc94d4cc91a5f8027b726f99e097c9c1921aa6688245b99f3085

Observation b35739bc-9d62-4f04-bc2a-0f1c5a8ea958 · outbound

This paper cites Animate-X: Universal Character Image Animation with Enhanced Motion Representation.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation Animate-X: Universal Character Image Animation with Enhanced Motion Representation

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:34.211475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:34.211475Z digest=sha256:600639330aaf2a8e72ab25b2919654dfbc816a24f01673e2356c4a233c24ac88

Observation cc4fa633-46fd-49cd-8158-098ff232e83d · outbound

This paper cites Cycle3D: High-quality and Consistent Image-to-3D Generation via Generation-Reconstruction Cycle.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation Cycle3D: High-quality and Consistent Image-to-3D Generation via Generation-Reconstruction Cycle

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:34.324476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:34.324476Z digest=sha256:93bdcf664e182d94085becb401cdaf2bbdb1c0bdeb070453826bde3d97261125

Observation 71052253-bbdd-4a14-ae89-66771b695258 · outbound

This paper cites Gemma 3 Technical Report.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation Gemma 3 Technical Report

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:34.446500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:34.446500Z digest=sha256:3a05f110912b9a67445c98e55a4a5d600f6e2ca1d1d0c1b464a9000ce17e9468

Observation c79d273d-30ad-4186-bafe-ebb6ac539c9f · outbound

This paper cites an unresolved cited work.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation Unresolved cited work

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:34.550294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:34.550294Z digest=sha256:fe550f7f7f8493b146ccaf0fdd2381cc6304384338b28e90ba03f08e75914aba

Observation 3e55a3bf-075d-4375-a7ec-717765318ab1 · outbound

This paper cites an unresolved cited work.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation Unresolved cited work

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:34.646831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:34.646831Z digest=sha256:8140b3a18754b3e51afd4d301e8daa95524015f3daf7243f1cc4ad08b249eb0f

Observation 6f59a7a3-5f8d-48b7-b9e1-e0984804100c · outbound

This paper cites Paddleocr.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation Paddleocr

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:34.766190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:34.766190Z digest=sha256:3a362f6686c0ddb2281245ee9f3bac4999d1e44fe88e53c00a948c5551fc852d

Observation fa802251-b8ba-4cdf-a54d-ce8af139e7e7 · outbound

This paper cites Wan: Open and Advanced Large-Scale Video Generative Models.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation Wan: Open and Advanced Large-Scale Video Generative Models

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:34.883459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:34.883459Z digest=sha256:2ae38ac611f1cabf3ea9b74dc133db2b46a6c68e8b4a80cad282b87b87e48907

Observation 6e3c765a-189b-4dfd-b82d-eaca53f409bd · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:34.953854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:34.953854Z digest=sha256:ff3bc952c2581567f53a794618502067de690663318d2325c8154a2cc1164ced

Observation 42a426d6-43e7-472f-a8c0-8647655e4834 · outbound

This paper cites Koala-36M: A Large-scale Video Dataset Improving Consistency between Fine-grained Conditions and Video Content.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation Koala-36M: A Large-scale Video Dataset Improving Consistency between Fine-grained Conditions and Video Content

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:35.043662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:35.043662Z digest=sha256:a0b435f4ea70d5644a6a13ee8d6a4a53858fd615ad331f395faa648e39f174bb

Observation 956c6d0d-814c-4d06-b946-94e4a5fb6d59 · outbound

This paper cites InstantID: Zero-shot Identity-Preserving Generation in Seconds.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation InstantID: Zero-shot Identity-Preserving Generation in Seconds

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:35.115504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:35.115504Z digest=sha256:b04fcd87f808726953f0a691d8fa2aaf08a1894ef8fb118af8afdbaa100e452d

Observation 4fd5de33-cbb1-4744-9e07-f6a94bc67674 · outbound

This paper cites VidProM: A Million-scale Real Prompt-Gallery Dataset for Text-to-Video Diffusion Models.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation VidProM: A Million-scale Real Prompt-Gallery Dataset for Text-to-Video Diffusion Models

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:35.179016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:35.179016Z digest=sha256:3ab401ea1da459e9c47f660b95ac4e713aca1ae6c40e4bf7ae7907c28413ac03

Observation c5788791-343a-4a42-b610-91b72e515481 · outbound

This paper cites Swap Attention in Spatiotemporal Diffusions for Text-to-Video Generation.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation Swap Attention in Spatiotemporal Diffusions for Text-to-Video Generation

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:35.266101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:35.266101Z digest=sha256:b3a123694456967e1e4edaf1d761fd315567f52bd5f3b7272f22b60df0fc236e

Observation 8daffc84-86b1-4a83-8ec4-10a2186913ec · outbound

This paper cites InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:35.298215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:35.298215Z digest=sha256:d57f690d828f3f8bee5aa754014ce45090e8f2c610bc4ca38bc36f60ef18761b

Observation 27bd54d9-7a9b-4961-a550-0366d6b1b154 · outbound

This paper cites Is Your World Simulator a Good Story Presenter? A Consecutive Events-Based Benchmark for Future Long Video Generation.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation Is Your World Simulator a Good Story Presenter? A Consecutive Events-Based Benchmark for Future Long Video Generation

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:35.359517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:35.359517Z digest=sha256:0847667a30491b1c5aa502fc7ef1169252708f7d693549e5ab5d0ab09ab5f383

Observation 2e7d5a2d-e248-40f9-abe4-65937c05c91b · outbound

This paper cites EchoVideo: Identity-Preserving Human Video Generation by Multimodal Feature Fusion.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation EchoVideo: Identity-Preserving Human Video Generation by Multimodal Feature Fusion

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:35.395109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:35.395109Z digest=sha256:1a5a41f2f61c8a436be5e99b3545944bd953dbe3076e41a0f7f9dcef5d40a980

Observation 7119b8e0-fdfa-4aae-aab3-d8766f96a204 · outbound

This paper cites Dreamvideo: Composing your dream videos with customized subject and motion.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation Dreamvideo: Composing your dream videos with customized subject and motion

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:35.430457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:35.430457Z digest=sha256:f421a442479072a0d43720c04cee0e3a246ab8bc93d8dad6682ab447188e9f46

Observation e1b967fb-9f0c-44cd-b544-e8908a293dc7 · outbound

This paper cites Exploring video quality assessment on user generated contents from aesthetic and technical perspectives.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation Exploring video quality assessment on user generated contents from aesthetic and technical perspectives

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:35.489028Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:35.489028Z digest=sha256:3ad5ec4c3f45e024a2e45ea65f8469409645621bc3a9272f5941f9a59e4635e1

Observation c48c9115-7c42-4700-92cb-3b8a15840695 · outbound

This paper cites Towards A Better Metric for Text-to-Video Generation.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation Towards A Better Metric for Text-to-Video Generation

Reference 101

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:35.615058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:35.615058Z digest=sha256:e50952cd435bb5bf9a6fc4320447f7d77c8595d81da4884d777d535b1ba5450d

Observation 0b5ae721-ba87-40ff-bab8-71e9c71c1738 · outbound

This paper cites MotionBooth: Motion-Aware Customized Text-to-Video Generation.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation MotionBooth: Motion-Aware Customized Text-to-Video Generation

Reference 102

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:35.677086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:35.677086Z digest=sha256:7a2e0096e54afe80dd3cea45829280df849b4a182cbe5b43577a01baa4d51de3

Pith citing papers

Observation 3aca3f54-c5c7-42cd-9089-7b36a2819a38 · inbound

Learning Zero-Shot Subject-Driven Video Generation Using 1% Compute cites this paper.

Learning Zero-Shot Subject-Driven Video Generation Using 1% Compute OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-05-22T17:51:54.359831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-22T17:51:41.947939Z digest=sha256:71046a6ebf451c0f9ae0897947efa3bf8253af110a74bda91b0ff862035d58b7

Observation c203b860-2132-44a4-b42f-76f32dfc4acd · inbound

UniWorld-V1: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation cites this paper.

UniWorld-V1: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-12T17:34:27.064326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T17:34:26.951644Z digest=sha256:84d0cfe7ad4c0d9eae2a9c3cc0faae47f68d5d84a2d9fa651cf0957e891e8550

Observation 1f147232-5df2-4fc5-99ef-ccfb5336fef7 · inbound

Identity-Preserving Text-to-Video Generation via Training-Free Prompt, Image, and Guidance Enhancement cites this paper.

Identity-Preserving Text-to-Video Generation via Training-Free Prompt, Image, and Guidance Enhancement OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-05T12:42:42.185508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:42:42.185508Z digest=sha256:c112869cb90a226e94aace6294d4fc130e3697183f0bee2f7ad4b62661107a7b

Observation df606aab-91b1-478b-acf8-815f5d58b22b · inbound

Rethinking Position Embedding as a Context Controller for Multi-Reference and Multi-Shot Video Generation cites this paper.

Rethinking Position Embedding as a Context Controller for Multi-Reference and Multi-Shot Video Generation OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation

Reference 39

Resolution
unresolved
no resolver link, observed 2026-07-13T12:23:05.876881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T12:23:05.876881Z digest=sha256:7f2b015e281fb8e3991e5eb56e1c9a41ae17914745a131698cf317a9b692b4c5

Observation 47eeac02-5dbc-4b30-b9c4-37208fa942b8 · inbound

OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation cites this paper.

OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation

Reference 70

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:11:00.839751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T15:09:02.727887Z digest=sha256:7843d614a9d9876373580fdc6bf4ab2714cea1ad117d5bc7533bf70fb6e11708

Observation ca89bad4-82c4-42aa-8181-608e8e4bff9b · inbound

LIVE: Leveraging Image Manipulation Priors for Instruction-based Video Editing cites this paper.

LIVE: Leveraging Image Manipulation Priors for Instruction-based Video Editing OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation

Reference 55

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T07:21:55.095118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T07:21:41.483427Z digest=sha256:6ea988fd00223e4b106ed627d56f1782a42a5f93784c08df0801debd0a09faa5

Observation 146ce315-82b4-4402-80e4-5c6faf90d9aa · inbound

MuSS: A Large-Scale Dataset and Cinematic Narrative Benchmark for Multi-Shot Subject-to-Video Generation cites this paper.

MuSS: A Large-Scale Dataset and Cinematic Narrative Benchmark for Multi-Shot Subject-to-Video Generation OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:11:19.210744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T06:28:42.129881Z digest=sha256:ed712d5441e677b1c9780e777ed20f840e4ca1d250d9aaa4c16f4d48c93db2c4

Observation 3c22f376-4c71-479b-ae69-5a2cacaa96a2 · inbound

MuSS: A Large-Scale Dataset and Cinematic Narrative Benchmark for Multi-Shot Subject-to-Video Generation cites this paper.

MuSS: A Large-Scale Dataset and Cinematic Narrative Benchmark for Multi-Shot Subject-to-Video Generation OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-12T00:51:15.130701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T00:50:10.509727Z digest=sha256:00a20f6fbfd62eb1f277427db21ff9de679a084c5224fbb250089d03a9798091

Observation 3b2377c7-e313-49fd-a656-7d491fbca3fb · inbound

EntityBench: Towards Entity-Consistent Long-Range Multi-Shot Video Generation cites this paper.

EntityBench: Towards Entity-Consistent Long-Range Multi-Shot Video Generation OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-15T03:19:44.040337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T03:16:03.450742Z digest=sha256:bcd4f631c74773bf1d6f443f9830d27717379b1e4d2e94aeb662a737483d0c60

Observation 6884ece6-c3fc-4215-9ccc-eaebec221846 · inbound

FashionChameleon: Towards Real-Time and Interactive Human-Garment Video Customization cites this paper.

FashionChameleon: Towards Real-Time and Interactive Human-Garment Video Customization OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-20T20:13:43.506965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-20T20:11:33.277349Z digest=sha256:afa10342ca1ac4b419d695b3ea67a6fed894bc28c30f7d772a9a6c34d1e34684

Observation d0708f02-1251-4c41-98d5-c8e04e91a0cc · inbound

FashionChameleon: Towards Real-Time and Interactive Human-Garment Video Customization cites this paper.

FashionChameleon: Towards Real-Time and Interactive Human-Garment Video Customization OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-06-30T19:25:00.604747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T19:23:07.044659Z digest=sha256:3bd70d01e6f4dfc463a199fd31f777bbbe6d20baaa5510290be916f1e62cd227

Observation 3f51a193-c34a-4d5e-9ebd-0abde07ca1f5 · inbound

Omni-Customizer: End-to-End MultiModal Customization for Joint Audio-Video Generation cites this paper.

Omni-Customizer: End-to-End MultiModal Customization for Joint Audio-Video Generation OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-20T14:13:21.404522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-20T14:08:30.802619Z digest=sha256:4640f8f8ca16bab053538b7cbf0e91db14dbf49fb2d07a93f087c8d2e3e2e9e2

Observation ec6a5880-b1f9-4524-8903-030cab407046 · inbound

MSAVBench: Towards Comprehensive and Reliable Evaluation of Multi-Shot Audio-Video Generation cites this paper.

MSAVBench: Towards Comprehensive and Reliable Evaluation of Multi-Shot Audio-Video Generation OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation

Reference 74

Resolution
verified exact
arxiv_id, observed 2026-05-20T05:18:03.111547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-20T05:17:11.690484Z digest=sha256:d332302dd235f057657669b153b230cadf9fc20f04488a8c0a53c05c4b7585cc

Observation 5717e354-4435-4328-84c4-ea7fe842b65e · inbound

MSAVBench: Towards Comprehensive and Reliable Evaluation of Multi-Shot Audio-Video Generation cites this paper.

MSAVBench: Towards Comprehensive and Reliable Evaluation of Multi-Shot Audio-Video Generation OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation

Reference 74

Resolution
verified exact
arxiv_id, observed 2026-06-30T18:04:58.012646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T18:00:39.556424Z digest=sha256:aeb0371533a5a951b6041411de7635da9f2402a37f15a0775193aaf871fb9570

Observation cd359f9b-2198-495b-99ac-3411e362971e · inbound

Bernini: Latent Semantic Planning for Video Diffusion cites this paper.

Bernini: Latent Semantic Planning for Video Diffusion OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation

Reference 91

Resolution
verified exact
arxiv_id, observed 2026-05-22T06:41:10.560161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-22T06:39:47.124605Z digest=sha256:a9fdb29e66a3c639f3d36ec696b19c66657aa3cb74210ad028edf78f972d2cab

Observation 4b4fd2de-3168-4ea7-a51a-7561b92f26d8 · inbound

Tail-Aware HiFloat4: W4A4 Post-Training Quantization for Wan2.2 cites this paper.

Tail-Aware HiFloat4: W4A4 Post-Training Quantization for Wan2.2 OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-06-29T18:23:50.628549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T18:20:47.823852Z digest=sha256:5ce5d933e19acf0d7fda5067a632dcfeb5356347a7c9026e8e013421b10d8352

Observation e4cabb42-3932-46e4-9b28-ebaf7eacccd8 · inbound

Collaborative Few-Step Distillation and Low-Bit Quantization for Wan2.2 Dual-Expert Video Diffusion Models cites this paper.

Collaborative Few-Step Distillation and Low-Bit Quantization for Wan2.2 Dual-Expert Video Diffusion Models OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-06-28T19:02:34.320535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T18:58:13.799615Z digest=sha256:d3113efd89ea9673f411a30ec9227d2208ae1f0c26ead4e561b0672b064dc78e

Observation a615470e-d2b6-4529-8694-59cd4b3071a9 · inbound

HarmoView: Harmonizing Multi-View Constraints for Identity-Consistent Video Generation cites this paper.

HarmoView: Harmonizing Multi-View Constraints for Identity-Consistent Video Generation OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T04:37:36.638712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T13:49:50.272650Z digest=sha256:dce3f1ad8cde35d553e81a26d40884ec94e70b2eb71df3bc30067c49690c1541

Observation 90ac3907-af02-4025-9542-4cd1e31b21b3 · inbound

DomainShuttle: Freeform Open Domain Subject-driven Text-to-video Generation cites this paper.

DomainShuttle: Freeform Open Domain Subject-driven Text-to-video Generation OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-07-04T21:10:09.696509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-25T19:00:23.260939Z digest=sha256:5fd156b11e2ede2d279ccb080b3bf3b70d2c57bab5fe6f54b23a81316646469d

Observation 415ecded-5571-409f-9c9f-beb34b30e50a · inbound

Aura: Consistent Multi-Subject Video Generation via VLM-Grounded Semantic Alignment cites this paper.

Aura: Consistent Multi-Subject Video Generation via VLM-Grounded Semantic Alignment OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-07-11T20:11:31.576642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T20:11:31.576642Z digest=sha256:21af8e7fa25f431c602b4f1163db492f00bdd696edfe4aba40204b1f2a5a5c89

Observation 276855ba-c617-4634-a6d8-21ed7da399bb · inbound

Vera: Identity-Faithful Human Subject-to-Video Generation cites this paper.

Vera: Identity-Faithful Human Subject-to-Video Generation OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T10:24:52.920746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:24:52.920746Z digest=sha256:bd3c81b55a32181c38de3af2298378ffc7526bc9767293f0c80b074ba5f2f1b5

Observation 1fdba910-7052-43bc-a0ce-292612471ffe · inbound

RefCaptioner: Multi-Reference Image-Grounded Video Captioning cites this paper.

RefCaptioner: Multi-Reference Image-Grounded Video Captioning OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation

Reference 37

Resolution
unresolved
no resolver link, observed 2026-07-31T05:08:20.144127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T05:08:20.144127Z digest=sha256:c927c918c1f7134d9107a0dbf0b183b06dcd76fab602410f2a14baaed88416a4