Pith. sign in

Paper Citation Record · LEDGER

Phantom: Subject-consistent video generation via cross-modal alignment

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 38 inbound Pith citation observations for arXiv:2502.11079.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.11079 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 38 of 38 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 38 of 38 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:59:32.071589Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T03:19:30.378145Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation e8ca3bea-1703-4a77-b703-ae6b0d629a7b · inbound

VACE: All-in-One Video Creation and Editing cites this paper.

VACE: All-in-One Video Creation and Editing Phantom: Subject-consistent video generation via cross-modal alignment

Reference 35

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T00:53:54.003853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T00:53:53.855965Z digest=sha256:1cbccfbfb5244f662a1f98bda77da0ad62e0c76ece7fc8e7a9e4f3951dd12e2c

Observation dab58e75-016b-4c38-b61a-e1ac198ec9ce · inbound

Learning Zero-Shot Subject-Driven Video Generation Using 1% Compute cites this paper.

Learning Zero-Shot Subject-Driven Video Generation Using 1% Compute Phantom: Subject-consistent video generation via cross-modal alignment

Reference 33

Resolution
metadata mismatch
arxiv_id, observed 2026-05-22T17:51:54.313500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-22T17:51:41.947939Z digest=sha256:22032a2108c47164a8194e04ebc45f937b512954d7795705af937679dade4f69

Observation 8b4584ab-6c87-4bf6-a062-9600001f2d39 · inbound

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation cites this paper.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation Phantom: Subject-consistent video generation via cross-modal alignment

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:32.071589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:32.071589Z digest=sha256:c91e0af18095b26a97ab86cf32d833627270515562fa9effe32074a7a82b3a8a

Observation 820661d1-558e-4c66-b64a-1d735b45ff39 · inbound

AnimeShooter: A Multi-Shot Animation Dataset for Reference-Guided Video Generation cites this paper.

AnimeShooter: A Multi-Shot Animation Dataset for Reference-Guided Video Generation Phantom: Subject-consistent video generation via cross-modal alignment

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T11:12:10.052472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:12:10.052472Z digest=sha256:aea4bbf7de48d6993dc7f896d448632c434f2cdcf1963aea41f3989b8dfb9a6d

Observation e62295d4-8fa2-4396-8deb-fac1b7d52bf6 · inbound

PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and Enhancement cites this paper.

PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and Enhancement Phantom: Subject-consistent video generation via cross-modal alignment

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T05:30:32.974027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:30:32.974027Z digest=sha256:df40bb4214174f1fe3d8b745999a3da90acdf05b7bedfc499586401e860d88db

Observation 762ce83a-a2b1-4b2b-b0d3-7452fde4b615 · inbound

DreamActor-H1: High-Fidelity Human-Product Demonstration Video Generation via Motion-designed Diffusion Transformers cites this paper.

DreamActor-H1: High-Fidelity Human-Product Demonstration Video Generation via Motion-designed Diffusion Transformers Phantom: Subject-consistent video generation via cross-modal alignment

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T04:27:39.046915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:27:39.046915Z digest=sha256:b37339f0865f73c5430b902da9295d319cdd7b4ca54432f0eb1460e90c28a1dc

Observation cde2f2fb-60aa-483a-8704-ee42bd252af6 · inbound

Phantom-Data : Towards a General Subject-Consistent Video Generation Dataset cites this paper.

Phantom-Data : Towards a General Subject-Consistent Video Generation Dataset Phantom: Subject-consistent video generation via cross-modal alignment

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:09.885970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:09.885970Z digest=sha256:a582698c1cecee9d4509931f54937c57d4372a7afd506de8159c190872270b41

Observation 78f0d332-7e44-4678-bd3e-d7f2777dd5a5 · inbound

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality cites this paper.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Phantom: Subject-consistent video generation via cross-modal alignment

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:27.649430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:51:27.649430Z digest=sha256:d18d143c36f33b39a4b92587ad4a13d7d9cb87ead15c21428459c14f705b96c1

Observation 5684ad4b-79c4-42a7-8685-928f47c503ff · inbound

FantasyPortrait: Enhancing Multi-Character Portrait Animation with Expression-Augmented Diffusion Transformers cites this paper.

FantasyPortrait: Enhancing Multi-Character Portrait Animation with Expression-Augmented Diffusion Transformers Phantom: Subject-consistent video generation via cross-modal alignment

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T16:39:36.070911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:39:36.070911Z digest=sha256:3ab85043af9073072cc10c8b67f156bb7df585d3ebbf3a657f8cef9583569c5a

Observation d07100f1-754d-40de-b195-94151d481fd2 · inbound

FantasyTalking2: Timestep-Layer Adaptive Preference Optimization for Audio-Driven Portrait Animation cites this paper.

FantasyTalking2: Timestep-Layer Adaptive Preference Optimization for Audio-Driven Portrait Animation Phantom: Subject-consistent video generation via cross-modal alignment

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-05T20:06:47.800330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:06:47.800330Z digest=sha256:a9c01c79c16e39b3a5da155bde7dd1e6805da30b9f83d1186dd58dedc42ccb56

Observation eeeba291-e2d0-47b9-8c95-6822177149b3 · inbound

Identity-Preserving Text-to-Video Generation via Training-Free Prompt, Image, and Guidance Enhancement cites this paper.

Identity-Preserving Text-to-Video Generation via Training-Free Prompt, Image, and Guidance Enhancement Phantom: Subject-consistent video generation via cross-modal alignment

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T12:42:40.992340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:42:40.992340Z digest=sha256:a54d2ea17d8261b0032b0fe62fc0bf99e5120e24cf04036f485983e8b1532af4

Observation 91ed0832-9ef6-4f2a-a878-5ce4e3355301 · inbound

UniVideo: Unified Understanding, Generation, and Editing for Videos cites this paper.

UniVideo: Unified Understanding, Generation, and Editing for Videos Phantom: Subject-consistent video generation via cross-modal alignment

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T10:49:48.589583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:49:48.589583Z digest=sha256:56b6170164e5735aa80c234c07b1855d71c20f3a8a4a55fb921c00a6106a85fc

Observation e5d041c1-8351-482f-9b6d-34fad7652cad · inbound

One-to-All Animation: Alignment-Free Character Animation and Image Pose Transfer cites this paper.

One-to-All Animation: Alignment-Free Character Animation and Image Pose Transfer Phantom: Subject-consistent video generation via cross-modal alignment

Reference 27

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T18:25:28.396016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-21T18:25:22.486891Z digest=sha256:de199f3b13c75a6bc6edaafdb1519cb3939092bb009c034be545a164dcd5d9fe

Observation b1e13b4c-d732-427f-b594-34f56251ba95 · inbound

CustomX: Unified Character, Action, and Scene Customization in Video World Models cites this paper.

CustomX: Unified Character, Action, and Scene Customization in Video World Models Phantom: Subject-consistent video generation via cross-modal alignment

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-03T15:28:55.722892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T15:28:55.722892Z digest=sha256:07092ebb02e580d01358b1431b2cc69d569bc58703f7954b37c48cc50d217a7d

Observation 04a016b9-f2f9-4cfa-a631-94fba03f29ea · inbound

EchoTorrent: Towards Swift, Sustained, and Streaming Multi-Modal Video Generation cites this paper.

EchoTorrent: Towards Swift, Sustained, and Streaming Multi-Modal Video Generation Phantom: Subject-consistent video generation via cross-modal alignment

Reference 40

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T22:20:22.455970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T22:20:16.320171Z digest=sha256:51ffffa1629361351bcc2f43bc476e5389f5974b96fb162a71b32deff110c124

Observation b16c105f-f6af-4215-ad0b-6fb297f91cdc · inbound

MVHOI: Bridge Multi-view Condition to Complex Human-Object Interaction Video Reenactment via 3D Foundation Model cites this paper.

MVHOI: Bridge Multi-view Condition to Complex Human-Object Interaction Video Reenactment via 3D Foundation Model Phantom: Subject-consistent video generation via cross-modal alignment

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T05:51:18.232473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T05:51:18.232473Z digest=sha256:ec1c8ce547e358be10607893b074b3a71e7a45555de4f581ded2049faca95ecc

Observation 60704d15-2282-4390-b433-5a5fe85931b1 · inbound

Vanast: Virtual Try-On with Human Image Animation via Synthetic Triplet Supervision cites this paper.

Vanast: Virtual Try-On with Human Image Animation via Synthetic Triplet Supervision Phantom: Subject-consistent video generation via cross-modal alignment

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-10T21:55:49.210641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T20:33:25.422350Z digest=sha256:4e0ce27930cf40a3c0d290aa5064704daf3732f7f43abd2bb295fc01880f1ca6

Observation a5ba3bef-7a30-4b49-b8ba-dfc6e2dd5f94 · inbound

Evolution of Video Generative Foundations cites this paper.

Evolution of Video Generative Foundations Phantom: Subject-consistent video generation via cross-modal alignment

Reference 205

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T00:05:51.328734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T18:41:38.616611Z digest=sha256:9601abdc5835d518a367242efd27a7c37f17d9bd86d1bb9b3a10ebd8e9a60777

Observation d4177208-82fc-49ec-9bf1-bed27d5d2af9 · inbound

Prompt Relay: Inference-Time Temporal Control for Multi-Event Video Generation cites this paper.

Prompt Relay: Inference-Time Temporal Control for Multi-Event Video Generation Phantom: Subject-consistent video generation via cross-modal alignment

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T10:01:00.416886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T15:42:35.693519Z digest=sha256:8fab4dc95f7ec8a046e8786c9286c5a142d235b22ae974c8e26a0323cde85ffd

Observation c5df77d7-6e3f-49ed-8acf-5e3c5ddccbe4 · inbound

Script-a-Video: Deep Structured Audio-visual Captions via Factorized Streams and Relational Grounding cites this paper.

Script-a-Video: Deep Structured Audio-visual Captions via Factorized Streams and Relational Grounding Phantom: Subject-consistent video generation via cross-modal alignment

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T11:11:03.692220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T15:07:45.595260Z digest=sha256:4be369f332f7a010ca2d4aa5b22a14e09899b230e3c5f1d0024171a658b289d4

Observation 8990056c-c55a-4965-b135-a1b03ef9f68b · inbound

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey cites this paper.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Phantom: Subject-consistent video generation via cross-modal alignment

Reference 141

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T08:30:56.950219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T16:36:33.264166Z digest=sha256:2f9e4d2d4417da4328ea108cb90e1a0b19b734a9aadf9e7586c20ac9107eca0d

Observation a504db6f-6ef3-4462-bd37-eae9a89ad9b7 · inbound

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey cites this paper.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Phantom: Subject-consistent video generation via cross-modal alignment

Reference 120

Resolution
unresolved
no resolver link, observed 2026-07-12T22:04:31.302192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:04:31.302192Z digest=sha256:af663da6b541dee353074a5e5d80be38c10f758d920e5bf823904bc69aa47122

Observation 78cd8348-c81b-4c6c-8519-b0c3160d0004 · inbound

OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation cites this paper.

OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation Phantom: Subject-consistent video generation via cross-modal alignment

Reference 42

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T11:06:05.893388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T15:09:02.727887Z digest=sha256:d0ce8310401c76826395c1295755b600d63d27cedfa6eae9d9b687d22b0bb08e

Observation 8020a8c1-a8f7-4b1e-bd94-51dde555e8a1 · inbound

ReImagine: Rethinking Controllable High-Quality Human Video Generation via Image-First Synthesis cites this paper.

ReImagine: Rethinking Controllable High-Quality Human Video Generation via Image-First Synthesis Phantom: Subject-consistent video generation via cross-modal alignment

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:01:26.592322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T02:24:57.882447Z digest=sha256:de37d2ce272edf0fcd3df557e292054e465596984abea579b364d3c7f35d64d6

Observation 9c0c0c03-b5a1-479c-bb2b-3fd35485b6c0 · inbound

FaithfulFaces: Pose-Faithful Facial Identity Preservation for Text-to-Video Generation cites this paper.

FaithfulFaces: Pose-Faithful Facial Identity Preservation for Text-to-Video Generation Phantom: Subject-consistent video generation via cross-modal alignment

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T17:21:09.635204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-08T17:40:00.224358Z digest=sha256:c908970c2cf739ed91d7d987310adf4791983baf3815735c326ac3c7ccd0195e

Observation 871d318c-ba5c-438c-ae15-f110cd8a7bac · inbound

FashionChameleon: Towards Real-Time and Interactive Human-Garment Video Customization cites this paper.

FashionChameleon: Towards Real-Time and Interactive Human-Garment Video Customization Phantom: Subject-consistent video generation via cross-modal alignment

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T20:13:43.610050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T20:11:33.277349Z digest=sha256:6beb27ec584598d6a88aee14342ed7e194ff73fbe8bd8898d009f5b2052ba7b4

Observation 66e35ba8-be81-4f43-982a-63b0a576b155 · inbound

FashionChameleon: Towards Real-Time and Interactive Human-Garment Video Customization cites this paper.

FashionChameleon: Towards Real-Time and Interactive Human-Garment Video Customization Phantom: Subject-consistent video generation via cross-modal alignment

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T19:25:00.562996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T19:23:07.044659Z digest=sha256:2e2590c20528a789adef9a3002eef766344a2072e8d15f8fab0ce1a20f060b7d

Observation b6a1476b-f6fc-4200-b0fc-087b41a9a896 · inbound

Bernini: Latent Semantic Planning for Video Diffusion cites this paper.

Bernini: Latent Semantic Planning for Video Diffusion Phantom: Subject-consistent video generation via cross-modal alignment

Reference 45

Resolution
metadata mismatch
arxiv_id, observed 2026-05-22T06:41:10.542258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-22T06:39:47.124605Z digest=sha256:9ce2700f29e94e1a0677806885e855194672b605593d09a5ef3b3fdcabc74f21

Observation 048d38ed-a3bc-4ad6-a5c0-aa4a4bded87a · inbound

EM-Vid: Training-Free Entity-Centric Memory for Efficient and Consistent Multi-Shot Video Generation cites this paper.

EM-Vid: Training-Free Entity-Centric Memory for Efficient and Consistent Multi-Shot Video Generation Phantom: Subject-consistent video generation via cross-modal alignment

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-25T05:00:22.425250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T04:57:13.479165Z digest=sha256:542946b9775a1595809d27870ff502e472758f0a9bc959158b2ef99c0b9e7c0b

Observation 54d0a592-7bba-47d0-b7f3-a32634311248 · inbound

MetaWorld: Scaling Multi-Agent Video World Model from Single-view Video Data cites this paper.

MetaWorld: Scaling Multi-Agent Video World Model from Single-view Video Data Phantom: Subject-consistent video generation via cross-modal alignment

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-07-01T22:56:20.244659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T14:52:30.406683Z digest=sha256:d8297c0b9fe66589c920adddbe4024e15a151caaac8f389cf4c7d5d34254a576

Observation 610c77d1-d687-469a-abaf-59b53dac25f1 · inbound

Streaming Video Generation with Streaming Force Control cites this paper.

Streaming Video Generation with Streaming Force Control Phantom: Subject-consistent video generation via cross-modal alignment

Reference 37

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T17:07:12.387385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T22:14:32.465663Z digest=sha256:c0f7a4c2bbe0f8549e644a7a098b369c3f4c8b2dbce06e6d5e71181fb7620223

Observation f48bf4c6-21ea-43b5-b88f-86cac9cca637 · inbound

HarmoView: Harmonizing Multi-View Constraints for Identity-Consistent Video Generation cites this paper.

HarmoView: Harmonizing Multi-View Constraints for Identity-Consistent Video Generation Phantom: Subject-consistent video generation via cross-modal alignment

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T04:37:36.648289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T13:49:50.272650Z digest=sha256:113288e4fda6be9dc8a6312bc59eac8ca90e55078e64a60bd9f2da4167fda0e8

Observation 24afdf48-3e48-4b2e-a930-603545850b29 · inbound

ARGUS: Stacked Multi-View Identity Mosaic Injection for Subject-Preserving Video Generation cites this paper.

ARGUS: Stacked Multi-View Identity Mosaic Injection for Subject-Preserving Video Generation Phantom: Subject-consistent video generation via cross-modal alignment

Reference 29

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T08:17:45.947052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T10:46:56.871174Z digest=sha256:e926d1fe25cdbd917af1d29be759db181276436a862e00b93da991307e51d077

Observation 68f86b7d-16e3-4865-a3fa-70fb1297d568 · inbound

GroundShot: Visually Consistent Multi-Shot Long Video Generation via Entity-Grounded Shot Scheduling cites this paper.

GroundShot: Visually Consistent Multi-Shot Long Video Generation via Entity-Grounded Shot Scheduling Phantom: Subject-consistent video generation via cross-modal alignment

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T03:19:30.381652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T18:15:51.862192Z digest=sha256:5f8baf54e651bf116dfdc4ca77e00cfb35d2328c1a3b9bd9f26fb0a5bb29b150

Observation 73bda7db-f3f1-424c-89bf-43f5dcb8f4a5 · inbound

GroundShot: Visually Consistent Multi-Shot Long Video Generation via Entity-Grounded Shot Scheduling cites this paper.

GroundShot: Visually Consistent Multi-Shot Long Video Generation via Entity-Grounded Shot Scheduling Phantom: Subject-consistent video generation via cross-modal alignment

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-02T10:48:05.537673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:48:05.537673Z digest=sha256:4122ba5ae4854cfb6bdec205e025eab096d670683da181c47d57a39e227ced54

Observation ecf247e4-33a2-4343-bb22-f847989a56a7 · inbound

A Good Talk Does not Look Like a Summary, It Teaches You! Measuring Takeaways from Paper-to-Video Talks cites this paper.

A Good Talk Does not Look Like a Summary, It Teaches You! Measuring Takeaways from Paper-to-Video Talks Phantom: Subject-consistent video generation via cross-modal alignment

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T15:45:48.892408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-30T00:59:50.478584Z digest=sha256:fb2938664649e83c2a2b9b6320777fee07876cf59665725cb5d89267bbbf8876

Observation 95c4ed9a-fded-45a9-8f3e-332d66048293 · inbound

Keyframe-Anchored Identity Preservation for Sequential-Action Video Generation cites this paper.

Keyframe-Anchored Identity Preservation for Sequential-Action Video Generation Phantom: Subject-consistent video generation via cross-modal alignment

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T16:30:02.172800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:30:02.172800Z digest=sha256:b3e90f5256ef528ac7d465055f9d55c86e68d6d39734e2551584d2a2cf926d77

Observation d111140e-0430-4dbc-a018-a4527dfbfc31 · inbound

ID-V2V: Identity-Preserving Video Restylization cites this paper.

ID-V2V: Identity-Preserving Video Restylization Phantom: Subject-consistent video generation via cross-modal alignment

Reference 191

Resolution
unresolved
no resolver link, observed 2026-08-01T04:27:31.566712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T04:27:31.566712Z digest=sha256:d15024483141070dca46eef13a6b1ec044bd7de70d1c1cb3533b1fdae24f766e