Pith. sign in

Paper Citation Record · LEDGER

Phantom: Subject-consistent video generation via cross-modal alignment

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 38 inbound Pith citation observations for arXiv:2502.11079.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.11079 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 38 of 38 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 38 of 38 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:59:32.071589Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T03:19:30.378145Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation e8ca3bea-1703-4a77-b703-ae6b0d629a7b · inbound

VACE: All-in-One Video Creation and Editing cites this paper.

VACE: All-in-One Video Creation and Editing Phantom: Subject-consistent video generation via cross-modal alignment

Reference 35

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T00:53:54.003853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T00:53:53.855965Z digest=sha256:361013a176224599218e0dfcf5449df6fac23957fc5f1092fd8c1da168f945fb

Observation dab58e75-016b-4c38-b61a-e1ac198ec9ce · inbound

Learning Zero-Shot Subject-Driven Video Generation Using 1% Compute cites this paper.

Learning Zero-Shot Subject-Driven Video Generation Using 1% Compute Phantom: Subject-consistent video generation via cross-modal alignment

Reference 33

Resolution
metadata mismatch
arxiv_id, observed 2026-05-22T17:51:54.313500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-22T17:51:41.947939Z digest=sha256:01d8f13ba0bde2d84a30654862c7e0fd09da0718bdb878e98227a85c5a496378

Observation 8b4584ab-6c87-4bf6-a062-9600001f2d39 · inbound

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation cites this paper.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation Phantom: Subject-consistent video generation via cross-modal alignment

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:32.071589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:32.071589Z digest=sha256:7adc94d84e549a309b2dfbd3305fb4ee8925cad6c63728ee2a2a225791b8f655

Observation 820661d1-558e-4c66-b64a-1d735b45ff39 · inbound

AnimeShooter: A Multi-Shot Animation Dataset for Reference-Guided Video Generation cites this paper.

AnimeShooter: A Multi-Shot Animation Dataset for Reference-Guided Video Generation Phantom: Subject-consistent video generation via cross-modal alignment

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T11:12:10.052472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:12:10.052472Z digest=sha256:85eaf7e1baf5b285c6cfe28f400f6f3f24c5ac8feecfd10129d654adc549b696

Observation e62295d4-8fa2-4396-8deb-fac1b7d52bf6 · inbound

PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and Enhancement cites this paper.

PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and Enhancement Phantom: Subject-consistent video generation via cross-modal alignment

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T05:30:32.974027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:30:32.974027Z digest=sha256:acc49ad79a70c52c355ac28c54913bb42d4d00c530b67ae908e7dd04367ef5a4

Observation 762ce83a-a2b1-4b2b-b0d3-7452fde4b615 · inbound

DreamActor-H1: High-Fidelity Human-Product Demonstration Video Generation via Motion-designed Diffusion Transformers cites this paper.

DreamActor-H1: High-Fidelity Human-Product Demonstration Video Generation via Motion-designed Diffusion Transformers Phantom: Subject-consistent video generation via cross-modal alignment

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T04:27:39.046915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:27:39.046915Z digest=sha256:6ee3ca3e1b7e8c862ae5cf369744edf9662729114b139c9e4e8d86b4c0ded83f

Observation cde2f2fb-60aa-483a-8704-ee42bd252af6 · inbound

Phantom-Data : Towards a General Subject-Consistent Video Generation Dataset cites this paper.

Phantom-Data : Towards a General Subject-Consistent Video Generation Dataset Phantom: Subject-consistent video generation via cross-modal alignment

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:09.885970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:09.885970Z digest=sha256:a2ede7a27234a207fb23ce400e3290d6a682f9fccf368746aa3c9a6a13202c07

Observation 78f0d332-7e44-4678-bd3e-d7f2777dd5a5 · inbound

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality cites this paper.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Phantom: Subject-consistent video generation via cross-modal alignment

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:27.649430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:51:27.649430Z digest=sha256:e9b33f4a055253a59f61e6bb72fd351d7b885bd9f571080b282a869a4d180aac

Observation 5684ad4b-79c4-42a7-8685-928f47c503ff · inbound

FantasyPortrait: Enhancing Multi-Character Portrait Animation with Expression-Augmented Diffusion Transformers cites this paper.

FantasyPortrait: Enhancing Multi-Character Portrait Animation with Expression-Augmented Diffusion Transformers Phantom: Subject-consistent video generation via cross-modal alignment

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T16:39:36.070911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:39:36.070911Z digest=sha256:48e06ef7302fd7633da34d719a7fa44bef427ae41440ecd92bed258398b717b4

Observation d07100f1-754d-40de-b195-94151d481fd2 · inbound

FantasyTalking2: Timestep-Layer Adaptive Preference Optimization for Audio-Driven Portrait Animation cites this paper.

FantasyTalking2: Timestep-Layer Adaptive Preference Optimization for Audio-Driven Portrait Animation Phantom: Subject-consistent video generation via cross-modal alignment

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-05T20:06:47.800330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:06:47.800330Z digest=sha256:b27e5c1f5c7b0c4280b2ea036a749cad6eafc7da5da39062b30f569fbbf05500

Observation eeeba291-e2d0-47b9-8c95-6822177149b3 · inbound

Identity-Preserving Text-to-Video Generation via Training-Free Prompt, Image, and Guidance Enhancement cites this paper.

Identity-Preserving Text-to-Video Generation via Training-Free Prompt, Image, and Guidance Enhancement Phantom: Subject-consistent video generation via cross-modal alignment

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T12:42:40.992340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:42:40.992340Z digest=sha256:3273ae7883c1a0b04be097769bc7de8f28111a7b4930cd8ed5928bc16d85cb67

Observation 91ed0832-9ef6-4f2a-a878-5ce4e3355301 · inbound

UniVideo: Unified Understanding, Generation, and Editing for Videos cites this paper.

UniVideo: Unified Understanding, Generation, and Editing for Videos Phantom: Subject-consistent video generation via cross-modal alignment

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T10:49:48.589583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:49:48.589583Z digest=sha256:fd1d0f8ca1433487dea615b89d2d853568019b2a7c4af58cf8acdb5ac0064ab7

Observation e5d041c1-8351-482f-9b6d-34fad7652cad · inbound

One-to-All Animation: Alignment-Free Character Animation and Image Pose Transfer cites this paper.

One-to-All Animation: Alignment-Free Character Animation and Image Pose Transfer Phantom: Subject-consistent video generation via cross-modal alignment

Reference 27

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T18:25:28.396016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-21T18:25:22.486891Z digest=sha256:008c66e3fe30d2baacd05233e5df2e3dd7701ef50f47ad534b95a8fe80d881d4

Observation b1e13b4c-d732-427f-b594-34f56251ba95 · inbound

CustomX: Unified Character, Action, and Scene Customization in Video World Models cites this paper.

CustomX: Unified Character, Action, and Scene Customization in Video World Models Phantom: Subject-consistent video generation via cross-modal alignment

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-03T15:28:55.722892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T15:28:55.722892Z digest=sha256:edd7862ee98a2d796a77a80641b8868d8a3f8c9d46ed28a9fdfc14f0f211058e

Observation 04a016b9-f2f9-4cfa-a631-94fba03f29ea · inbound

EchoTorrent: Towards Swift, Sustained, and Streaming Multi-Modal Video Generation cites this paper.

EchoTorrent: Towards Swift, Sustained, and Streaming Multi-Modal Video Generation Phantom: Subject-consistent video generation via cross-modal alignment

Reference 40

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T22:20:22.455970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T22:20:16.320171Z digest=sha256:15ccbf78cf6d31ed52e25c23917e19503936c7d3cf65d4f89bb62a76b7d67d55

Observation b16c105f-f6af-4215-ad0b-6fb297f91cdc · inbound

MVHOI: Bridge Multi-view Condition to Complex Human-Object Interaction Video Reenactment via 3D Foundation Model cites this paper.

MVHOI: Bridge Multi-view Condition to Complex Human-Object Interaction Video Reenactment via 3D Foundation Model Phantom: Subject-consistent video generation via cross-modal alignment

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T05:51:18.232473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T05:51:18.232473Z digest=sha256:27785584f4decb8b33bd470e9aafdb0392607636d205384bd8d8c676c71771dd

Observation 60704d15-2282-4390-b433-5a5fe85931b1 · inbound

Vanast: Virtual Try-On with Human Image Animation via Synthetic Triplet Supervision cites this paper.

Vanast: Virtual Try-On with Human Image Animation via Synthetic Triplet Supervision Phantom: Subject-consistent video generation via cross-modal alignment

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-10T21:55:49.210641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T20:33:25.422350Z digest=sha256:fcb0a9b806b4bad4625aa9d06d9daaef246ecc5b25753a77ff1c94865bf25842

Observation a5ba3bef-7a30-4b49-b8ba-dfc6e2dd5f94 · inbound

Evolution of Video Generative Foundations cites this paper.

Evolution of Video Generative Foundations Phantom: Subject-consistent video generation via cross-modal alignment

Reference 205

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T00:05:51.328734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T18:41:38.616611Z digest=sha256:4ff008aa391b7e18c8691011eee5b683b752bed71d774836808b446d61ae3753

Observation d4177208-82fc-49ec-9bf1-bed27d5d2af9 · inbound

Prompt Relay: Inference-Time Temporal Control for Multi-Event Video Generation cites this paper.

Prompt Relay: Inference-Time Temporal Control for Multi-Event Video Generation Phantom: Subject-consistent video generation via cross-modal alignment

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T10:01:00.416886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T15:42:35.693519Z digest=sha256:48d81f0a8aae9f0d423b1f26eba9dffbfb20be6bc29e6383c652af92cad093ff

Observation c5df77d7-6e3f-49ed-8acf-5e3c5ddccbe4 · inbound

Script-a-Video: Deep Structured Audio-visual Captions via Factorized Streams and Relational Grounding cites this paper.

Script-a-Video: Deep Structured Audio-visual Captions via Factorized Streams and Relational Grounding Phantom: Subject-consistent video generation via cross-modal alignment

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T11:11:03.692220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T15:07:45.595260Z digest=sha256:1c3e8f90299f99d1aaa57fc7e709e3a5bfd34e4134a74d8ad74fc0d54894715f

Observation 8990056c-c55a-4965-b135-a1b03ef9f68b · inbound

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey cites this paper.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Phantom: Subject-consistent video generation via cross-modal alignment

Reference 141

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T08:30:56.950219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T16:36:33.264166Z digest=sha256:dfb56713ce177ac869317d220a004f758dd6f0c34ebe193663143e185a66cccf

Observation a504db6f-6ef3-4462-bd37-eae9a89ad9b7 · inbound

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey cites this paper.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Phantom: Subject-consistent video generation via cross-modal alignment

Reference 120

Resolution
unresolved
no resolver link, observed 2026-07-12T22:04:31.302192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:04:31.302192Z digest=sha256:e31c8b22da26c71dd11efe25b5170b5d15b36a717e805b2d7eb6a626c646fb6e

Observation 78cd8348-c81b-4c6c-8519-b0c3160d0004 · inbound

OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation cites this paper.

OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation Phantom: Subject-consistent video generation via cross-modal alignment

Reference 42

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T11:06:05.893388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T15:09:02.727887Z digest=sha256:cb808d0b3649309e6e302d97ebd5ba3da549290c7a1f7146cf5c59692dad87bd

Observation 8020a8c1-a8f7-4b1e-bd94-51dde555e8a1 · inbound

ReImagine: Rethinking Controllable High-Quality Human Video Generation via Image-First Synthesis cites this paper.

ReImagine: Rethinking Controllable High-Quality Human Video Generation via Image-First Synthesis Phantom: Subject-consistent video generation via cross-modal alignment

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:01:26.592322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T02:24:57.882447Z digest=sha256:3a16a42b9be9ed9daa492438b4bd138aedc4acd5b4592bf8b3f2738a6797b4e8

Observation 9c0c0c03-b5a1-479c-bb2b-3fd35485b6c0 · inbound

FaithfulFaces: Pose-Faithful Facial Identity Preservation for Text-to-Video Generation cites this paper.

FaithfulFaces: Pose-Faithful Facial Identity Preservation for Text-to-Video Generation Phantom: Subject-consistent video generation via cross-modal alignment

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T17:21:09.635204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T17:40:00.224358Z digest=sha256:fed27cbf47b40686ba4ed2b4887cd18fc44de1eba92fd09b0d5dbe9aaa7329e4

Observation 871d318c-ba5c-438c-ae15-f110cd8a7bac · inbound

FashionChameleon: Towards Real-Time and Interactive Human-Garment Video Customization cites this paper.

FashionChameleon: Towards Real-Time and Interactive Human-Garment Video Customization Phantom: Subject-consistent video generation via cross-modal alignment

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T20:13:43.610050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T20:11:33.277349Z digest=sha256:d3940e549ec6175375a8a4e70509b4226d606b256ce3c0767e2696d43d692623

Observation 66e35ba8-be81-4f43-982a-63b0a576b155 · inbound

FashionChameleon: Towards Real-Time and Interactive Human-Garment Video Customization cites this paper.

FashionChameleon: Towards Real-Time and Interactive Human-Garment Video Customization Phantom: Subject-consistent video generation via cross-modal alignment

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T19:25:00.562996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T19:23:07.044659Z digest=sha256:c62d09be4f32262bd064955ccdaad64ca18ca070b2b2ad1474053ec7f0db895b

Observation b6a1476b-f6fc-4200-b0fc-087b41a9a896 · inbound

Bernini: Latent Semantic Planning for Video Diffusion cites this paper.

Bernini: Latent Semantic Planning for Video Diffusion Phantom: Subject-consistent video generation via cross-modal alignment

Reference 45

Resolution
metadata mismatch
arxiv_id, observed 2026-05-22T06:41:10.542258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-22T06:39:47.124605Z digest=sha256:799f4f5d13059e371bd1818448bf347b550a6e3da04a1cd106220bddc4ba511d

Observation 048d38ed-a3bc-4ad6-a5c0-aa4a4bded87a · inbound

EM-Vid: Training-Free Entity-Centric Memory for Efficient and Consistent Multi-Shot Video Generation cites this paper.

EM-Vid: Training-Free Entity-Centric Memory for Efficient and Consistent Multi-Shot Video Generation Phantom: Subject-consistent video generation via cross-modal alignment

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-25T05:00:22.425250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-25T04:57:13.479165Z digest=sha256:4321c12697f5e7dd6ec4d3fa1da9e9e6555c274c7baef892335cdc8d726e7e2d

Observation 54d0a592-7bba-47d0-b7f3-a32634311248 · inbound

MetaWorld: Scaling Multi-Agent Video World Model from Single-view Video Data cites this paper.

MetaWorld: Scaling Multi-Agent Video World Model from Single-view Video Data Phantom: Subject-consistent video generation via cross-modal alignment

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-07-01T22:56:20.244659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T14:52:30.406683Z digest=sha256:62035abe014530be7c71f1931e6d5b6005cf65640a5eaf246bd0862297d9b13e

Observation 610c77d1-d687-469a-abaf-59b53dac25f1 · inbound

Streaming Video Generation with Streaming Force Control cites this paper.

Streaming Video Generation with Streaming Force Control Phantom: Subject-consistent video generation via cross-modal alignment

Reference 37

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T17:07:12.387385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T22:14:32.465663Z digest=sha256:778518bb6fc91c193ae91ffdfa53d0f260582fd0f746568f45f02399066eb27d

Observation f48bf4c6-21ea-43b5-b88f-86cac9cca637 · inbound

HarmoView: Harmonizing Multi-View Constraints for Identity-Consistent Video Generation cites this paper.

HarmoView: Harmonizing Multi-View Constraints for Identity-Consistent Video Generation Phantom: Subject-consistent video generation via cross-modal alignment

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T04:37:36.648289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T13:49:50.272650Z digest=sha256:f85233878d12660b6d93b26f045458523e3924973e18c9bbb3c367b6f941ff27

Observation 24afdf48-3e48-4b2e-a930-603545850b29 · inbound

ARGUS: Stacked Multi-View Identity Mosaic Injection for Subject-Preserving Video Generation cites this paper.

ARGUS: Stacked Multi-View Identity Mosaic Injection for Subject-Preserving Video Generation Phantom: Subject-consistent video generation via cross-modal alignment

Reference 29

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T08:17:45.947052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T10:46:56.871174Z digest=sha256:07921c04c0046c2b08945e96aa3e612515003d68168e704da0160d45a98727cf

Observation 68f86b7d-16e3-4865-a3fa-70fb1297d568 · inbound

GroundShot: Visually Consistent Multi-Shot Long Video Generation via Entity-Grounded Shot Scheduling cites this paper.

GroundShot: Visually Consistent Multi-Shot Long Video Generation via Entity-Grounded Shot Scheduling Phantom: Subject-consistent video generation via cross-modal alignment

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T03:19:30.381652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-26T18:15:51.862192Z digest=sha256:e0b175324b88c9e4a40e3e1108e7087f2a89b06fcb56833b9b16eb356ae64c2e

Observation 73bda7db-f3f1-424c-89bf-43f5dcb8f4a5 · inbound

GroundShot: Visually Consistent Multi-Shot Long Video Generation via Entity-Grounded Shot Scheduling cites this paper.

GroundShot: Visually Consistent Multi-Shot Long Video Generation via Entity-Grounded Shot Scheduling Phantom: Subject-consistent video generation via cross-modal alignment

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-02T10:48:05.537673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:48:05.537673Z digest=sha256:719ad55ac63133642c3b6e2da4961f6ddfdb10b49211c5f3335a45345d46507e

Observation ecf247e4-33a2-4343-bb22-f847989a56a7 · inbound

A Good Talk Does not Look Like a Summary, It Teaches You! Measuring Takeaways from Paper-to-Video Talks cites this paper.

A Good Talk Does not Look Like a Summary, It Teaches You! Measuring Takeaways from Paper-to-Video Talks Phantom: Subject-consistent video generation via cross-modal alignment

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T15:45:48.892408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-30T00:59:50.478584Z digest=sha256:df82d13461863ed2bfea28c2dfcd74d0bca9d876ce5b8f4adaacd31fc62f1676

Observation 95c4ed9a-fded-45a9-8f3e-332d66048293 · inbound

Keyframe-Anchored Identity Preservation for Sequential-Action Video Generation cites this paper.

Keyframe-Anchored Identity Preservation for Sequential-Action Video Generation Phantom: Subject-consistent video generation via cross-modal alignment

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T16:30:02.172800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:30:02.172800Z digest=sha256:a672fc47a4077a1fa0ac7d8f89cb6c0f4d65ceec2e31d5874019de322ba7024a

Observation d111140e-0430-4dbc-a018-a4527dfbfc31 · inbound

ID-V2V: Identity-Preserving Video Restylization cites this paper.

ID-V2V: Identity-Preserving Video Restylization Phantom: Subject-consistent video generation via cross-modal alignment

Reference 191

Resolution
unresolved
no resolver link, observed 2026-08-01T04:27:31.566712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T04:27:31.566712Z digest=sha256:a70743b50db10a8b4afb7e8f224b6aef0e6e99dabd5ea41d9e56937d27d2727c