Pith. sign in

Paper Citation Record · LEDGER

GroupVideo: Multi-Identity Customized Text-to-Video Generation

As of 9 August 2026, this Paper Citation Record lists 83 of 83 outbound references and 0 inbound Pith citation observations for arXiv:2607.21027.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.21027 v1

Coverage vector

measured 83 of 83 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-01T08:45:56.546648Z

measured 83 of 83 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

83 of 83 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved83
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 04979821-5378-4331-b7cf-13139b70f22a · outbound

This paper cites an unresolved cited work.

GroupVideo: Multi-Identity Customized Text-to-Video Generation Unresolved cited work

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-01T08:45:47.849989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:45:47.849989Z digest=sha256:a6619972f5ada4a98d69a739d7ecb8fff29e6f5acabdc1f831d2faf1e3f081b8

Observation dde5b637-a5a7-4d63-8eb3-682a9d00f610 · outbound

This paper cites An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual Inversion.

GroupVideo: Multi-Identity Customized Text-to-Video Generation An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual Inversion

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-01T08:45:47.939453Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:45:47.939453Z digest=sha256:95b2324a2f812fd95dd46196721ce4306e993e8af342941e8a5342fe77ce9540

Observation 52ff2417-9cd0-4b06-9177-b408efce37e3 · outbound

This paper cites an unresolved cited work.

GroupVideo: Multi-Identity Customized Text-to-Video Generation Unresolved cited work

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-01T08:45:48.039742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:45:48.039742Z digest=sha256:c72c452be92076b0a48272feb479d182ea57075f32e860d6cb2dca36a3916230

Observation 8b06610e-e8fc-4373-8eca-63b02eb97a07 · outbound

This paper cites ID-Animator: Zero-Shot Identity-Preserving Human Video Generation.

GroupVideo: Multi-Identity Customized Text-to-Video Generation ID-Animator: Zero-Shot Identity-Preserving Human Video Generation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-01T08:45:48.164156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:45:48.164156Z digest=sha256:56d2ba93b4b37958ba3a22a8526233b229b92ec2bac5fc484a692519c1adec9c

Observation 1da02510-44b9-4241-a697-8be32731aab8 · outbound

This paper cites EchoVideo: Identity-Preserving Human Video Generation by Multimodal Feature Fusion.

GroupVideo: Multi-Identity Customized Text-to-Video Generation EchoVideo: Identity-Preserving Human Video Generation by Multimodal Feature Fusion

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-01T08:45:48.274281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:45:48.274281Z digest=sha256:514161e18e0004d3c98b17cd0f050fa23546b799b0ce2d8129ad0d49e365e220

Observation 879256db-3409-48ed-9da5-af0ab32ada93 · outbound

This paper cites an unresolved cited work.

GroupVideo: Multi-Identity Customized Text-to-Video Generation Unresolved cited work

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-01T08:45:48.340361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:45:48.340361Z digest=sha256:9ad40a439e3f2068cc96adf52932c1101d114cd6abc0428862ca3cbe77bc4556

Observation 8ddf9296-090a-4599-bc4f-54f614f9b393 · outbound

This paper cites Ingredients: Blending Custom Photos with Video Diffusion Transformers.

GroupVideo: Multi-Identity Customized Text-to-Video Generation Ingredients: Blending Custom Photos with Video Diffusion Transformers

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-01T08:45:48.400345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:45:48.400345Z digest=sha256:fe362de239c438babef290792027f8e254c15efc29dce05bc9d34a5a618f1da6

Observation 3302ffb3-3434-41a4-8921-7fe0f811fdaf · outbound

This paper cites Peebles and S.

GroupVideo: Multi-Identity Customized Text-to-Video Generation Peebles and S

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-01T08:45:48.486586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:45:48.486586Z digest=sha256:79ee892e1e9c42297af6c6afeae5d399e13275ff690ed58bab129737faf5b83f

Observation a1526efe-a4c7-4a29-b777-6ff343080701 · outbound

This paper cites Generative adversarial networks.

GroupVideo: Multi-Identity Customized Text-to-Video Generation Generative adversarial networks

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-01T08:45:48.573792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:45:48.573792Z digest=sha256:3b2591aad4f44dd900c1afc3fb82b3642e978925ee184ade166a66d11ed08e56

Observation 3eef3856-9cad-442b-8d17-6a90ef3e7dee · outbound

This paper cites Karras, S.

GroupVideo: Multi-Identity Customized Text-to-Video Generation Karras, S

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-01T08:45:48.636063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:45:48.636063Z digest=sha256:2eddac51cfabcbb75d4b5182bb190c8fac2358cd17bb5bca4f5de47effbeb4db

Observation 44b772a2-7647-48eb-b168-a940d5d48f2e · outbound

This paper cites an unresolved cited work.

GroupVideo: Multi-Identity Customized Text-to-Video Generation Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T08:45:48.721612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:45:48.721612Z digest=sha256:a1022d8323fef7c07b59edb5818d2d945d3a206c8651ad635ef4cc0426b7d53d

Observation b7fc6a97-fa8e-461c-99b6-7a3560d986ef · outbound

This paper cites K ¨oksal, K.

GroupVideo: Multi-Identity Customized Text-to-Video Generation K ¨oksal, K

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-01T08:45:48.807772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:45:48.807772Z digest=sha256:130aeafe08ec83313fe0e5237af51993c21b4479f562d7e461314bc109271d20

Observation 41c282ae-151f-450d-9534-5bd66bca3731 · outbound

This paper cites Taming transformers for high- resolution image synthesis.

GroupVideo: Multi-Identity Customized Text-to-Video Generation Taming transformers for high- resolution image synthesis

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-01T08:45:48.900449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:45:48.900449Z digest=sha256:7b302944961f9af8cb9ca5c6abbcf0dd852730883863b4ff09d3405cfe2b39ba

Observation f01819d9-7891-4554-8b92-a003e39b4b78 · outbound

This paper cites Scaling Autoregressive Models for Content-Rich Text-to-Image Generation.

GroupVideo: Multi-Identity Customized Text-to-Video Generation Scaling Autoregressive Models for Content-Rich Text-to-Image Generation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T08:45:48.961920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:45:48.961920Z digest=sha256:ab4ab806ff7ec435f2a069c6f839f115ac75f398168a46b4579543074c44f3fd

Observation e4b3e649-1a7f-4ef7-b35d-16343ed9a549 · outbound

This paper cites Rombach, A.

GroupVideo: Multi-Identity Customized Text-to-Video Generation Rombach, A

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-01T08:45:49.051032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:45:49.051032Z digest=sha256:b61540818b4f45a399b96f8412d5d860b20948af0867f85e4cc549f5ac8dd98c

Observation e4cc1a4d-e696-4dc2-a9c3-0de8f436ec98 · outbound

This paper cites L.; Ghasemipour, K.; Gontijo Lopes, R.; Karagol Ayan, B.; Salimans, T.; et al.

GroupVideo: Multi-Identity Customized Text-to-Video Generation L.; Ghasemipour, K.; Gontijo Lopes, R.; Karagol Ayan, B.; Salimans, T.; et al

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-01T08:45:49.142905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:45:49.142905Z digest=sha256:2dfc1830ebe235e64e116e4a5bffc0dd9648975fb3763942a462ef88975ff348

Observation 8c975695-e6fb-4676-b5d3-3eb7b97980c3 · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

GroupVideo: Multi-Identity Customized Text-to-Video Generation SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-01T08:45:49.235738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:45:49.235738Z digest=sha256:d3718d740deb9e670b5eaa33780bd31947a29151d7dbcac2b3803131292d44a6

Observation 9cea2026-2787-4f02-9c29-187d50d1322a · outbound

This paper cites Unialignment: Semantic alignment for unified image generation, understanding, manipulation and perception.

GroupVideo: Multi-Identity Customized Text-to-Video Generation Unialignment: Semantic alignment for unified image generation, understanding, manipulation and perception

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T08:45:49.318127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:45:49.318127Z digest=sha256:cd1574b09a287cd0ca65f025c39da97e5ab69da7897e17d2e30b2897f41137bf

Observation f585b9ff-f8e4-47ff-8fbd-d1299f8f12e7 · outbound

This paper cites 3SGen: Unified Subject, Style, and Structure-Driven Image Generation with Adaptive Task-specific Memory, 2025, arXiv preprint arXiv:2512.19271.

GroupVideo: Multi-Identity Customized Text-to-Video Generation 3SGen: Unified Subject, Style, and Structure-Driven Image Generation with Adaptive Task-specific Memory, 2025, arXiv preprint arXiv:2512.19271

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-01T08:45:49.409780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:45:49.409780Z digest=sha256:5530a7de319fb0d7ee29944f709153e9746bf03990641673f7519b92e95ccbb6

Observation 51798d8f-8776-43b8-8ee3-a4225cc2c005 · outbound

This paper cites Fine-Grained Text-to-Image Synthesis with Semantic Refinement.

GroupVideo: Multi-Identity Customized Text-to-Video Generation Fine-Grained Text-to-Image Synthesis with Semantic Refinement

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-01T08:45:49.496661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:45:49.496661Z digest=sha256:4b1e7485413d1c891ccb1a1e1768d7294946d24a6a2655385428b9fafdaee9b4

Observation 0e4e8d98-ae7c-4924-8477-eb7e2d9bcc00 · outbound

This paper cites an unresolved cited work.

GroupVideo: Multi-Identity Customized Text-to-Video Generation Unresolved cited work

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-01T08:45:49.584438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:45:49.584438Z digest=sha256:5cf7c545b752f142eec553b3fafb5a21f77b662c0fcdd8fc419c25abc08ad2ca

Observation 6b12dabc-3e86-4776-971c-d1e79deffa23 · outbound

This paper cites an unresolved cited work.

GroupVideo: Multi-Identity Customized Text-to-Video Generation Unresolved cited work

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-01T08:45:49.664719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:45:49.664719Z digest=sha256:e4b51ee30a3b395e54f925a46f2893426f9923f10ac6891b37ea6c9d6433cd57

Observation a793c566-9794-46f4-bcc3-b3b1f0fe9f64 · outbound

This paper cites Video generation models as world simulators.

GroupVideo: Multi-Identity Customized Text-to-Video Generation Video generation models as world simulators

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-01T08:45:49.755672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:45:49.755672Z digest=sha256:6871a1bb7cec98b7c5c52ae291f6e3780a690c7d1d8ef4ac4a571ed4986d7647

Observation b5965f3f-70a2-4187-b623-eed03b805114 · outbound

This paper cites Latte: Latent Diffusion Transformer for Video Generation.

GroupVideo: Multi-Identity Customized Text-to-Video Generation Latte: Latent Diffusion Transformer for Video Generation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-01T08:45:49.844026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:45:49.844026Z digest=sha256:7bff159300185847d29865982cad562b7e53c205bc016a629ed6bc489703106a

Observation 21201fe6-024b-481b-ac1b-ce2d3c57562e · outbound

This paper cites Open-Sora Plan: Open-Source Large Video Generation Model.

GroupVideo: Multi-Identity Customized Text-to-Video Generation Open-Sora Plan: Open-Source Large Video Generation Model

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-01T08:45:49.939631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:45:49.939631Z digest=sha256:d982e4379ab2355fd9387243ccf50ff4a5eca6759d122a8837bc972b5101b134

Observation f4477a6b-15ab-41a0-b874-d7c967e631e1 · outbound

This paper cites Open-Sora: Democratizing Efficient Video Production for All.

GroupVideo: Multi-Identity Customized Text-to-Video Generation Open-Sora: Democratizing Efficient Video Production for All

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-01T08:45:50.028879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:45:50.028879Z digest=sha256:36d189f27d487f5275a6b5f7f1e2ccda005645b075e1081dc183ce76dcafbe23

Observation a86cc048-0fde-4905-978f-df6072d88e4f · outbound

This paper cites CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer.

GroupVideo: Multi-Identity Customized Text-to-Video Generation CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-01T08:45:50.148925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:45:50.148925Z digest=sha256:8b31050ac61efe6fce3bab3313b9678b8d02a1972bac9042c5229095103469ea

Observation 41685536-516c-413a-ae04-3a1e07d9f2a3 · outbound

This paper cites HunyuanVideo: A Systematic Framework For Large Video Generative Models.

GroupVideo: Multi-Identity Customized Text-to-Video Generation HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-01T08:45:50.326139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:45:50.326139Z digest=sha256:aefa51b450a3eb55ac76461d7bfcde5f842ae2b1b8284c7e33fe6381db654435

Observation 4601cdfb-62c4-4bba-a418-da37901b710f · outbound

This paper cites TokLIP: Marry Visual Tokens to CLIP for Multimodal Comprehension and Generation.

GroupVideo: Multi-Identity Customized Text-to-Video Generation TokLIP: Marry Visual Tokens to CLIP for Multimodal Comprehension and Generation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-01T08:45:50.500497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:45:50.500497Z digest=sha256:11ac63031feb0a6d1b8ee9b8cd9d147957c64e3c1fd49b1a34b34adf5b0a4877

Observation be5a43f0-8ed9-4f93-b84c-0ea0d2e072ef · outbound

This paper cites IV-CoT: Implicit Visual Chain-of-Thought for Structure-Aware Text-to-Image Generation.

GroupVideo: Multi-Identity Customized Text-to-Video Generation IV-CoT: Implicit Visual Chain-of-Thought for Structure-Aware Text-to-Image Generation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-01T08:45:50.621647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:45:50.621647Z digest=sha256:1f9e6009f432b281528c59048fdf34edb66d0f8c9b864d89cce6bc9e247d6373

Observation 6c21d9fe-24ea-4e25-9686-17b3e41d81fa · outbound

This paper cites Concept-Guided Tokenization: Closing the Gap Between Reconstruction and Generation, 2026.

GroupVideo: Multi-Identity Customized Text-to-Video Generation Concept-Guided Tokenization: Closing the Gap Between Reconstruction and Generation, 2026

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-01T08:45:50.787703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:45:50.787703Z digest=sha256:9b2ff1c140e0102db55ffae676820220ab5ce3f8c906eb47f418b09d1cf6e9fa

Observation fffe54f4-e2f7-4a6d-864d-de6d7ca1df70 · outbound

This paper cites Shaping a Stabilized Video by Mitigating Unintended Changes for Concept-Augmented Video Editing.

GroupVideo: Multi-Identity Customized Text-to-Video Generation Shaping a Stabilized Video by Mitigating Unintended Changes for Concept-Augmented Video Editing

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-01T08:45:50.904043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:45:50.904043Z digest=sha256:8ee9d602f244b2eda3079020fc4dd13d1bd8ca6f224394dfc3206d3699bdb65c

Observation 3b3fbf76-5f74-44de-a6af-511c5700fc85 · outbound

This paper cites an unresolved cited work.

GroupVideo: Multi-Identity Customized Text-to-Video Generation Unresolved cited work

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-01T08:45:51.023640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:45:51.023640Z digest=sha256:cd75a9976a7894965ba6a4144e8eb65449a534617eb470cfcfe2c21d89a85ed2

Observation abcee6ac-ee23-4347-8b50-5636d92a92a7 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

GroupVideo: Multi-Identity Customized Text-to-Video Generation LoRA: Low-Rank Adaptation of Large Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-01T08:45:51.195504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:45:51.195504Z digest=sha256:643b0d3c33d97ee122de88c74504a4e979edd1191b113df63314aa1c91aab8a5

Observation 51697e3a-1efa-4a92-b605-d00230f453c9 · outbound

This paper cites Zhang, T.

GroupVideo: Multi-Identity Customized Text-to-Video Generation Zhang, T

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-01T08:45:51.321655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:45:51.321655Z digest=sha256:6d0fd3f9205e0281f4249083b58f19bbb928f585b37bf7b966e8f9841a546ec4

Observation 4a915594-91db-4e6e-8dd0-c7e2e7f2adfd · outbound

This paper cites an unresolved cited work.

GroupVideo: Multi-Identity Customized Text-to-Video Generation Unresolved cited work

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-01T08:45:51.449963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:45:51.449963Z digest=sha256:c9309898ccad900a317dcb98edcc14007651477c848c04d9ffc6e1d2b5695287

Observation 33a1c856-56d2-479f-a4dc-9b817171a0bc · outbound

This paper cites InstantID: Zero-shot Identity-Preserving Generation in Seconds.

GroupVideo: Multi-Identity Customized Text-to-Video Generation InstantID: Zero-shot Identity-Preserving Generation in Seconds

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-01T08:45:51.599832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:45:51.599832Z digest=sha256:8b4daf77b165bcf7ff3d3983d838bead20cec07f5262c65167b99fb0e6aff45f

Observation 68941400-d86c-445c-957a-a5bf51b6365a · outbound

This paper cites IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models.

GroupVideo: Multi-Identity Customized Text-to-Video Generation IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-01T08:45:51.690564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:45:51.690564Z digest=sha256:7e66f28cd299dc5725d3e679bdaed274ea29795b32e7e30e5d28ecca7ae0a131

Observation 84972062-0a74-46ce-8da5-5d235bd1b3d1 · outbound

This paper cites PuLID: Pure and Lightning ID Customization via Contrastive Alignment.

GroupVideo: Multi-Identity Customized Text-to-Video Generation PuLID: Pure and Lightning ID Customization via Contrastive Alignment

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-01T08:45:51.815191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:45:51.815191Z digest=sha256:8a90d15c1fd67853610718594c7b4f220ac0a0445c74eff4f36665f776b7f838

Observation bb98f02e-7e56-4e1a-99c9-443d1b355f8f · outbound

This paper cites Blip-diffusion: Pre-trained subject represen- tation for controllable text-to-image generation and editing.

GroupVideo: Multi-Identity Customized Text-to-Video Generation Blip-diffusion: Pre-trained subject represen- tation for controllable text-to-image generation and editing

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-01T08:45:51.894743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:45:51.894743Z digest=sha256:da851f89b0580aedf599a9adc9a9c432db4362da085b6c61575d4444ef488e62

Observation 55f2d88f-c001-4102-afb0-77c31c4a3512 · outbound

This paper cites an unresolved cited work.

GroupVideo: Multi-Identity Customized Text-to-Video Generation Unresolved cited work

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-01T08:45:51.984244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:45:51.984244Z digest=sha256:7dbc9ea63a9a60cb6de9cd65ad23c5d532934a5ec3661d7b931e08c99c4190ef

Observation a0dfd4fd-405f-4f8e-9b2b-2aa1266af31c · outbound

This paper cites Dreamidentity: Enhanced editability for efficient face-identity preserved image generation.

GroupVideo: Multi-Identity Customized Text-to-Video Generation Dreamidentity: Enhanced editability for efficient face-identity preserved image generation

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-01T08:45:52.070503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:45:52.070503Z digest=sha256:fd273698a07a5a5f2571f772c3e86fd5104efaa3c1679415239bcc5ae5426016

Observation bc93d4e3-6330-4613-a818-f1cc88b5bbcc · outbound

This paper cites H.; Chechik, G.; and Cohen- Or, D.

GroupVideo: Multi-Identity Customized Text-to-Video Generation H.; Chechik, G.; and Cohen- Or, D

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-01T08:45:52.134146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:45:52.134146Z digest=sha256:881f0d152482952e2cbeb282dbb885e2bad9d0e3a82a549448ea84a64bf32d35

Observation b3f4765c-914b-43e8-952f-ebd0c5f221dc · outbound

This paper cites X-portrait: Expressive portrait animation with hierarchical motion attention.

GroupVideo: Multi-Identity Customized Text-to-Video Generation X-portrait: Expressive portrait animation with hierarchical motion attention

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-01T08:45:52.223745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:45:52.223745Z digest=sha256:18c7749b3400a774ebde01d24642cf8ecb0ccdcf7e746450b139376328d38266

Observation 688ab11b-1772-4358-9b41-dc34424caa9f · outbound

This paper cites an unresolved cited work.

GroupVideo: Multi-Identity Customized Text-to-Video Generation Unresolved cited work

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-01T08:45:52.324072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:45:52.324072Z digest=sha256:6840760609c6ea6e61b02278e1b9e92b1a9d824cdd407c1c6d53bb4685257ff2

Observation f5380f45-5080-41d0-b506-697fd6fa03d2 · outbound

This paper cites PoseGen: In-Context LoRA Finetuning for Pose-Controllable Long Human Video Generation.

GroupVideo: Multi-Identity Customized Text-to-Video Generation PoseGen: In-Context LoRA Finetuning for Pose-Controllable Long Human Video Generation

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-01T08:45:52.409992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:45:52.409992Z digest=sha256:9f017e3e38eb11a77880fe93671115f534cfd18d7ce5e7e9d8d1ddd8100be39b

Observation 6e99865e-3d8d-4c27-badb-daddaa3011ae · outbound

This paper cites AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning.

GroupVideo: Multi-Identity Customized Text-to-Video Generation AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-01T08:45:52.490771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:45:52.490771Z digest=sha256:6371a66b4f9b5d6226e27f3f64f2a80b4dce02b18ad88b5fee9afe4042f14cd0

Observation 53d9f9ea-7673-48ff-ad73-4811b39ac0cc · outbound

This paper cites I2VGen-XL: High-Quality Image-to-Video Synthesis via Cascaded Diffusion Models.

GroupVideo: Multi-Identity Customized Text-to-Video Generation I2VGen-XL: High-Quality Image-to-Video Synthesis via Cascaded Diffusion Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-01T08:45:52.604815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:45:52.604815Z digest=sha256:2b018691719c2924ef53946f355115caeb118ca4fb91111d1aafe2bacd2b5303

Observation b59748e8-59e6-48e1-af7a-4b698411acaa · outbound

This paper cites Dynamicrafter: Animating open-domain images with video diffusion priors.

GroupVideo: Multi-Identity Customized Text-to-Video Generation Dynamicrafter: Animating open-domain images with video diffusion priors

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-01T08:45:52.768788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:45:52.768788Z digest=sha256:33d5fa7e0c1ea8001cc16a4946fba7fa99f0999da4095843c31051bf774f21ec

Observation 5ab24a74-92fa-4fcd-b1c4-98f73a03c691 · outbound

This paper cites Magic-Me: Identity-Specific Video Customized Diffusion.

GroupVideo: Multi-Identity Customized Text-to-Video Generation Magic-Me: Identity-Specific Video Customized Diffusion

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-01T08:45:52.867257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:45:52.867257Z digest=sha256:f78f877d1c2baee13bb4b841e3a7e3707061b2386112683bd14c94bddbfe0564

Observation 50a4f281-3658-4746-8cb1-ab31f5f95573 · outbound

This paper cites an unresolved cited work.

GroupVideo: Multi-Identity Customized Text-to-Video Generation Unresolved cited work

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-01T08:45:52.951559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:45:52.951559Z digest=sha256:8e371143f5c990fb27f42215e106966c61bd5b7ba3f39ae905f25dd6736dd0f6

Observation a50c2573-385d-4397-81d8-52e243695af6 · outbound

This paper cites an unresolved cited work.

GroupVideo: Multi-Identity Customized Text-to-Video Generation Unresolved cited work

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-01T08:45:53.102394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:45:53.102394Z digest=sha256:40f9dccdde8f764ba28a01f27e7d1487da9f63c5735353e9239c981a07865a4b

Observation 94e0d4ef-e82d-44c7-93e2-a57112bba0f3 · outbound

This paper cites Jiang et al., ”VideoBooth: Diffusion-based Video Generation with Image Prompts,” 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, W A, USA, pp.

GroupVideo: Multi-Identity Customized Text-to-Video Generation Jiang et al., ”VideoBooth: Diffusion-based Video Generation with Image Prompts,” 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, W A, USA, pp

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-01T08:45:53.192894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:45:53.192894Z digest=sha256:093efe5facfdb9d00805390108acfed0b85a4dacd2e2c2b7c064a82e63d81cd1

Observation 61130903-dadb-44e7-8806-5df1efc8ae0f · outbound

This paper cites Multi- concept customization of text-to-image diffusion.

GroupVideo: Multi-Identity Customized Text-to-Video Generation Multi- concept customization of text-to-image diffusion

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-01T08:45:53.296362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:45:53.296362Z digest=sha256:61532c4c51d25729e371d8afa05de0e76ea9f8365c61e2e4602489c0b0ba7f34

Observation 50a5ab12-0aaa-403d-bbe6-a3c9a066e022 · outbound

This paper cites Cones 2: Customizable image synthesis with multiple subjects.

GroupVideo: Multi-Identity Customized Text-to-Video Generation Cones 2: Customizable image synthesis with multiple subjects

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-01T08:45:53.442151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:45:53.442151Z digest=sha256:100db9c96c00fd05b4884c1ee7c245b0106e8b1ee1a06f653c01078982a4a3a1

Observation 9e718114-07ac-4539-a70b-c6a1291695b0 · outbound

This paper cites an unresolved cited work.

GroupVideo: Multi-Identity Customized Text-to-Video Generation Unresolved cited work

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-01T08:45:53.551537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:45:53.551537Z digest=sha256:c5dac150ffd9851935d19aa7dc968f09b2551a8a6a5533ceddff8aaaf77db725

Observation 0ae6c2b2-d320-4abc-aac2-7e701455f22c · outbound

This paper cites Jiang, Q.

GroupVideo: Multi-Identity Customized Text-to-Video Generation Jiang, Q

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-01T08:45:53.668440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:45:53.668440Z digest=sha256:1bc582afe2693241c69a167e226db09c7f0dc33be3431cf0c416b1a31f4c2fde

Observation 57edd10e-546e-436e-8861-bd09e26538d4 · outbound

This paper cites InstantFamily: Masked Attention for Zero-shot Multi-ID Image Generation.

GroupVideo: Multi-Identity Customized Text-to-Video Generation InstantFamily: Masked Attention for Zero-shot Multi-ID Image Generation

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-01T08:45:53.750808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:45:53.750808Z digest=sha256:f37f6a462b9ab944795d204fe4654153fa41fc4b0b0c51ea904290e6df83f13a

Observation 5c2f5e89-be5e-4505-ab83-4167fa72519e · outbound

This paper cites Z.; Shi, Y .; Chen, Y .; Fan, Z.; Xiao, W.; Zhao, R.; Chang, S.; Wu, W.; et al.

GroupVideo: Multi-Identity Customized Text-to-Video Generation Z.; Shi, Y .; Chen, Y .; Fan, Z.; Xiao, W.; Zhao, R.; Chang, S.; Wu, W.; et al

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-01T08:45:53.864276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:45:53.864276Z digest=sha256:a69267bf80ce8b4e54b41012b085d9c840a6fbdd6bf772354f50ace7e682785d

Observation a4488120-ede3-4c7b-aad5-77c802a95fe2 · outbound

This paper cites T.; Durand, F.; and Han, S.

GroupVideo: Multi-Identity Customized Text-to-Video Generation T.; Durand, F.; and Han, S

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-01T08:45:53.980171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:45:53.980171Z digest=sha256:410f5d30c36f9290ec7b830fdedd4af99331f84080efb4ab03c4f7e59553dba1

Observation 9674dc1c-ebe5-4b26-9a5a-c5c319b477c9 · outbound

This paper cites Wang et al., ”StableIdentity: Inserting Anybody into Anywhere at First Sight,” in IEEE Transactions on Multimedia, 2025.

GroupVideo: Multi-Identity Customized Text-to-Video Generation Wang et al., ”StableIdentity: Inserting Anybody into Anywhere at First Sight,” in IEEE Transactions on Multimedia, 2025

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-01T08:45:54.104811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:45:54.104811Z digest=sha256:50e10f1cdb91b32efeb54afeb004e41e34a7327c6de89ebe40c953380be78732

Observation 4eee05e7-9df3-4ded-bc7f-9388351152a5 · outbound

This paper cites Moa: Mixture-of-attention for subject-context disentanglement in person- alized image generation.

GroupVideo: Multi-Identity Customized Text-to-Video Generation Moa: Mixture-of-attention for subject-context disentanglement in person- alized image generation

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-01T08:45:54.187980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:45:54.187980Z digest=sha256:389ab1cd5e62426d69b1e725b8ae0299e5ea943ccdcc3468932a29d870a54c71

Observation 9c30dbaf-a471-4e77-affc-d426df0c0b79 · outbound

This paper cites UniPortrait: A Unified Framework for Identity-Preserving Single- and Multi-Human Image Personalization.

GroupVideo: Multi-Identity Customized Text-to-Video Generation UniPortrait: A Unified Framework for Identity-Preserving Single- and Multi-Human Image Personalization

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-01T08:45:54.264478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:45:54.264478Z digest=sha256:419d1d75100eb4ba91eec4a5b80cebf09d8d4d2ea4b04a6b1dcb5b1dd547a5a9

Observation 660a1039-2c13-4019-a163-942f9a440ec2 · outbound

This paper cites an unresolved cited work.

GroupVideo: Multi-Identity Customized Text-to-Video Generation Unresolved cited work

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-01T08:45:54.387514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:45:54.387514Z digest=sha256:4c3c22017dd3ff2b44814838350fcbab2c8d02660386e0c4ffc4098b80debe6a

Observation 9263479e-2623-41ce-8c90-3255123cf0d5 · outbound

This paper cites Prompt disentanglement via language guidance and representation alignment for domain generalization.

GroupVideo: Multi-Identity Customized Text-to-Video Generation Prompt disentanglement via language guidance and representation alignment for domain generalization

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-01T08:45:54.496271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:45:54.496271Z digest=sha256:f6f8bf1abde7c342d9be5ea7a34f2bc476e030ec7248daccb902639b67133d17

Observation 5b1a376d-3df0-45e6-86d9-d2df267195c9 · outbound

This paper cites Isolating Interference Factors for Robust Cloth-Changing Person Re-Identification.

GroupVideo: Multi-Identity Customized Text-to-Video Generation Isolating Interference Factors for Robust Cloth-Changing Person Re-Identification

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-01T08:45:54.623013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:45:54.623013Z digest=sha256:96052657bae3a3c045aced9ea38a42585ac187b661f329b194ebb9324606e2d2

Observation 1045384a-ebfc-405b-b663-7e630c8beb6b · outbound

This paper cites Semantic-Aligned Learning with Collaborative Refinement for Unsupervised VI-ReID.

GroupVideo: Multi-Identity Customized Text-to-Video Generation Semantic-Aligned Learning with Collaborative Refinement for Unsupervised VI-ReID

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-01T08:45:54.748922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:45:54.748922Z digest=sha256:0083b1fa06a09ad796d855ee7a4f8c40c684411e4815e8c5cc9534034a47c0af

Observation 5c604d6c-64b5-455d-a0ca-d5e11d37c6a1 · outbound

This paper cites Arcface: Additive an- gular margin loss for deep face recognition.

GroupVideo: Multi-Identity Customized Text-to-Video Generation Arcface: Additive an- gular margin loss for deep face recognition

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-01T08:45:54.876312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:45:54.876312Z digest=sha256:563a3b2dba83d43e6c2fb8d500c2311d18e2b40f951a678cf0c57f30ab2bef17

Observation c3ceb86b-0cfd-4972-b063-d3adfe16fa23 · outbound

This paper cites Blip-2: Bootstrapping language- image pre-training with frozen image encoders and large language models.

GroupVideo: Multi-Identity Customized Text-to-Video Generation Blip-2: Bootstrapping language- image pre-training with frozen image encoders and large language models

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-01T08:45:55.017512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:45:55.017512Z digest=sha256:41539f307868bcf585eb47f314ccc9f8c50d9ad9ad890b880582540af6ba9a95

Observation e7cbcd72-ae7e-44a3-b208-c31a67f8affa · outbound

This paper cites C.; Cai, W.; and Wu, W.

GroupVideo: Multi-Identity Customized Text-to-Video Generation C.; Cai, W.; and Wu, W

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-01T08:45:55.192855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:45:55.192855Z digest=sha256:63980a37f62fed370648b0c8788cbe04610668175e5c49d221627c9aca23231c

Observation 905c32c9-bc04-4fd0-b72b-e918be86cffc · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

GroupVideo: Multi-Identity Customized Text-to-Video Generation Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-01T08:45:55.326589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:45:55.326589Z digest=sha256:e58ef97a97f33a7bc3ecb4bdbd7702296aad915d7798643e5ba9af9834d7f1cd

Observation 88eed811-e3ed-4617-b5b7-34feddf62511 · outbound

This paper cites You only look once: Unified, real-time object detection.

GroupVideo: Multi-Identity Customized Text-to-Video Generation You only look once: Unified, real-time object detection

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-01T08:45:55.498705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:45:55.498705Z digest=sha256:99ced1c7b59bb3a4edc13326e6e0e09c0adc5cb5d250635231bd12db300881e1

Observation 8f765b00-60ec-4d34-be47-d10aebbc786d · outbound

This paper cites an unresolved cited work.

GroupVideo: Multi-Identity Customized Text-to-Video Generation Unresolved cited work

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-01T08:45:55.631554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:45:55.631554Z digest=sha256:784eaf38f1d495c869b5be45bcbf744ab4b9462ee45dc9984feeb7967a98ab7d

Observation 95705e5c-a329-434e-8266-0bec69185265 · outbound

This paper cites CogVLM: Visual Expert for Pretrained Language Models.

GroupVideo: Multi-Identity Customized Text-to-Video Generation CogVLM: Visual Expert for Pretrained Language Models

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-01T08:45:55.863782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:45:55.863782Z digest=sha256:ed61448b041f1e9b645039db2391f4595eb78a2d7074a011e37bf369ddf44fc9

Observation 48c89310-2f36-4c5b-817c-ef31ff4fec05 · outbound

This paper cites Concat-ID: Towards Universal Identity-Preserving Video Synthesis.

GroupVideo: Multi-Identity Customized Text-to-Video Generation Concat-ID: Towards Universal Identity-Preserving Video Synthesis

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-01T08:45:55.949098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:45:55.949098Z digest=sha256:984d84464dbcec57ca459b0777ca693c4a2bb7fddf3797262d615c7794055cbc

Observation 6a0ab88e-903f-40c6-8814-7232c89e803e · outbound

This paper cites an unresolved cited work.

GroupVideo: Multi-Identity Customized Text-to-Video Generation Unresolved cited work

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-01T08:45:56.022224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:45:56.022224Z digest=sha256:3e124c0e16ccda659d3b9030cebf1f13a007302edbeb204d2b3397fc14c42d51

Observation 31b07728-66f6-412e-9c74-2fb5a5f7264f · outbound

This paper cites Facenet: A unified embedding for face recognition and clustering.

GroupVideo: Multi-Identity Customized Text-to-Video Generation Facenet: A unified embedding for face recognition and clustering

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-01T08:45:56.070192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:45:56.070192Z digest=sha256:d65b52ad3c6751cfa15c0f072579a2eab1b0aa1688085d71e1f63f1b3f516cd2

Observation b7b3befd-d387-4e60-a430-ca0a169c8d65 · outbound

This paper cites CLIPScore: A Reference-free Evaluation Metric for Image Captioning.

GroupVideo: Multi-Identity Customized Text-to-Video Generation CLIPScore: A Reference-free Evaluation Metric for Image Captioning

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-01T08:45:56.145247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:45:56.145247Z digest=sha256:29a46919f50c33bceede5e827e86456ab32079523bbb1d55e6de6247c49e2ff1

Observation 098e84f8-7205-4552-aed7-c545aed5d8c3 · outbound

This paper cites Gans trained by a two time-scale update rule converge to a local nash equilibrium.

GroupVideo: Multi-Identity Customized Text-to-Video Generation Gans trained by a two time-scale update rule converge to a local nash equilibrium

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-01T08:45:56.233802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:45:56.233802Z digest=sha256:5aa7dfd3876b736e0bbf32b3a141f06f235823aef6179a913496bacd93c39288

Observation 35792e90-8689-4c82-9045-a7cca4861254 · outbound

This paper cites Vbench: Comprehensive benchmark suite for video generative models.

GroupVideo: Multi-Identity Customized Text-to-Video Generation Vbench: Comprehensive benchmark suite for video generative models

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-01T08:45:56.369880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:45:56.369880Z digest=sha256:a8acb70b1ef5ba9e3f3a37244a3a2a0529b014dbc0847e23465f861f022fcebb

Observation 120d1f0b-33e7-4426-bb2f-fa0eaa08e5ef · outbound

This paper cites Wan: Open and Advanced Large-Scale Video Generative Models.

GroupVideo: Multi-Identity Customized Text-to-Video Generation Wan: Open and Advanced Large-Scale Video Generative Models

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-01T08:45:56.546648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:45:56.546648Z digest=sha256:96f6bb7ad215a18457bb8ee6aa5c4626e01de9fbf600960445c98106e0303938

Observation e46cf379-e30f-47d9-8488-fb037c77e821 · outbound

This paper cites SAM 2: Segment Anything in Images and Videos.

GroupVideo: Multi-Identity Customized Text-to-Video Generation SAM 2: Segment Anything in Images and Videos

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-01T08:45:55.760903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:45:55.760903Z digest=sha256:b0152dce4cd29d6fca46bc9fd54bfe41085d4e4fcccb633109444b154ba25bf5

Observation fbf45139-c00d-49f8-bc20-0138b7e78cc1 · outbound

This paper cites PoseGen: In-Context LoRA Finetuning for Pose-Controllable Long Human Video Generation.

GroupVideo: Multi-Identity Customized Text-to-Video Generation PoseGen: In-Context LoRA Finetuning for Pose-Controllable Long Human Video Generation

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-01T08:45:52.437732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:45:52.437732Z digest=sha256:39be00f7d45c6d61c972f79ca68dbf81f45c92c1139a909be3d9c21e8d760c7a

Pith citing papers

No inbound Pith citation observations are available.