Pith. sign in

Paper Citation Record · LEDGER

Emu Video: Factorizing Text-to-Video Generation by Explicit Image Conditioning

As of 16 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 58 inbound Pith citation observations for arXiv:2311.10709.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2311.10709 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 58 of 58 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 58 of 58 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T00:57:28.806572Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T12:36:57.125908Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation ee8b18b6-b8ba-42f6-bc35-867b0c0204fc · inbound

Sora: A Review on Background, Technology, Limitations, and Opportunities of Large Vision Models cites this paper.

Sora: A Review on Background, Technology, Limitations, and Opportunities of Large Vision Models Emu Video: Factorizing Text-to-Video Generation by Explicit Image Conditioning

Reference 196

Resolution
verified exact
arxiv_id, observed 2026-05-13T13:43:11.281861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-13T13:43:11.024069Z digest=sha256:53cf2a56b03b936691f2ab76360e50290f32de70619cd5c9e438137408665145

Observation 3cb89a74-e556-407c-ab97-f64cabe95f79 · inbound

RoboDreamer: Learning Compositional World Models for Robot Imagination cites this paper.

RoboDreamer: Learning Compositional World Models for Robot Imagination Emu Video: Factorizing Text-to-Video Generation by Explicit Image Conditioning

Reference 83

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T20:47:30.309561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-05-15T20:47:30.225116Z digest=sha256:9405109a18b3bbd12c4a2a69ecdac6543a625fa5c0f177115657342e71432234

Observation cb5caf97-71c3-4f11-abcf-dfb6933fb220 · inbound

CAT3D: Create Anything in 3D with Multi-View Diffusion Models cites this paper.

CAT3D: Create Anything in 3D with Multi-View Diffusion Models Emu Video: Factorizing Text-to-Video Generation by Explicit Image Conditioning

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-19T21:29:51.149871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-19T21:29:51.062396Z digest=sha256:b18f37fdb96fed7c3d2fa8e066c0ab711343818b2e21b1d57c50eb7601f71a38

Observation 3dae8dae-6b11-4043-a03c-219153506474 · inbound

Diffusion Models Are Real-Time Game Engines cites this paper.

Diffusion Models Are Real-Time Game Engines Emu Video: Factorizing Text-to-Video Generation by Explicit Image Conditioning

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-05-16T12:04:44.249890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-05-16T12:04:44.162343Z digest=sha256:e01b772f5d99165d8616691c65888fc9119d21a5c56b824904c9f2454157b511

Observation 12b802e2-f91b-4393-b4ce-c66e0fd20f86 · inbound

Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation cites this paper.

Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation Emu Video: Factorizing Text-to-Video Generation by Explicit Image Conditioning

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-15T12:17:01.365911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-15T12:17:01.294466Z digest=sha256:a49b2a331cbced4c628fec98c19d741e765604c6e92d01f35fc9f4dc6ae9c14d

Observation c66e20ca-dc1f-4a25-88e9-b04316bbde4c · inbound

VBench++: Comprehensive and Versatile Benchmark Suite for Video Generative Models cites this paper.

VBench++: Comprehensive and Versatile Benchmark Suite for Video Generative Models Emu Video: Factorizing Text-to-Video Generation by Explicit Image Conditioning

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T16:22:43.898673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:22:43.898673Z digest=sha256:1b151e9fbb39e5e492b0fef24d49d0fa63690253efdcfbd00db8356fff523b10

Observation 0eb5f3c8-f23e-4dfe-b862-5d4e6a6e63e7 · inbound

KFC-W: Generating 3D-Consistent Videos from Unposed Internet Photos cites this paper.

KFC-W: Generating 3D-Consistent Videos from Unposed Internet Photos Emu Video: Factorizing Text-to-Video Generation by Explicit Image Conditioning

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-23T16:48:12.448381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-23T16:47:09.717923Z digest=sha256:e055c43cb3fce0433c85299e6ed0b821097c562d024a9aa47f18441e5fa07325

Observation e5049670-8073-4e81-a47b-7d23aaf1a366 · inbound

Generative Omnimatte: Learning to Decompose Video into Layers cites this paper.

Generative Omnimatte: Learning to Decompose Video into Layers Emu Video: Factorizing Text-to-Video Generation by Explicit Image Conditioning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T12:56:32.823672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:56:32.823672Z digest=sha256:356e722afe95f467dc8aba170b9ef10012a4992c64bcc474df0d43fe45f068d4

Observation 754f10b5-a258-4902-ad67-7b47db1fafff · inbound

$\textit{Revelio}$: Interpreting and leveraging semantic information in diffusion models cites this paper.

$\textit{Revelio}$: Interpreting and leveraging semantic information in diffusion models Emu Video: Factorizing Text-to-Video Generation by Explicit Image Conditioning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T14:23:52.756008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:23:52.756008Z digest=sha256:14831f4b7ca550341bdef12d804457f1497be21de3861a2d8c92382a34f46af5

Observation f0625430-8b40-479d-b21c-e00cb918bcaa · inbound

Towards Precise Scaling Laws for Video Diffusion Transformers cites this paper.

Towards Precise Scaling Laws for Video Diffusion Transformers Emu Video: Factorizing Text-to-Video Generation by Explicit Image Conditioning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T12:58:57.137762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:58:57.137762Z digest=sha256:3d979182170aa3e559d96b1553e08c2d506c991e3feaae9577e43031af4d1744

Observation e8d0ba44-3b67-4b21-a455-a9981a39633c · inbound

Video-Guided Foley Sound Generation with Multimodal Controls cites this paper.

Video-Guided Foley Sound Generation with Multimodal Controls Emu Video: Factorizing Text-to-Video Generation by Explicit Image Conditioning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T12:00:01.084536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:00:01.084536Z digest=sha256:e17d9d14195eb4b562478d5845d4776c3bcf29b3384ee203d3ef5aad8c7fa29e

Observation e2956df3-1384-4350-818a-87ae39263b43 · inbound

I2VControl: Disentangled and Unified Video Motion Synthesis Control cites this paper.

I2VControl: Disentangled and Unified Video Motion Synthesis Control Emu Video: Factorizing Text-to-Video Generation by Explicit Image Conditioning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T12:36:49.208146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:36:49.208146Z digest=sha256:394ee138a8da5aeec4b31b99e6ecbb1f3f4bef7a257e56f789490cbb93813131

Observation fcd1968c-67b0-4456-bd70-fd12bf6b028d · inbound

ROICtrl: Boosting Instance Control for Visual Generation cites this paper.

ROICtrl: Boosting Instance Control for Visual Generation Emu Video: Factorizing Text-to-Video Generation by Explicit Image Conditioning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T11:44:39.856443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:44:39.856443Z digest=sha256:b3f95454f0655afd506fff1e44ebe3ba0c198192631e543a33c9c7d6ae5bc89e

Observation 0936cd5f-c0a4-4279-b6ab-c5135c9b3780 · inbound

ReconDreamer: Crafting World Models for Driving Scene Reconstruction via Online Restoration cites this paper.

ReconDreamer: Crafting World Models for Driving Scene Reconstruction via Online Restoration Emu Video: Factorizing Text-to-Video Generation by Explicit Image Conditioning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T10:09:30.551819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:09:30.551819Z digest=sha256:3735c4375f19d0e0811b4a569486a39e4892b3c5e8bb7dda87dc0cf4fb8118b8

Observation 15fc769f-c302-4013-a2d4-a16049c5069e · inbound

MoTrans: Customized Motion Transfer with Text-driven Video Diffusion Models cites this paper.

MoTrans: Customized Motion Transfer with Text-driven Video Diffusion Models Emu Video: Factorizing Text-to-Video Generation by Explicit Image Conditioning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T04:31:13.377340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:31:13.377340Z digest=sha256:bcf8c22bdea875e486fea5273a4e9b8f6e8312e0c2e5a078e226a7ce5875ecac

Observation 9b9ddeb9-6603-4097-b682-afc21f3c449a · inbound

World-consistent Video Diffusion with Explicit 3D Modeling cites this paper.

World-consistent Video Diffusion with Explicit 3D Modeling Emu Video: Factorizing Text-to-Video Generation by Explicit Image Conditioning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T00:59:22.373007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:59:22.373007Z digest=sha256:52f9806dbab5fc8b23221b7a51acaa26bc948a35a6035e34a3492db84626d865

Observation 71c850b4-401d-4bf5-9ee1-348ffdf0f099 · inbound

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions cites this paper.

ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions Emu Video: Factorizing Text-to-Video Generation by Explicit Image Conditioning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T00:04:08.767955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:04:08.767955Z digest=sha256:7115fa65828b17c3e8d35293ae2c7842894acbe5eefe7b7532389973ac6d0fad

Observation 1886df4d-cb7a-4ed0-b01e-aea208aaa508 · inbound

Motion Prompting: Controlling Video Generation with Motion Trajectories cites this paper.

Motion Prompting: Controlling Video Generation with Motion Trajectories Emu Video: Factorizing Text-to-Video Generation by Explicit Image Conditioning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T23:15:09.578329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:15:09.578329Z digest=sha256:42d9d04653f5b592190fdec1d569f968321b40d4bc5a6229ad01d9d6c2b67da7

Observation 026324f5-ada2-4885-bae7-ef28f7f22124 · inbound

Navigation World Models cites this paper.

Navigation World Models Emu Video: Factorizing Text-to-Video Generation by Explicit Image Conditioning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T22:18:09.768273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:18:09.768273Z digest=sha256:c91252668adefa9498441d386ca595aa70c758040e254bd1151f1f41703c035b

Observation 681ba316-c6ff-45ab-9fc8-5b1b6a570203 · inbound

HunyuanVideo: A Systematic Framework For Large Video Generative Models cites this paper.

HunyuanVideo: A Systematic Framework For Large Video Generative Models Emu Video: Factorizing Text-to-Video Generation by Explicit Image Conditioning

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-23T07:42:43.581649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-23T07:41:58.617477Z digest=sha256:ca2aaedc7494f756ba49b6b875093d3cb6625d7ed886cebc710da6f1b02f7263

Observation ec7def8e-cb2a-45f5-80aa-6db2cea31c5e · inbound

Factorized Video Autoencoders for Efficient Generative Modelling cites this paper.

Factorized Video Autoencoders for Efficient Generative Modelling Emu Video: Factorizing Text-to-Video Generation by Explicit Image Conditioning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:03.687129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:03.687129Z digest=sha256:af0cf2b0f85c08154792cf10954f3e659cee843e4e0df1fc46ab22a4f59a9690

Observation 30e2ec5f-24e8-411b-aac1-02ccbc6817ed · inbound

SynCamMaster: Synchronizing Multi-Camera Video Generation from Diverse Viewpoints cites this paper.

SynCamMaster: Synchronizing Multi-Camera Video Generation from Diverse Viewpoints Emu Video: Factorizing Text-to-Video Generation by Explicit Image Conditioning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T18:36:18.856072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T18:36:18.856072Z digest=sha256:be077cee7a3fbda327da32a0bf8754378c9dd97b6dcdd872ee6d17642c13a5c4

Observation 25442122-0eaa-4598-bfad-f4a1aed5da86 · inbound

DrivingRecon: Large 4D Gaussian Reconstruction Model For Autonomous Driving cites this paper.

DrivingRecon: Large 4D Gaussian Reconstruction Model For Autonomous Driving Emu Video: Factorizing Text-to-Video Generation by Explicit Image Conditioning

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T17:26:29.097766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:26:29.097766Z digest=sha256:856588a169d2c6eaaa133469a1321f638006b252c50b747a904a3da8eaf2b7b5

Observation e1c7108d-86e9-429d-a487-cdda0b005945 · inbound

LinGen: Towards High-Resolution Minute-Length Text-to-Video Generation with Linear Computational Complexity cites this paper.

LinGen: Towards High-Resolution Minute-Length Text-to-Video Generation with Linear Computational Complexity Emu Video: Factorizing Text-to-Video Generation by Explicit Image Conditioning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T16:43:08.238185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:43:08.238185Z digest=sha256:a8c561be23518f92ad7a523f1e9adcd4f808e7859a7a87253e1437a0175ce622

Observation 9c3d6d94-b945-458c-b211-343982de1c5e · inbound

Video Diffusion Transformers are In-Context Learners cites this paper.

Video Diffusion Transformers are In-Context Learners Emu Video: Factorizing Text-to-Video Generation by Explicit Image Conditioning

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T15:40:19.814698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:40:19.814698Z digest=sha256:2e1038e7c0b2c58e9446071cf14d5ef505740c01158d3955e625c537fa7a4578

Observation ed0f69db-665f-483b-b7db-d425581e2d33 · inbound

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control cites this paper.

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Emu Video: Factorizing Text-to-Video Generation by Explicit Image Conditioning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T15:16:34.031763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:16:34.031763Z digest=sha256:0f1dc07419e99f0624e73cc9d5ba14b041ab569f0c39ae52bffe7fd897b8304f

Observation 48ecadf3-075f-4935-9102-38e3f58ebf4e · inbound

MotionBridge: Dynamic Video Inbetweening with Flexible Controls cites this paper.

MotionBridge: Dynamic Video Inbetweening with Flexible Controls Emu Video: Factorizing Text-to-Video Generation by Explicit Image Conditioning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T13:23:35.238481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:23:35.238481Z digest=sha256:850600de42b77f25fdab8c61f484ae4414638824fea41af03b57b1ed6239a7ec

Observation d08671be-7965-4847-96e7-52a5dd099746 · inbound

Flowing from Words to Pixels: A Noise-Free Framework for Cross-Modality Evolution cites this paper.

Flowing from Words to Pixels: A Noise-Free Framework for Cross-Modality Evolution Emu Video: Factorizing Text-to-Video Generation by Explicit Image Conditioning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T11:38:24.477461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:38:24.477461Z digest=sha256:1122f4f3285a2f8a7a4091e60c84098b01b4129b3e734c11c241db54bfdaa125

Observation b742eb0c-580f-4f5b-b34d-c306fa973399 · inbound

VideoMaker: Zero-shot Customized Video Generation with the Inherent Force of Video Diffusion Models cites this paper.

VideoMaker: Zero-shot Customized Video Generation with the Inherent Force of Video Diffusion Models Emu Video: Factorizing Text-to-Video Generation by Explicit Image Conditioning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T00:10:16.462262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:10:16.462262Z digest=sha256:e85bb97a9e694cf227923f84770833cee60fbfdbf8cf4856cdcb7673d2bbf365

Observation bcb5023a-3e47-4494-b766-3ca7526411e2 · inbound

Ingredients: Blending Custom Photos with Video Diffusion Transformers cites this paper.

Ingredients: Blending Custom Photos with Video Diffusion Transformers Emu Video: Factorizing Text-to-Video Generation by Explicit Image Conditioning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T22:25:25.605678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:25:25.605678Z digest=sha256:9c3fd07777a8fc0c56ba88c1314e2d4fd2872b6ac2126edd04759ecdd71e4595

Observation f1465803-95ba-4758-900e-5490b1021b54 · inbound

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation cites this paper.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation Emu Video: Factorizing Text-to-Video Generation by Explicit Image Conditioning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T22:01:56.958288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:01:56.958288Z digest=sha256:e5e0be58309c97bff1245a162a5ad15349df60df6905f40f794bf36295dc5f13

Observation 0fb7dced-ad04-48dd-8b2a-487eaa9b2b66 · inbound

Generative Physical AI in Vision: A Survey cites this paper.

Generative Physical AI in Vision: A Survey Emu Video: Factorizing Text-to-Video Generation by Explicit Image Conditioning

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-10T18:53:00.095026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:53:00.095026Z digest=sha256:d6755adc58353707733e9d6049e7ef192ae4d22cc5ab13589748d04e4da46841

Observation 1c0bdc77-daf9-4d5d-a6c7-979593719cff · inbound

Diffusion Autoencoders are Scalable Image Tokenizers cites this paper.

Diffusion Autoencoders are Scalable Image Tokenizers Emu Video: Factorizing Text-to-Video Generation by Explicit Image Conditioning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-09T22:56:36.059679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T22:56:36.059679Z digest=sha256:3023572d71f20ab45a3c10d684be0ce759dfc9a3005bc358db810377ad370a0f

Observation e7d2e53e-9376-4a61-a257-5282fcc8e6c7 · inbound

Goku: Flow Based Video Generative Foundation Models cites this paper.

Goku: Flow Based Video Generative Foundation Models Emu Video: Factorizing Text-to-Video Generation by Explicit Image Conditioning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-08T21:07:32.265159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T21:07:32.265159Z digest=sha256:3ee67f3fb3c0425ccae6833a92899b790fe94a99e35aa9746aaaa189bde7a145

Observation 73e4a5ff-ba75-45d7-9cc4-40c385a0e376 · inbound

Next Block Prediction: Video Generation via Semi-Autoregressive Modeling cites this paper.

Next Block Prediction: Video Generation via Semi-Autoregressive Modeling Emu Video: Factorizing Text-to-Video Generation by Explicit Image Conditioning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-08T11:51:13.196287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T11:51:13.196287Z digest=sha256:717afe286eacf6b23d8bb28cc5904e2ae9912041155959284903533d6d15d68e

Observation 9411a428-63b1-4420-979f-a9456d8543e8 · inbound

Movie Weaver: Tuning-Free Multi-Concept Video Personalization with Anchored Prompts cites this paper.

Movie Weaver: Tuning-Free Multi-Concept Video Personalization with Anchored Prompts Emu Video: Factorizing Text-to-Video Generation by Explicit Image Conditioning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-09T11:21:20.449710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:21:20.449710Z digest=sha256:7f863c1c2ef9652af9b806854f2e3541f2b3d8e2f497a1c65fa54edbeda41bda

Observation 9622a6ce-725c-4a86-812f-9090e3a01a44 · inbound

RealCam-I2V: Real-World Image-to-Video Generation with Interactive Complex Camera Control cites this paper.

RealCam-I2V: Real-World Image-to-Video Generation with Interactive Complex Camera Control Emu Video: Factorizing Text-to-Video Generation by Explicit Image Conditioning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T19:40:21.199001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T19:40:21.199001Z digest=sha256:f15ea82f05995dc19645606ac17b3766dfbf67bd15b2408aab55c0e7fb4aabbd

Observation ae12a699-9f90-485d-98fc-2c21d24d3013 · inbound

Unified Video Action Model cites this paper.

Unified Video Action Model Emu Video: Factorizing Text-to-Video Generation by Explicit Image Conditioning

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-13T17:50:29.739035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-13T17:50:29.675358Z digest=sha256:56559e96261d76e7c611e0a7cfe5d28bd1b3c06db67ad65c5f82eae9be5f0bba

Observation 2e1445b9-2383-4218-86fc-de697303dfc7 · inbound

Scene-Action Prompt Fusion for Coherent Text-to-Video Storytelling cites this paper.

Scene-Action Prompt Fusion for Coherent Text-to-Video Storytelling Emu Video: Factorizing Text-to-Video Generation by Explicit Image Conditioning

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-23T00:12:17.853066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-23T00:07:39.286486Z digest=sha256:7548b7ba185c55ea5a0be7dc1b5ce75a4d91fa2cf96ba0d9a30af68dec154686

Observation 19c71e63-3ee7-4435-aa98-11c0b565e9ce · inbound

T2S: High-resolution Time Series Generation with Text-to-Series Diffusion Models cites this paper.

T2S: High-resolution Time Series Generation with Text-to-Series Diffusion Models Emu Video: Factorizing Text-to-Video Generation by Explicit Image Conditioning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-16T00:57:28.806572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:57:28.806572Z digest=sha256:3972e03d19f9fcbfc7d76f8490e87d8354ce49addaff022ba2fb773d8743a43e

Observation e049380e-28a3-4eb2-b774-cb6c72371dfd · inbound

Character-Centered Dialogue Generation from Scene-Level Prompts cites this paper.

Character-Centered Dialogue Generation from Scene-Level Prompts Emu Video: Factorizing Text-to-Video Generation by Explicit Image Conditioning

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-22T13:34:53.714162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-22T13:31:43.083678Z digest=sha256:5de86e4c29e320ae8f8732bb8685c2e5c02787abad6a7be8a3aff41e1b4367ec

Observation e17a50bd-73a3-493e-989c-9ce2743ee4c2 · inbound

EPiC: Efficient Video Camera Control Learning with Precise Anchor-Video Guidance cites this paper.

EPiC: Efficient Video Camera Control Learning with Precise Anchor-Video Guidance Emu Video: Factorizing Text-to-Video Generation by Explicit Image Conditioning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T13:28:10.219773Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:28:10.219773Z digest=sha256:5437f811c834e660e9cf0d9b51cb5355754019ef09854d8fec8782ead6f83b73

Observation 70454a36-adb2-49a2-80ba-a43db546785d · inbound

Medical World Model: Generative Simulation of Tumor Evolution for Treatment Planning cites this paper.

Medical World Model: Generative Simulation of Tumor Evolution for Treatment Planning Emu Video: Factorizing Text-to-Video Generation by Explicit Image Conditioning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T11:31:49.473885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:31:49.473885Z digest=sha256:edfb4349d46eb6fc8848221b1c4fa6f344bdbd84244f329ea3c2a73c1424c4cf

Observation 6f2b9577-c8c0-446a-9ca7-5e76c575eab1 · inbound

Edit360: 2D Image Edits to 3D Assets from Any Angle cites this paper.

Edit360: 2D Image Edits to 3D Assets from Any Angle Emu Video: Factorizing Text-to-Video Generation by Explicit Image Conditioning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T04:29:56.147229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:29:56.147229Z digest=sha256:78decaa4a4fd2f2c4656dfd939ae1bc6571f0cfab486a067e7b7224963cceb48

Observation 761b04be-0f23-4ca1-be7d-3d9049da1c52 · inbound

From 2D to 3D Cognition: A Brief Survey of General World Models cites this paper.

From 2D to 3D Cognition: A Brief Survey of General World Models Emu Video: Factorizing Text-to-Video Generation by Explicit Image Conditioning

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T22:57:41.606777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:57:41.606777Z digest=sha256:c0f4496a850507614a9567bfdebc2e29c2d6eeed233cb11f67140150577f014f

Observation 520d1e4f-28c8-4986-9d4f-61ed11c87aea · inbound

Consistent Zero-shot 3D Texture Synthesis Using Geometry-aware Diffusion and Temporal Video Models cites this paper.

Consistent Zero-shot 3D Texture Synthesis Using Geometry-aware Diffusion and Temporal Video Models Emu Video: Factorizing Text-to-Video Generation by Explicit Image Conditioning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T22:41:38.554878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:41:38.554878Z digest=sha256:2c373286a44d177ebeac3509189199e88808cf80ba12b02b2dcaf43b99eb4306

Observation 3f5dd659-901c-45e7-aeb2-0eab12855c8c · inbound

LiON-LoRA: Rethinking LoRA Fusion to Unify Controllable Spatial and Temporal Generation for Video Diffusion cites this paper.

LiON-LoRA: Rethinking LoRA Fusion to Unify Controllable Spatial and Temporal Generation for Video Diffusion Emu Video: Factorizing Text-to-Video Generation by Explicit Image Conditioning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T19:25:28.048695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:25:28.048695Z digest=sha256:d1bb763fde52167e0a56113498b34827262da9692782495a28b93bc3e768994f

Observation aca4e05c-c417-4e7a-8288-279394392d66 · inbound

Waver: Wave Your Way to Lifelike Video Generation cites this paper.

Waver: Wave Your Way to Lifelike Video Generation Emu Video: Factorizing Text-to-Video Generation by Explicit Image Conditioning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T17:45:16.195972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:45:16.195972Z digest=sha256:fc47aa57975c2d634a41d6ac8c6a303141120c55d92497e13e3fd76f0381c990

Observation fb7500c6-0093-41dd-9a46-b97ba90ea362 · inbound

Enhancing Physical Plausibility in Video Generation by Reasoning the Implausibility cites this paper.

Enhancing Physical Plausibility in Video Generation by Reasoning the Implausibility Emu Video: Factorizing Text-to-Video Generation by Explicit Image Conditioning

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-18T12:56:24.364040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-18T12:55:42.679016Z digest=sha256:cb4c211fce862b0b4196cd9efb885acb5878a95518dd504dd362cef9b8c087c1

Observation 397bf284-b29c-40b2-9a61-7039541f728b · inbound

Rolling Sink: Bridging Limited-Horizon Training and Open-Ended Testing in Autoregressive Video Diffusion cites this paper.

Rolling Sink: Bridging Limited-Horizon Training and Open-Ended Testing in Autoregressive Video Diffusion Emu Video: Factorizing Text-to-Video Generation by Explicit Image Conditioning

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T07:07:29.800902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-16T07:02:38.876518Z digest=sha256:4d39b1bc3cba20ba80ff46f4c8346dbc98f5fdd6b2aeb7d91d4f3c28f9c6ae83

Observation b44a179a-fd78-4e80-866a-b537cbe73adc · inbound

Evolution of Video Generative Foundations cites this paper.

Evolution of Video Generative Foundations Emu Video: Factorizing Text-to-Video Generation by Explicit Image Conditioning

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:05:51.865212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-10T18:41:38.616611Z digest=sha256:a2372cbc8de362893e4a2fd8107f51f42d3722f7cf8779ebbd6216b3eb31675b

Observation 8aae5da8-096e-45ae-86a1-e987c5ee7132 · inbound

CityRAG: Stepping Into a City via Spatially-Grounded Video Generation cites this paper.

CityRAG: Stepping Into a City via Spatially-Grounded Video Generation Emu Video: Factorizing Text-to-Video Generation by Explicit Image Conditioning

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T13:11:24.307510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-10T02:10:54.935053Z digest=sha256:4ce35db640d172f54e54639e3b75fe3e504ad057596229da2b47cf060352d490

Observation d2706235-b757-457b-8195-09d8ec60f07d · inbound

DCR: Counterfactual Attractor Guidance for Rare Compositional Generation cites this paper.

DCR: Counterfactual Attractor Guidance for Rare Compositional Generation Emu Video: Factorizing Text-to-Video Generation by Explicit Image Conditioning

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:01:09.626148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-08T13:06:56.964670Z digest=sha256:f5a0e31e5e033d71c94f980a5a2ad6a1e2157c9b65944ef39ff4bf1325be6e82

Observation eb7a8750-1c8c-42cb-8804-28603fa479e0 · inbound

Head Forcing: Long Autoregressive Video Generation via Head Heterogeneity cites this paper.

Head Forcing: Long Autoregressive Video Generation via Head Heterogeneity Emu Video: Factorizing Text-to-Video Generation by Explicit Image Conditioning

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T02:33:32.158143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-15T02:31:48.593354Z digest=sha256:3ac7f1ceef11e0c35fdde4b3e365e80a8155ba2384d0e7bb3c237ca98c65aa15

Observation 150be9d5-dcc2-4da5-ab76-cf4c28c30415 · inbound

StreamingEffect: Real-Time Human-Centric Video Effect Generation cites this paper.

StreamingEffect: Real-Time Human-Centric Video Effect Generation Emu Video: Factorizing Text-to-Video Generation by Explicit Image Conditioning

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-19T20:12:44.842188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-19T20:11:47.359922Z digest=sha256:5c1847f6672cfdcb98a0b5333cee2eee5d883754a01527b421cdbee807bcd742

Observation f3495b01-7d04-45d5-9b88-553a08e7dddf · inbound

Image-to-Video Diffusion: From Foundations to Open Frontiers cites this paper.

Image-to-Video Diffusion: From Foundations to Open Frontiers Emu Video: Factorizing Text-to-Video Generation by Explicit Image Conditioning

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-20T15:08:24.886841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T15:06:02.084336Z digest=sha256:bd6c7c9c6b6e1e37cbb04ce85f88eb868c2ed45c77f7254e63a02fa38645469f

Observation c51df560-1140-455f-b80a-4d35e944fb67 · inbound

Physics in 2-Steps: Locking Motion Priors Before Visual Refinement Erases Them cites this paper.

Physics in 2-Steps: Locking Motion Priors Before Visual Refinement Erases Them Emu Video: Factorizing Text-to-Video Generation by Explicit Image Conditioning

Reference 86

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T12:36:57.127543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-06-28T01:58:00.336605Z digest=sha256:5a3b5e3cc0fbfc912da2d79dcccb0c355790079098a8479498b0168a9f7c8f09

Observation 055ba8a3-e502-43ba-b1ce-7629263d1522 · inbound

TPD: Temporal Prior Decoupling for Text-to-Video Diffusion Models cites this paper.

TPD: Temporal Prior Decoupling for Text-to-Video Diffusion Models Emu Video: Factorizing Text-to-Video Generation by Explicit Image Conditioning

Reference 48

Resolution
unresolved
no resolver link, observed 2026-07-30T23:29:03.450404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T23:29:03.450404Z digest=sha256:21a3493283df83865bcefb4973797e11f3ea73ae275841bd52eea7284a562eae