Pith. sign in

Paper Citation Record · LEDGER

Context as Memory: Scene-Consistent Interactive Long Video Generation with Memory Retrieval

As of 19 August 2026, this Paper Citation Record lists 36 of 36 outbound references and 28 inbound Pith citation observations for arXiv:2506.03141.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.03141 v2

Coverage vector

measured 36 of 36 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:12:17.628244Z

measured 64 of 64 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 28 of 28 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T05:06:52.053244Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T17:58:47.760773Z

Reference resolution

36 of 36 outbound references displayed

  • verified exact0
  • verified fuzzy5
  • unresolved31
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 96500013-c643-49cf-bbf2-11d951ccd5f1 · outbound

This paper cites Vidu: a Highly Consistent, Dynamic and Skilled Text-to-Video Generator with Diffusion Models.

Context as Memory: Scene-Consistent Interactive Long Video Generation with Memory Retrieval Vidu: a Highly Consistent, Dynamic and Skilled Text-to-Video Generator with Diffusion Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T11:12:17.523287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:12:17.523287Z digest=sha256:7729cb889a163336dcf888db8e449ef6ea620cdb6b7baff9246648cc9e86768d

Observation 468cb3db-974f-4d36-a19a-cea2cd2c4080 · outbound

This paper cites Diffusion Forcing: Next-token Prediction Meets Full-Sequence Diffusion.

Context as Memory: Scene-Consistent Interactive Long Video Generation with Memory Retrieval Diffusion Forcing: Next-token Prediction Meets Full-Sequence Diffusion

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T11:12:17.526708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:12:17.526708Z digest=sha256:0106b10e91b9cfd8a5f1b9579243d5c7e8b97c40c6ec8c8a790ea9ecd84c4e74

Observation 5207c38b-a15d-4fc2-9cad-8683c84282c9 · outbound

This paper cites Autoregressive Video Generation without Vector Quantization.

Context as Memory: Scene-Consistent Interactive Long Video Generation with Memory Retrieval Autoregressive Video Generation without Vector Quantization

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T11:12:17.529993Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:12:17.529993Z digest=sha256:f91fcca76dc9683f9446d5915c3e114ddab0f10286e664fba32f56ab6a54a92c

Observation e17f759b-4999-4706-b9c5-0c7da8281d6d · outbound

This paper cites The Matrix: Infinite-Horizon World Generation with Real-Time Moving Control.

Context as Memory: Scene-Consistent Interactive Long Video Generation with Memory Retrieval The Matrix: Infinite-Horizon World Generation with Real-Time Moving Control

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T11:12:17.533297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:12:17.533297Z digest=sha256:df6f55ebbdad3b8c79cc0e032149593f1839caa2378ea012e48a464c1a545c38

Observation c6a3478b-e369-4ba5-a83e-ec5dc458715d · outbound

This paper cites Vista: A Generalizable Driving World Model with High Fidelity and Versatile Controllability.

Context as Memory: Scene-Consistent Interactive Long Video Generation with Memory Retrieval Vista: A Generalizable Driving World Model with High Fidelity and Versatile Controllability

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T11:12:17.536837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:12:17.536837Z digest=sha256:eaadccc394ce14b19d0013f92db9c0a6613f3687ba90c9d41505a5d812bc2fe4

Observation 39b33f8d-ef86-4377-880c-95334fc5b127 · outbound

This paper cites Long-Context Autoregressive Video Modeling with Next-Frame Prediction.

Context as Memory: Scene-Consistent Interactive Long Video Generation with Memory Retrieval Long-Context Autoregressive Video Modeling with Next-Frame Prediction

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T11:12:17.540549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:12:17.540549Z digest=sha256:5f98c9d89a85ad4169daa408de59242bed636d3d01099e55d6db52b1211cf9fa

Observation 94f46396-c3a3-488a-8088-5f17cf2bac18 · outbound

This paper cites Long Context Tuning for Video Generation.

Context as Memory: Scene-Consistent Interactive Long Video Generation with Memory Retrieval Long Context Tuning for Video Generation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T11:12:17.543716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:12:17.543716Z digest=sha256:0bbccbf0b39a6d3f3815bf7add55dabf06a45726a04334bfdf20ac6bec3c5011

Observation 4968c0f2-f214-4591-a6d7-c0a464770805 · outbound

This paper cites CameraCtrl: Enabling Camera Control for Text-to-Video Generation.

Context as Memory: Scene-Consistent Interactive Long Video Generation with Memory Retrieval CameraCtrl: Enabling Camera Control for Text-to-Video Generation

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T11:12:17.546924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:12:17.546924Z digest=sha256:3c29954ebc3660123a0269fb90a135197940c5444c99ae9f9caeeab3f26cfe49

Observation 8bc79a47-42df-46fc-b00b-d217d8829960 · outbound

This paper cites Nature 638, 8051 (2025), 656–663.

Context as Memory: Scene-Consistent Interactive Long Video Generation with Memory Retrieval Nature 638, 8051 (2025), 656–663

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:12:18.190896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T11:12:17.559224Z digest=sha256:3da0e8bd2c2fb5ba1f3bd7df6ba855f0e1da675bdeb8cb6ab1e648cc99c25b56

Observation 1557b2fb-9902-45a5-bb9e-73018f3f936e · outbound

This paper cites VideoPoet: A Large Language Model for Zero-Shot Video Generation.

Context as Memory: Scene-Consistent Interactive Long Video Generation with Memory Retrieval VideoPoet: A Large Language Model for Zero-Shot Video Generation

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T11:12:17.562082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:12:17.562082Z digest=sha256:487e299a1067c5b824e2b08c07eed37ef7c8791f5776af22f1ab55f134d930bd

Observation f164a619-6614-4125-8079-d425f15bb81e · outbound

This paper cites HunyuanVideo: A Systematic Framework For Large Video Generative Models.

Context as Memory: Scene-Consistent Interactive Long Video Generation with Memory Retrieval HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T11:12:17.565245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:12:17.565245Z digest=sha256:60fa562d01823f9f096f465ba00d9a01a4d3aaa9f72acbca54a6d20fd8c041ac

Observation 773edda9-b519-40e9-bd44-3fbd4b748581 · outbound

This paper cites Autoregressive Image Generation without Vector Quantization.

Context as Memory: Scene-Consistent Interactive Long Video Generation with Memory Retrieval Autoregressive Image Generation without Vector Quantization

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T11:12:17.568138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:12:17.568138Z digest=sha256:b5fb081f6ff5286ce81dade308affd1b3e1dd4e330fddb26e2024776fa02fe00

Observation bebabb30-3ed8-4b82-8837-93653fb3088a · outbound

This paper cites Flow Matching for Generative Modeling.

Context as Memory: Scene-Consistent Interactive Long Video Generation with Memory Retrieval Flow Matching for Generative Modeling

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T11:12:17.571141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:12:17.571141Z digest=sha256:48b79720743e1e98519e2327d72c51a77b0dc9e386e33844ef0299bcc57dc0ec

Observation 3034e4c1-579f-4738-8759-035a7747bfac · outbound

This paper cites Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow.

Context as Memory: Scene-Consistent Interactive Long Video Generation with Memory Retrieval Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T11:12:17.574359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:12:17.574359Z digest=sha256:325fed21150a4767858783dd0edf24ae51494cdc3b287596be8b6aef5914f607

Observation f1bdfee4-7e77-4aaf-bde2-6aa6c01532dc · outbound

This paper cites You See it, You Got it: Learning 3D Creation on Pose-Free Videos at Scale.

Context as Memory: Scene-Consistent Interactive Long Video Generation with Memory Retrieval You See it, You Got it: Learning 3D Creation on Pose-Free Videos at Scale

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T11:12:17.577727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:12:17.577727Z digest=sha256:0e06d4bd898fcbd610368dcc723c25e14b8c407d378dbf822cf341cc18ebbbbe

Observation ad3d3edd-91d0-4cee-b070-a8c5e36db9c5 · outbound

This paper cites WorldSimBench: Towards Video Generation Models as World Simulators.

Context as Memory: Scene-Consistent Interactive Long Video Generation with Memory Retrieval WorldSimBench: Towards Video Generation Models as World Simulators

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T11:12:17.580891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:12:17.580891Z digest=sha256:b0123339f72767553f46a0091ca7a127c40516c95c55f779f36a28cb02deb578

Observation 8d5e1cb6-f92b-4d6b-b121-90883bf3ce44 · outbound

This paper cites GEN3C: 3D-Informed World-Consistent Video Generation with Precise Camera Control.

Context as Memory: Scene-Consistent Interactive Long Video Generation with Memory Retrieval GEN3C: 3D-Informed World-Consistent Video Generation with Precise Camera Control

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T11:12:17.584083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:12:17.584083Z digest=sha256:e893b0e3eb80209088f2ccf6c7bca889b64017fb629169458461cfbb75e19425

Observation 9ecc685f-7530-4d00-b580-cb169fdbe5cd · outbound

This paper cites GAIA-2: A Controllable Multi-View Generative World Model for Autonomous Driving.

Context as Memory: Scene-Consistent Interactive Long Video Generation with Memory Retrieval GAIA-2: A Controllable Multi-View Generative World Model for Autonomous Driving

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T11:12:17.587192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:12:17.587192Z digest=sha256:af75ef6e13e315dc8e63d03686aafef8a6a41e3e39f689cd219ac1a812a42141

Observation dda9aaae-1482-42cc-a9f4-eda9127e58af · outbound

This paper cites History-Guided Video Diffusion.

Context as Memory: Scene-Consistent Interactive Long Video Generation with Memory Retrieval History-Guided Video Diffusion

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T11:12:17.590181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:12:17.590181Z digest=sha256:078d0316c4ec548d3f00d46cbf93d143a6b230811043ad475207dcfb3d3642b1

Observation b2d53f0e-589e-42a4-9d88-b4cb0224f1ca · outbound

This paper cites Neurocomputing 568 (2024), 127063.

Context as Memory: Scene-Consistent Interactive Long Video Generation with Memory Retrieval Neurocomputing 568 (2024), 127063

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:12:18.116788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T11:12:17.599097Z digest=sha256:442d641e148019519ddf38dcab26a5ef55b01570e24d33dc596f9063e0a54791

Observation a001fb6c-167b-497d-abd4-6a2425a18c5d · outbound

This paper cites Diffusion Models Are Real-Time Game Engines.

Context as Memory: Scene-Consistent Interactive Long Video Generation with Memory Retrieval Diffusion Models Are Real-Time Game Engines

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T11:12:17.602243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:12:17.602243Z digest=sha256:d021f8094661721d09649e38dcf339216c4a8c4a3c34a81f1f4ef84a356e4973

Observation 3954512d-2626-4f56-b70d-85bff10e5ceb · outbound

This paper cites Wan: Open and Advanced Large-Scale Video Generative Models.

Context as Memory: Scene-Consistent Interactive Long Video Generation with Memory Retrieval Wan: Open and Advanced Large-Scale Video Generative Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T11:12:17.605426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:12:17.605426Z digest=sha256:c126e6fa703aa4c0c7f37ae07b52d5cbb982150ef12b445bf318154e74cf56a3

Observation 51c0d825-0d6b-49cc-b3b1-3d36c4abce7b · outbound

This paper cites arXiv preprint arXiv:2504.12369 (2025).

Context as Memory: Scene-Consistent Interactive Long Video Generation with Memory Retrieval arXiv preprint arXiv:2504.12369 (2025)

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T11:12:17.608684Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:12:17.608684Z digest=sha256:f5b5b1a95d2f10d2a316a1f32b084e9c4f94dcff988ff91e02af8306782ae51e

Observation 63f2037b-dfa4-421c-8dff-7cc6ac4c35ff · outbound

This paper cites DynamiCrafter: Animating Open-domain Images with Video Diffusion Priors.

Context as Memory: Scene-Consistent Interactive Long Video Generation with Memory Retrieval DynamiCrafter: Animating Open-domain Images with Video Diffusion Priors

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T11:12:17.611750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:12:17.611750Z digest=sha256:9f04e1e3f700236a6397776af17147a74e659ca394360371a55990b986a21dd7

Observation 070cf774-429d-4527-8ac3-3d3ac9f4e7b8 · outbound

This paper cites VideoGPT: Video Generation using VQ-VAE and Transformers.

Context as Memory: Scene-Consistent Interactive Long Video Generation with Memory Retrieval VideoGPT: Video Generation using VQ-VAE and Transformers

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T11:12:17.615226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:12:17.615226Z digest=sha256:270b0f63bf8ac21f8134c899631bcd59db2d496de128e10386cabea5be230ddb

Observation 380ebcbe-bcee-4c43-920b-fcc981c7c491 · outbound

This paper cites Learning Interactive Real-World Simulators.

Context as Memory: Scene-Consistent Interactive Long Video Generation with Memory Retrieval Learning Interactive Real-World Simulators

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T11:12:17.618634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:12:17.618634Z digest=sha256:712e80652f32e67e006f3eff5f52084da277247e1888461e58063eccd1be32de

Observation f7ac7720-375c-4570-8838-c8086d939f8b · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

Context as Memory: Scene-Consistent Interactive Long Video Generation with Memory Retrieval MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T11:12:17.621640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:12:17.621640Z digest=sha256:72e8b2f0e83169b2635d2e93a7ffa5837d4d5ce77117df25b90815ff87ce2b6e

Observation f2e97719-1b60-47d7-b361-26030eecb2f6 · outbound

This paper cites arXiv preprint arXiv:2504.12626 (2025).

Context as Memory: Scene-Consistent Interactive Long Video Generation with Memory Retrieval arXiv preprint arXiv:2504.12626 (2025)

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T11:12:17.625179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:12:17.625179Z digest=sha256:3922968c39b77bad724029a63902aa229d1865a27cda6f6c33b7120bc6bdcc21

Observation f6a2a3d5-53e1-45bd-b213-d0d9de257a71 · outbound

This paper cites Stereo Magnification: Learning View Synthesis using Multiplane Images.

Context as Memory: Scene-Consistent Interactive Long Video Generation with Memory Retrieval Stereo Magnification: Learning View Synthesis using Multiplane Images

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-07T11:12:17.628244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:12:17.628244Z digest=sha256:9de262f9759c60b2e2e7f7d65d379b811676dfa1557a74f950c83bd336f90bf6

Observation a5be92bd-a567-4562-8033-355fce2ec885 · outbound

This paper cites Advances in neural information processing systems (2019).

Context as Memory: Scene-Consistent Interactive Long Video Generation with Memory Retrieval Advances in neural information processing systems (2019)

Reference 2019

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:12:18.149202Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T11:12:17.593059Z digest=sha256:23cc16f6596f78eb6409132a733f0b275e1544a6aa75ce45701ab3ff13137395

Observation 77ac232f-da3f-453a-afa3-9b1dbe1500d8 · outbound

This paper cites Advances in neural information processing systems (2020).

Context as Memory: Scene-Consistent Interactive Long Video Generation with Memory Retrieval Advances in neural information processing systems (2020)

Reference 2020

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:12:18.253056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T11:12:17.550328Z digest=sha256:fe8aa5c9534dd12acdb3174b88d190f26bbe2ff972a7ffff2156a7239ef3a7c9

Observation c5935c72-52a4-4402-a5b3-4dd5f6b98ae5 · outbound

This paper cites International Conference on Learning Representations (2021).

Context as Memory: Scene-Consistent Interactive Long Video Generation with Memory Retrieval International Conference on Learning Representations (2021)

Reference 2021

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:12:18.132297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T11:12:17.595978Z digest=sha256:ff332c8884eb663cc3ad1ed4c62d3b9a4a9eab66ec97308dd7751b359ed8855a

Observation a511b16c-3105-449a-bf73-f9b0187af0c6 · outbound

This paper cites Classifier-Free Diffusion Guidance.

Context as Memory: Scene-Consistent Interactive Long Video Generation with Memory Retrieval Classifier-Free Diffusion Guidance

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-07T11:12:17.553177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:12:17.553177Z digest=sha256:3534544641231c2d974fbab3b06737dcd9811e71cb7b88528a2d2e64752fb93b

Observation a4be7c35-fe9b-4b9b-9f88-0027e84dbf49 · outbound

This paper cites GAIA-1: A Generative World Model for Autonomous Driving.

Context as Memory: Scene-Consistent Interactive Long Video Generation with Memory Retrieval GAIA-1: A Generative World Model for Autonomous Driving

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T11:12:17.556236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:12:17.556236Z digest=sha256:85f7c6411e892abf524af2c7064e27e353665e9df9135fdad6c3862e56737892

Observation 4c299b7d-e5e4-4399-8a2c-bc9d2752a7ab · outbound

This paper cites SynCamMaster: Synchronizing Multi-Camera Video Generation from Diverse Viewpoints.

Context as Memory: Scene-Consistent Interactive Long Video Generation with Memory Retrieval SynCamMaster: Synchronizing Multi-Camera Video Generation from Diverse Viewpoints

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T11:12:17.519375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:12:17.519375Z digest=sha256:f56e2acc9fba8a7f01b91fb5cde62de2ad9d59f80c5029c1313a67be483c6cea

Observation 477ec2b4-1a64-493d-8dd1-d564b4d5e006 · outbound

This paper cites ReCamMaster: Camera-Controlled Generative Rendering from A Single Video.

Context as Memory: Scene-Consistent Interactive Long Video Generation with Memory Retrieval ReCamMaster: Camera-Controlled Generative Rendering from A Single Video

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T11:12:17.515261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:12:17.515261Z digest=sha256:f438b7a3c73cf845c99052f6ef2adbedd0fb297b23c4d878ca695789d94dd653

Pith citing papers

Observation 71645466-37ac-464f-bbfc-e2b05e1791ba · inbound

CDP: Towards Robust Autoregressive Visuomotor Policy Learning via Causal Diffusion cites this paper.

CDP: Towards Robust Autoregressive Visuomotor Policy Learning via Causal Diffusion Context as Memory: Scene-Consistent Interactive Long Video Generation with Memory Retrieval

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T00:15:37.627688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:15:37.627688Z digest=sha256:035ae5144998bc92fc2c0c4b7254c4035cfcafcecfe32fcec7d349bf2952975a

Observation ce1a6f6f-7ba5-45e6-896e-8926d8ebb6d2 · inbound

Recurrent Autoregressive Diffusion: Global Memory Meets Local Attention cites this paper.

Recurrent Autoregressive Diffusion: Global Memory Meets Local Attention Context as Memory: Scene-Consistent Interactive Long Video Generation with Memory Retrieval

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-03T21:59:33.245663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T21:59:33.245663Z digest=sha256:a3b0a3dabf25e7cbef9abe7307a00d063210b137639bcfddf666ca9075aab3a1

Observation d1ff0185-3f70-425f-a363-857e497202a6 · inbound

Target-Bench: Can Video World Models Achieve Mapless Path Planning with Semantic Targets? cites this paper.

Target-Bench: Can Video World Models Achieve Mapless Path Planning with Semantic Targets? Context as Memory: Scene-Consistent Interactive Long Video Generation with Memory Retrieval

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:02:04.221770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-17T20:00:20.895672Z digest=sha256:9d166f79b6bd466ab36294930d47d82a65b28cf688a5027058a5a12b2c9680f6

Observation dde0d32c-cd0c-4c25-a363-3914c8c1b8c7 · inbound

Vision-Language Memory for Spatial Reasoning cites this paper.

Vision-Language Memory for Spatial Reasoning Context as Memory: Scene-Consistent Interactive Long Video Generation with Memory Retrieval

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-03T20:15:36.183372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:15:36.183372Z digest=sha256:b97ae46f1993dba1e30a0b4759891067a8c1e410e852c957ea2d5228bfbfbb2d

Observation 7d25cc99-4b09-45c2-b852-a67e1d1a6801 · inbound

WorldPlay: Towards Long-Term Geometric Consistency for Real-Time Interactive World Modeling cites this paper.

WorldPlay: Towards Long-Term Geometric Consistency for Real-Time Interactive World Modeling Context as Memory: Scene-Consistent Interactive Long Video Generation with Memory Retrieval

Reference 74

Resolution
verified exact
arxiv_id, observed 2026-05-15T14:29:56.497402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-15T14:29:56.348733Z digest=sha256:40d0c65a58dd4b9487da2e07e735a08b70363b2301a74201a632dba1684190f3

Observation be1124f5-f6ad-4afc-b78c-e2f9f667a430 · inbound

WorldPlay: Towards Long-Term Geometric Consistency for Real-Time Interactive World Modeling cites this paper.

WorldPlay: Towards Long-Term Geometric Consistency for Real-Time Interactive World Modeling Context as Memory: Scene-Consistent Interactive Long Video Generation with Memory Retrieval

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-03T16:11:18.301098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:11:18.301098Z digest=sha256:ec84705166d44c07dc401fc54452ae435d4e4bdb471d470bf8f57d949301f3b9

Observation 0b4530a8-a06c-45de-88d2-bb72ea16e39d · inbound

CustomX: Unified Character, Action, and Scene Customization in Video World Models cites this paper.

CustomX: Unified Character, Action, and Scene Customization in Video World Models Context as Memory: Scene-Consistent Interactive Long Video Generation with Memory Retrieval

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-03T15:29:00.649927Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T15:29:00.649927Z digest=sha256:8ac018d7849031ff5a13ca02c15d7ac7e54a6110ba340ce6ce20de39eb26762f

Observation 6fe4146b-c908-4b1e-93fc-c505b9f41773 · inbound

UCM: Unified Modeling of Camera Control and Memory with Time-aware Positional Encoding Warping for World Models cites this paper.

UCM: Unified Modeling of Camera Control and Memory with Time-aware Positional Encoding Warping for World Models Context as Memory: Scene-Consistent Interactive Long Video Generation with Memory Retrieval

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-02T20:36:13.810793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:36:13.810793Z digest=sha256:bafeb03a9b728ddaa1ce3834155dd5dfcfa1206727a3c6e5906eda5d5cba9717

Observation 08ab7c33-aa47-4062-9ef4-dc1d07f78abd · inbound

SuperLocalMemory V3.3: The Living Brain -- Biologically-Inspired Forgetting, Cognitive Quantization, and Multi-Channel Retrieval for Zero-LLM Agent Memory Systems cites this paper.

SuperLocalMemory V3.3: The Living Brain -- Biologically-Inspired Forgetting, Cognitive Quantization, and Multi-Channel Retrieval for Zero-LLM Agent Memory Systems Context as Memory: Scene-Consistent Interactive Long Video Generation with Memory Retrieval

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:30:50.645013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-10T19:48:34.606351Z digest=sha256:779ca0806a73e0386a35b10c1b648fec268f1526316296e9920446df8034a688

Observation f841b663-8e70-4ab6-be4e-e21884641ba2 · inbound

Evolution of Video Generative Foundations cites this paper.

Evolution of Video Generative Foundations Context as Memory: Scene-Consistent Interactive Long Video Generation with Memory Retrieval

Reference 281

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:05:51.270209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-10T18:41:38.616611Z digest=sha256:778b4d9560dbfee87e268b61756a82f3870677ed5e3743da831889e1142372f8

Observation 9845b9fb-5bfd-40eb-9c42-9c1f1c7f4afb · inbound

Scal3R: Scalable Test-Time Training for Large-Scale 3D Reconstruction cites this paper.

Scal3R: Scalable Test-Time Training for Large-Scale 3D Reconstruction Context as Memory: Scene-Consistent Interactive Long Video Generation with Memory Retrieval

Reference 94

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:50:59.685157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-10T16:56:14.380078Z digest=sha256:09d26dc18591042c8d884f4e3f22edfedbca7126dd3a67c877f79a834626f6b7

Observation 9bd845dc-2172-4866-a22a-5d35380d1722 · inbound

HY-World 2.0: A Multi-Modal World Model for Reconstructing, Generating, and Simulating 3D Worlds cites this paper.

HY-World 2.0: A Multi-Modal World Model for Reconstructing, Generating, and Simulating 3D Worlds Context as Memory: Scene-Consistent Interactive Long Video Generation with Memory Retrieval

Reference 84

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:45:27.154100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-10T13:45:24.961208Z digest=sha256:f73d6099e1ba5f75776ac27caa33c781061694f8fdcd68822fc5588f378459c7

Observation f4f4daf0-08e0-4e10-9967-4d9d4c0dbb9d · inbound

MultiWorld: Scalable Multi-Agent Multi-View Video World Models cites this paper.

MultiWorld: Scalable Multi-Agent Multi-View Video World Models Context as Memory: Scene-Consistent Interactive Long Video Generation with Memory Retrieval

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-05-10T10:09:08.126250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-10T05:06:11.514186Z digest=sha256:56b34eee493660debc97fbfe16623625831339201bc7fc783013b1b5f3070aa7

Observation 036bfba9-9d4c-4994-835c-98b4c678c78f · inbound

Head Forcing: Long Autoregressive Video Generation via Head Heterogeneity cites this paper.

Head Forcing: Long Autoregressive Video Generation via Head Heterogeneity Context as Memory: Scene-Consistent Interactive Long Video Generation with Memory Retrieval

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-05-15T02:33:32.128967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-15T02:31:48.593354Z digest=sha256:8612a7ffc665e23b9e45d0186cf67f8791b5ab20299474a732b7c45a306a3a9d

Observation 90e471ac-c942-4b50-b2fc-8f0d726cf0c4 · inbound

Gamma-World: Generative Multi-Agent World Modeling Beyond Two Players cites this paper.

Gamma-World: Generative Multi-Agent World Modeling Beyond Two Players Context as Memory: Scene-Consistent Interactive Long Video Generation with Memory Retrieval

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-06-29T13:53:29.081619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-29T13:34:03.020051Z digest=sha256:ea5d833dc603bbc3b1db2e42d5b21293094d3024dcbcfebfb618f37c861d0812

Observation 11b18680-c829-415d-8f2e-fd1bbfbe7f9a · inbound

Light Interaction: Training-Free Inference Acceleration for Interactive Video World Models cites this paper.

Light Interaction: Training-Free Inference Acceleration for Interactive Video World Models Context as Memory: Scene-Consistent Interactive Long Video Generation with Memory Retrieval

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-07-01T19:26:00.130065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-28T22:43:32.749425Z digest=sha256:a4b22c531c519a6b470a195513901e47013095c454faa89f836bae08ad44b43a

Observation 32f981f9-c8bb-43fa-8478-fc258d0b67fd · inbound

DecMem: Towards Minute-Long Consistent World Generation with Decoupled Memory cites this paper.

DecMem: Towards Minute-Long Consistent World Generation with Decoupled Memory Context as Memory: Scene-Consistent Interactive Long Video Generation with Memory Retrieval

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-06-28T22:42:46.789949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-28T22:38:10.473453Z digest=sha256:ae0f1a18d17e79c9a06ca7e3f9528882c90c5a162a358b3e17db6e981646a784

Observation 8a09830f-5302-4e12-9481-88fd0609a359 · inbound

MetaWorld: Scaling Multi-Agent Video World Model from Single-view Video Data cites this paper.

MetaWorld: Scaling Multi-Agent Video World Model from Single-view Video Data Context as Memory: Scene-Consistent Interactive Long Video Generation with Memory Retrieval

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-07-01T22:56:20.240121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-28T14:52:30.406683Z digest=sha256:e28262b872471aa9aab680bf804ff84515239a39d572a4b4955aa053d72690a6

Observation 31c4f39c-d9b9-47c6-917c-871c65fc4759 · inbound

Echo-Memory: A Controlled Study of Memory in Action World Models cites this paper.

Echo-Memory: A Controlled Study of Memory in Action World Models Context as Memory: Scene-Consistent Interactive Long Video Generation with Memory Retrieval

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-07-03T00:57:29.617097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-27T16:58:37.552036Z digest=sha256:f06c86f3e23571f16a8e028086a70af29ea62f6950f6a37a8ff5d1d5670fe1e8

Observation 99edf4a0-a104-4d72-b421-d62804aad845 · inbound

Latent Spatial Memory for Video World Models cites this paper.

Latent Spatial Memory for Video World Models Context as Memory: Scene-Consistent Interactive Long Video Generation with Memory Retrieval

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-07-03T01:07:30.554029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-27T16:47:42.761342Z digest=sha256:284ef2842d82572cd8cd96d4c6a52c4be3aa2e3e588dcd6035cf34ac5aee2f96

Observation 1086bb37-ba36-442f-93c6-7b29e02e7aba · inbound

FadeMem: Distance-Aware Memory Consolidation for Autoregressive Video Diffusion cites this paper.

FadeMem: Distance-Aware Memory Consolidation for Autoregressive Video Diffusion Context as Memory: Scene-Consistent Interactive Long Video Generation with Memory Retrieval

Reference 82

Resolution
verified exact
arxiv_id, observed 2026-07-03T04:37:37.482634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-06-27T13:44:44.550033Z digest=sha256:af26843ab022c4989bc7478f4c9357bc343a6f9ef7e9d8e94ea88dcf5022f4bd

Observation 6d2f390b-73ae-473a-8785-784f244508a6 · inbound

PermaVid: Consistent Video Generation Across Edits via Disentangled Context Memory cites this paper.

PermaVid: Consistent Video Generation Across Edits via Disentangled Context Memory Context as Memory: Scene-Consistent Interactive Long Video Generation with Memory Retrieval

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-07-03T17:58:47.762544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-27T03:23:58.455778Z digest=sha256:cfe8bb4775b65acafe0fe19cd99459ee9c6faa6cc8365640b8cc163a65e6ec9f

Observation 7bf165a4-308a-412f-8e7b-143ac53ad726 · inbound

MemLearner: Learning to Query Context memory for Video World Models cites this paper.

MemLearner: Learning to Query Context memory for Video World Models Context as Memory: Scene-Consistent Interactive Long Video Generation with Memory Retrieval

Reference 66

Resolution
verified exact
arxiv_id, observed 2026-07-01T10:25:41.877273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-01T05:30:56.140465Z digest=sha256:d60f21c3afb6e89052ccce729068ac417be09442b8599b28f6945cbf029dcfe6

Observation 8e84fb0f-18ac-4197-8c51-0b5170118c76 · inbound

WorldDirector: Building Controllable World Simulators with Persistent Dynamic Memory cites this paper.

WorldDirector: Building Controllable World Simulators with Persistent Dynamic Memory Context as Memory: Scene-Consistent Interactive Long Video Generation with Memory Retrieval

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-07-03T14:28:31.047823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-03T14:28:24.922144Z digest=sha256:6aef71fa3c1fa5dd420daabd05fb68d7629eac1623e3790fd599dca7baba4d70

Observation d03d8117-f012-4c67-ad39-cdec4511a596 · inbound

SlotMem: Character-Addressable Internal Memory for Narrative Long Video Generation cites this paper.

SlotMem: Character-Addressable Internal Memory for Narrative Long Video Generation Context as Memory: Scene-Consistent Interactive Long Video Generation with Memory Retrieval

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-01T22:25:26.677846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T22:25:26.677846Z digest=sha256:6951ec1de3f43dfc757940e38c3e563c1db6e9c9a39e2c126fc668066640c4d3

Observation 8b85fad5-f630-4ea3-bbe1-05ccf3774b37 · inbound

StreamHOI: Interaction-aware Temporal Memory Adaptation for Streaming HOI Video Generation cites this paper.

StreamHOI: Interaction-aware Temporal Memory Adaptation for Streaming HOI Video Generation Context as Memory: Scene-Consistent Interactive Long Video Generation with Memory Retrieval

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-01T10:37:56.086456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:37:56.086456Z digest=sha256:1505b7621b380d4bb209b60f3aafe78bbe164a08010d53f74b10574e903dba65

Observation b9f5f915-f899-4ff6-bc6a-ddad2266ea8e · inbound

Wonder: Video World Model Done Better cites this paper.

Wonder: Video World Model Done Better Context as Memory: Scene-Consistent Interactive Long Video Generation with Memory Retrieval

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-01T00:52:27.986020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:52:27.986020Z digest=sha256:75feae4eda0fb9402ea10968365150d048180debd1064d24131397445c62ff36

Observation bb90fd8c-9d3c-4522-8872-b1243849f5be · inbound

Addressable Memory for Video World Models cites this paper.

Addressable Memory for Video World Models Context as Memory: Scene-Consistent Interactive Long Video Generation with Memory Retrieval

Reference 110

Resolution
unresolved
no resolver link, observed 2026-08-10T05:06:52.053244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T05:06:52.053244Z digest=sha256:b6fd027fcabd51ccd791d78306cf50a90792b130ff9c9f4765b9e05bce826959