REVIEW 5 cited by
VMem: Consistent Interactive Video Scene Generation with Surfel-Indexed View Memory
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We propose a novel memory module for building video generators capable of interactively exploring environments. Previous approaches have achieved similar results either by out-painting 2D views of a scene while incrementally reconstructing its 3D geometry-which quickly accumulates errors-or by using video generators with a short context window, which struggle to maintain scene coherence over the long term. To address these limitations, we introduce Surfel-Indexed View Memory (VMem), a memory module that remembers past views by indexing them geometrically based on the 3D surface elements (surfels) they have observed. VMem enables efficient retrieval of the most relevant past views when generating new ones. By focusing only on these relevant views, our method produces consistent explorations of imagined environments at a fraction of the computational cost required to use all past views as context. We evaluate our approach on challenging long-term scene synthesis benchmarks and demonstrate superior performance compared to existing methods in maintaining scene coherence and camera control.
Forward citations
Cited by 5 Pith papers
-
Addressable Memory for Video World Models
Video world models can recall revisited scenes far beyond their training horizon by storing compressed memory at fixed in-distribution positions and averaging keys in a rotation-free space.
-
I3DM: Implicit 3D-aware Memory Retrieval and Injection for Consistent Video Scene Generation
An implicit 3D-aware memory mechanism, I3DM, improves revisit consistency and camera control in video scene generation by retrieving historical frames with NVS features and injecting 3D-aligned conditioned latents.
-
Generation Models Know Space: Unleashing Implicit 3D Priors for Scene Understanding
VEGA-3D extracts intermediate spatiotemporal features from a pretrained video diffusion model and fuses them into MLLMs to improve geometric and embodied reasoning without explicit 3D supervision.
-
UCM: Unified Modeling of Camera Control and Memory with Time-aware Positional Encoding Warping for World Models
A video-generation world model that warps positional encodings of memory frames to target viewpoints achieves state-of-the-art long-term consistency and camera control.
-
CustomX: Unified Character, Action, and Scene Customization in Video World Models
AniX generates controllable videos of a user-supplied character performing typed actions inside a user-supplied 3D scene by fine-tuning a pre-trained video generator on small locomotion datasets.
Discussion (0). Continue with ORCID to comment.