REVIEW 9 cited by
LayerPano3D: Layered 3D Panorama for Hyper-Immersive Scene Generation
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
3D immersive scene generation is a challenging yet critical task in computer vision and graphics. A desired virtual 3D scene should 1) exhibit omnidirectional view consistency, and 2) allow for free exploration in complex scene hierarchies. Existing methods either rely on successive scene expansion via inpainting or employ panorama representation to represent large FOV scene environments. However, the generated scene suffers from semantic drift during expansion and is unable to handle occlusion among scene hierarchies. To tackle these challenges, we introduce Layerpano3D, a novel framework for full-view, explorable panoramic 3D scene generation from a single text prompt. Our key insight is to decompose a reference 2D panorama into multiple layers at different depth levels, where each layer reveals the unseen space from the reference views via diffusion prior. Layerpano3D comprises multiple dedicated designs: 1) We introduce a new panorama dataset Upright360, comprising 9k high-quality and upright panorama images, and finetune the advanced Flux model on Upright360 for high-quality, upright and consistent panorama generation. 2) We pioneer the Layered 3D Panorama as underlying representation to manage complex scene hierarchies and lift it into 3D Gaussians to splat detailed 360-degree omnidirectional scenes with unconstrained viewing paths. Extensive experiments demonstrate that our framework generates state-of-the-art 3D panoramic scene in both full view consistency and immersive exploratory experience. We believe that Layerpano3D holds promise for advancing 3D panoramic scene creation with numerous applications.
Forward citations
Cited by 9 Pith papers
-
Look Beyond: Two-Stage Scene View Generation via Panorama and Video Diffusion
Single-image novel view synthesis is decomposed into panorama outpainting plus keyframe-conditioned video diffusion, producing loop-consistent scene tours.
-
ScenePainter: Semantically Consistent Perpetual 3D Scene Generation with Concept Relation Alignment
ScenePainter introduces a SceneConceptGraph that encodes multi-level scene concepts and relations, and aligns an outpainting model with them to reduce semantic drift in perpetual 3D scene generation.
-
DiffDecompose: Layer-Wise Decomposition of Alpha-Composited Images via Diffusion Transformers
DiffDecompose recovers foreground and background layers from alpha-composited images using in-context diffusion with position encoding cloning, trained and evaluated on a new six-task synthetic dataset.
-
HoloTime: Taming Video Diffusion Models for Panoramic 4D Scene Generation
A framework that generates panoramic videos from one image and reconstructs them into 4D Gaussian scenes, with a new panoramic video dataset and improved depth alignment.
-
OmniDrag: Enabling Motion Control for Omnidirectional Image-to-Video Generation
A drag-style motion control method for 360 degree image-to-video generation, built on spherical trajectory estimation and joint fine-tuning of a pretrained video diffusion model.
-
Imagine360: Immersive 360 Video Generation from Perspective Anchor
Imagine360 generates 360-degree equirectangular videos from ordinary perspective video anchors using a dual-branch diffusion model with antipodal attention and elevation-aware handling.
-
TiP4GEN: Text to Immersive Panorama 4D Scene Generation
TiP4GEN generates motion-rich, geometry-consistent 360-degree 4D scenes from a global text prompt plus four local perspective prompts, using a dual-branch video diffusion model with bidirectional cross-attention and a...
-
HunyuanWorld 1.0: Generating Immersive, Explorable, and Interactive 3D Worlds from Words or Pixels
A staged pipeline generates layered, mesh-based 3D worlds from text or images by combining panoramic diffusion, semantic layer decomposition, and video-based expansion.
-
X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models
X-Prompt compresses in-context image examples into a few learned tokens and adds text-description tasks, enabling a Chameleon-style autoregressive model to handle multiple image generation tasks in one framework.
Discussion (0). Continue with ORCID to comment.