REVIEW 5 cited by
JointNet: Extending Text-to-Image Diffusion for Dense Distribution Modeling
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
We introduce JointNet, a novel neural network architecture for modeling the joint distribution of images and an additional dense modality (e.g., depth maps). JointNet is extended from a pre-trained text-to-image diffusion model, where a copy of the original network is created for the new dense modality branch and is densely connected with the RGB branch. The RGB branch is locked during network fine-tuning, which enables efficient learning of the new modality distribution while maintaining the strong generalization ability of the large-scale pre-trained diffusion model. We demonstrate the effectiveness of JointNet by using RGBD diffusion as an example and through extensive experiments, showcasing its applicability in a variety of applications, including joint RGBD generation, dense depth prediction, depth-conditioned image generation, and coherent tile-based 3D panorama generation.
Forward citations
Cited by 5 Pith papers
-
Enhancing In-context Panoramic Generation via Geometric-aware Pretraining
Geometry-aware RGB–depth pretraining plus a 1M multi-task panoramic dataset yields a unified in-context 360° generator with stronger FAED and seam consistency than prior methods.
-
FICGen: Frequency-Inspired Contextual Disentanglement for Layout-driven Degraded Image Generation
A frequency-guided layout-to-image generation framework, FICGen, improves fidelity, layout alignment, and detector trainability on degraded scenes across five benchmarks.
-
Diffusion-Guided Search via Exponential Tilting (DiffTilt): An Application to Falsification of Safety-Critical Systems
DiffTilt is a diffusion-guided falsification method that exponentially tilts a joint distribution over environments and executions to amplify rare safety failures, outperforming conditional sampling and often reducing...
-
Exploring Representation-Aligned Latent Space for Better Generation
Aligning VAE latents with DINOv2 semantic features improves latent diffusion model image generation on ImageNet by about 15% FID.
-
Joint Learning of Depth and Appearance for Portrait Image Animation
A single diffusion model jointly generates portrait RGB images and aligned depth maps, and its fine-tuned variants can estimate depth, edit from depth, relight, and produce audio-driven talking heads with depth.
Discussion (0). Continue with ORCID to comment.